nomic-embed-text-v1.5
Release in the Nomic Embed family · version nomic-embed-text-v1.5
A text embedding model that updates nomic-embed-text-v1 with Matryoshka representation learning, so its 768-dimension embeddings can be shortened to as few as 64 dimensions, or binarized, with a small loss in quality. Released in February 2024, it follows v1's long-context design, which Nomic describes as supporting 8,192 tokens.213
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Model hub: GGUF build (external site: huggingface.co)
- Release notes: v1.5 announcement (external site: nomic.ai)
- Paper: Nomic Embed technical report (arXiv 2402.01613) (external site: arxiv.org)
- Repository: contrastors training code (external site: github.com)
Availability and license
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | 1 |
| Inference codeIs code for running the model published? | Public | The card documents use with Sentence Transformers, Transformers, and Transformers.js, and says trust_remote_code is no longer needed from Transformers 5.5.0 and Sentence Transformers 5.3.0.1 |
| Training codeIs the code used to train the model published? | Public | Training code is in the contrastors repository.25 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Public | The card says the training data is released in full. The contrastors README explains how to download it from Nomic's storage after creating a Nomic Atlas account.125 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The card describes the stages: start from nomic-bert-2048, contrastive training on weakly related text pairs (for example forum question-answer pairs and review title-body pairs), then fine-tuning on labeled data with hard-example mining; the technical report gives details.1 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Unknown | Not assessed. |
What it is useful for
Embedding documents and queries for retrieval-augmented generation and search, and texts for clustering and classification, selected with the task prefixes search_document, search_query, clustering, and classification. Nomic's vision model nomic-embed-vision-v1.5 is aligned to its embedding space, so text and image embeddings can be compared.1
Run and use notes
- The card shows applying layer normalization before truncating to a smaller Matryoshka dimension and then normalizing, and explains how to extend the sequence length past 2,048 tokens with dynamic RoPE scaling.1
Organization context
Provenance and derivatives
Trained by Nomic from nomic-bert-2048, a long-context BERT variant Nomic developed for Nomic Embed, as an update to nomic-embed-text-v1.13
- Derived from: nomic-bert-2048 (external site: huggingface.co) — Nomic's long-context BERT starting model.
- Derived from: nomic-embed-text-v1 (external site: huggingface.co) — Previous release; v1.5 adds Matryoshka representation learning.
Other releases in the Nomic Embed family
- nomic-embed-text-v2-moeModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.