nomic-embed-text-v2-moe
Release in the Nomic Embed family · version nomic-embed-text-v2-moe
A multilingual text embedding model that uses a mixture-of-experts design (8 experts, top-2 routing), with 475M total and 305M active parameters. It covers about 100 languages, produces 768-dimensional embeddings that can be truncated to 256 through Matryoshka training, and accepts inputs of up to 512 tokens. Nomic released it in February 2025.123
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Model hub: GGUF build (external site: huggingface.co)
- Paper: Technical report (arXiv 2502.07972) (external site: arxiv.org)
- Repository: contrastors training code (external site: github.com)
Availability and license
Overall availability
Downloadable from Hugging Face without gating.1
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | 1 |
| Inference codeIs code for running the model published? | Public | The card documents use with Sentence Transformers and Transformers (with trust_remote_code for the custom architecture) and recommends Nomic's megablocks fork for GPU performance.1 |
| Training codeIs the code used to train the model published? | Public | The card and technical report point to the contrastors repository.126 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The card says training used 1.6 billion multilingual pairs and that the training data is released, pointing to the contrastors repository. That README documents access (with a Nomic Atlas account) to the nomic-embed-text-v1 dataset but does not separately describe the v2 multilingual data.16 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The card summarizes the recipe: weakly supervised contrastive pretraining, then supervised fine-tuning, with consistency filtering and Matryoshka representation learning; the technical report gives details.12 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The technical report's abstract says code, models, and evaluation data are open sourced; the card shows BEIR and MIRACL results.21 |
What it is useful for
Multilingual retrieval, including retrieval-augmented generation. Queries and documents must carry the task prefixes search_query and search_document.1
Run and use notes
- The card recommends the 256-dimension embeddings when storage or compute is a concern and notes that the mixture-of-experts architecture may need more resources than comparable dense models.1
Organization context
Provenance and derivatives
Trained by Nomic from nomic-xlm-2048, which Nomic made by swapping XLM-RoBERTa Base's learned position embeddings for rotary embeddings and training further on CC100. The released model is fine-tuned from the contrastively pretrained checkpoint nomic-embed-text-v2-moe-unsupervised.145
- Derived from: nomic-embed-text-v2-moe-unsupervised (external site: huggingface.co) — Contrastive-pretraining checkpoint from Nomic.
- Derived from: nomic-xlm-2048 (external site: huggingface.co) — Nomic's RoPE variant of FacebookAI/xlm-roberta-base.
Other releases in the Nomic Embed family
- nomic-embed-text-v1.5Model-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.