USASI
Model release

nomic-embed-text-v2-moe

Release in the Nomic Embed family · version nomic-embed-text-v2-moe

Maintained by Nomic AI1

A multilingual text embedding model that uses a mixture-of-experts design (8 experts, top-2 routing), with 475M total and 305M active parameters. It covers about 100 languages, produces 768-dimensional embeddings that can be truncated to 256 through Matryoshka training, and accepts inputs of up to 512 tokens. Nomic released it in February 2025.123

Last reviewedEntry updated Documented release Feb 2025

Availability and license

Overall availability

Public

Downloadable from Hugging Face without gating.1

Availability is separate from permission: read the license before using or redistributing.

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for nomic-embed-text-v2-moe
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?Public1
Inference codeIs code for running the model published?PublicThe card documents use with Sentence Transformers and Transformers (with trust_remote_code for the custom architecture) and recommends Nomic's megablocks fork for GPU performance.1
Training codeIs the code used to train the model published?PublicThe card and technical report point to the contrastors repository.126
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe card says training used 1.6 billion multilingual pairs and that the training data is released, pointing to the contrastors repository. That README documents access (with a Nomic Atlas account) to the nomic-embed-text-v1 dataset but does not separately describe the v2 multilingual data.16
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe card summarizes the recipe: weakly supervised contrastive pretraining, then supervised fine-tuning, with consistency filtering and Matryoshka representation learning; the technical report gives details.12
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PublicThe technical report's abstract says code, models, and evaluation data are open sourced; the card shows BEIR and MIRACL results.21

What it is useful for

Multilingual retrieval, including retrieval-augmented generation. Queries and documents must carry the task prefixes search_query and search_document.1

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The card recommends the 256-dimension embeddings when storage or compute is a concern and notes that the mixture-of-experts architecture may need more resources than comparable dense models.1

Organization context

Provenance and derivatives

Trained by Nomic from nomic-xlm-2048, which Nomic made by swapping XLM-RoBERTa Base's learned position embeddings for rotary embeddings and training further on CC100. The released model is fine-tuned from the contrastively pretrained checkpoint nomic-embed-text-v2-moe-unsupervised.145

Other releases in the Nomic Embed family

Nomic Embed family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

Developed and published by Nomic, Inc., which states that its headquarters is in New York City. The model starts from nomic-xlm-2048, Nomic's modified XLM-RoBERTa Base.185

Assessed Sep 29, 2026

Sources

  1. 1.
  2. 2.
  3. 3.
    Nomic Embed Text V2: An Open Source, Multilingual, Mixture-of-Experts Embedding Model (external site: simonwillison.net)

    Simon Willison · News report · published Feb 12, 2025 · accessed Sep 29, 2026

  4. 4.
  5. 5.
    nomic-ai/nomic-xlm-2048 model card (external site: huggingface.co)

    Nomic · Model card · accessed Sep 29, 2026

  6. 6.
    nomic-ai/contrastors (external site: github.com)

    Nomic · Repository · accessed Sep 29, 2026

  7. 7.
  8. 8.
    Careers | Nomic (external site: nomic.ai)

    Nomic · Official page · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project