USASI
DatasetDataset

The Pile

Project record

Maintained by EleutherAI364

The Pile is an 825 GiB English text dataset for language-model pretraining, assembled by EleutherAI from 22 component datasets that include web text (Pile-CC, from Common Crawl), academic and professional sources such as PubMed Central, arXiv, and FreeLaw, books, GitHub code, and dialogue. It was described in a December 2020 paper and a 2022 datasheet and is the training corpus of EleutherAI's Pythia models. Its current availability is limited; see access conditions.23411

Last reviewedEntry updated Documented release Unknown

Availability and license

Overall availability

Partial

The Pile website and replication-code README still point to a download hosted by the Eye, which the Pythia model card calls a community mirror; when this catalog checked on 2026-09-29, that host presented an expired certificate and the Pile download path returned a not-found error. The Hugging Face "EleutherAI/pile" repository contains only a dataset card and a loading script that downloads from that host. The 2025 Common Pile paper, whose authors include EleutherAI researchers, states that the use of unlicensed training data has resulted in DMCA takedowns of datasets such as the Pile. EleutherAI's Hugging Face organization still hosts a deduplicated text version without a dataset card, the pre-tokenized data used to train Pythia, and the validation and test splits (tagged with an MIT license).141167108129

Availability is separate from permission: read the license before using or redistributing.

The dataset has no single license. The Hugging Face card directs users to the license of each component subset, and the loading script lists "Unknown" as the license of every individual component it offers. The MIT license covers the replication code only.675

Public materials checklist

Items for a dataset under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for The Pile
ItemStatusNotes and evidence
AccessCan the data be obtained, and on what terms?PartialThe original download channel could not be verified; EleutherAI hosts a deduplicated version, pre-tokenized versions, and the validation and test splits on Hugging Face.11089
ProvenanceAre the data's origins documented?PublicThe paper, datasheet, and replication-code README list the 22 components with their sources, sizes, and sampling weights, and each document is tagged with its component.2349
DocumentationIs there a datasheet, card, or equivalent documentation?PublicA paper and a separate datasheet document the dataset.23
LicensingAre the licensing terms stated?PartialLicensing is left to each component, and many component licenses are listed as unknown.67
Stated limitationsDoes the documentation state known limitations or risks?PublicEleutherAI's Pythia card warns that the Pile is known to contain profanity and lewd or otherwise offensive text and points to the Pile paper's discussion of documented gender, religious, and racial biases.11

What it is useful for

Pretraining and studying language models on diverse text, and as a language-modeling benchmark (Pile bits per byte), as described on the Pile website. As the fixed training corpus of the Pythia suite, it also underpins research on memorization and training dynamics.112

Organization context

Provenance and derivatives

Assembled by EleutherAI from 22 existing or newly scraped sources, including Pile-CC (text extracted from Common Crawl), PubMed Central, Books3, OpenWebText2, arXiv, GitHub, FreeLaw, Stack Exchange, USPTO backgrounds, Project Gutenberg (PG-19), OpenSubtitles, English Wikipedia, and others. The datasheet describes the sources as a mix of original scrapes, data made available by the data owners, and third-party scrapes.43

Catalog records that name this entry in their provenance:

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The datasheet describes the Pile as created by EleutherAI, and its website, replication code, and Hugging Face copies are maintained under EleutherAI's domains and organizations. EleutherAI states it incorporated as a non-profit research institute in early 2023, and the IRS exempt-organization extract for the District of Columbia lists EleutherAI Institute in Washington, D.C., as a 501(c)(3) organization.31461314

Assessed Sep 29, 2026

Sources

  1. 1.
  2. 2.
    The Pile: An 800GB Dataset of Diverse Text for Language Modeling (arXiv 2101.00027) (external site: arxiv.org)

    EleutherAI (arXiv) · Paper · published Dec 31, 2020 · accessed Sep 29, 2026

  3. 3.
    Datasheet for the Pile (arXiv 2201.07311) (external site: arxiv.org)

    EleutherAI (arXiv) · Paper · published Jan 13, 2022 · accessed Sep 29, 2026

  4. 4.
  5. 5.
  6. 6.
    EleutherAI/pile dataset card (external site: huggingface.co)

    EleutherAI · Dataset card · accessed Sep 29, 2026

  7. 7.
    EleutherAI/pile loading script (pile.py) (external site: huggingface.co)

    EleutherAI · Repository · accessed Sep 29, 2026

  8. 8.
    EleutherAI/the_pile_deduplicated (external site: huggingface.co)

    EleutherAI · Dataset card · accessed Sep 29, 2026

  9. 9.
  10. 10.
    The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text (arXiv 2506.05209) (external site: arxiv.org)

    EleutherAI and collaborators (arXiv) · Paper · published Jun 2025 · accessed Sep 29, 2026

  11. 11.
    EleutherAI/pythia-12b model card (external site: huggingface.co)

    EleutherAI · Model card · accessed Sep 29, 2026

  12. 12.
  13. 13.
    About | EleutherAI (external site: eleuther.ai)

    EleutherAI · Official page · accessed Sep 29, 2026

  14. 14.

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project