datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.c-sac-corpora
C-SAC LibriTTS-R training subset
Deterministically selected and resampled speech from
mythicinfinity/libritts_r
for the C-SAC causal speech-codec program. The package retains source revision,
Parquet shard, row, utterance, transcript, and content hashes. LibriTTS-R is
distributed under CC BY 4.0; downstream users remain responsible for attribution.
Only prefixes with a hash-bound _COMPLETE.json sentinel are admissible.
DeMix_Corpora
Dataset Card for DeMix Corpora
DeMix
📄 Paper: Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
🤗 Dataset: DeMix Corpora
🐱 Github: Demix
Dataset Details
Dataset Description
DeMix Corpora (15T original tokens and 22T mixture tokens) serves as a comprehensive, high-quality, large-scale, and carefully mixed resource that can be directly employed for pre-training.
(2026.2.7: This is an… See the full description on the dataset page: https://huggingface.co/datasets/lucius1022/DeMix_Corpora.kalamaki_corporawmdp-corpora
Dataset Card for WMDP Corpora
The Weapons of Mass Destruction Proxy (WMDP) Corpora includes all of the corpora used to perform unlearning on WMDP-Bio and WMDP-Cyber.
See our paper, website, and GitHub for more details!
The corpora are also available at the following mirrors with password wmdpcorpora: 1, 2
The bio forget corpus must be requested separately; please visit this form.
cyber-retain-corpus and cyber-forget-corpus
The forget and retain corpora consist of… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-corpora.legalbench_corporate_lobbying
LegalBenchCorporateLobbying
An MTEB dataset
Massive Text Embedding Benchmark
The dataset includes bill titles and bill summaries related to corporate lobbying.
Task category
t2t
Domains
Legal, Written
Reference
https://huggingface.co/datasets/nguha/legalbench/viewer/corporate_lobbying
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_corporate_lobbying.ThinkChain-20M
We are excited to announce the release of SVECTOR-CORPORATION/ThinkChain-20M, a synthetic reasoning dataset containing over 22 million general reasoning questions and responses generated using Spec-T1. While multiple efforts exist to build open reasoning datasets for math and code tasks, there has been a gap in large datasets covering diverse non code/math topics such as social and natural sciences, education, creative writing, and general conversations. This dataset fills that gap.
Note: The… See the full description on the dataset page: https://huggingface.co/datasets/SVECTOR-CORPORATION/ThinkChain-20M.raw-text-corpus
📝 Zomi Raw Text Corpus (Community-Contributed)
The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks.
This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.medicine-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the instruction-augmented corpora in biomedicine domain used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train language models. The instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/medicine-instruction-augmented-corpora.calib-corpora
calib-corpora
A pool of calibration material, the recipes that turn it into a calibration set
for one specific model, and the measurement corpora those quants are scored
against.
This repository is not a corpus. Nothing here is meant to be fed to
llama-imatrix as-is except the files under builds/, and each of those was
made for one named model and is close to useless for any other.
Why it is built this way
The first version of this repository was a single… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/calib-corpora.ALARM-Corpora
Dataset Card for ALARM-Corpora
Dataset Summary
This is the dataset used in the ALARM: Audio-Language Alignment for Reasoning Models paper.
It consists of Audio Captions and Reasoning Language Model responses rephrased to sound like they were provided by an audio-understanding model.
For more details regarding the dataset and the instructions for obtaining audio files, please refer to our GitHub.
Dataset Statistics
Audio Type
# Elements (M)
# Hours (K)… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/ALARM-Corpora.lma-individual-project-corporamarketingVoiceodqa-wiki-corpora
Dataset Card for Open-Domain Question Answering Wikipedia Corpora
Dataset Description
Dataset Summary
The Wikipedia corpus variants provided can serve as knowledge sources for question-answering systems based on a retriever–reader pipeline. These corpus variants and their corresponding experiments are described further in the paper entitled:
Pre-Processing Matters! Improved Wikipedia Corpora for Open-Domain Question Answering.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/castorini/odqa-wiki-corpora.twsyllables
twsyllables — Taiwanese Mandarin syllable acoustics
Per-syllable acoustic reference data for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW):
37,947 measured syllable tokens, position-sensitive acoustic templates for 1,491
syllable×tone types, voice-onset-time norms for all 17 obstruent initials, and
a between-speaker variability model estimated over 271 speakers.
Every number was measured from native Taiwanese recordings by one reproducible
pipeline; no figure in this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twsyllables.leipzig_corpora_collection
Leipzig Corpora Collection
The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs.
The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.entity-references
Entity References Database
A comprehensive entity database for organizations, people, roles, and locations with embedding-based semantic search. Built from authoritative sources (GLEIF, SEC, Companies House, Wikidata) for entity linking and named entity disambiguation.
Dataset Summary
This dataset provides fast lookup and qualification of named entities using vector similarity search. It stores records from authoritative global sources with embeddings generated by… See the full description on the dataset page: https://huggingface.co/datasets/Corp-o-Rate-Community/entity-references.epiq-laer-corporate-benchmark
Epiq LAER CorporateBench
Harbor release of Epiq LAER CorporateBench: Enterprise Knowledge.
Epiq LAER CorporateBench is CorporateBench in its Harbor configuration. It packages the five CorporateBench capabilities as 128 scored Harbor tasks with 1,132 graded cases over the four synthetic companies of the original paper, with five development tasks alongside.
What is this?
LLMs are increasingly able to answer complex questions about enterprise-scale document… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/epiq-laer-corporate-benchmark.corporatebench
CorporateBench
Dataset release for CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases.
CorporateBench evaluates information extraction, retrieval, and question answering over four synthetic corporate corpora ranging from 353 to 232,692 released documents. The corpora are generated from temporally evolving knowledge bases, providing deterministic ground truth across related documents.
Dataset Viewer
https://corporatebench.epiqai.com/… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/corporatebench.mats-gf-provenance-corpora
Provenance-codeword training corpora
All training corpora from the eight-experiment provenance codewords
program (per-source activation codewords in Qwen3 models). Code, paper, and
reproduction scripts:
https://github.com/Sid-MB/mats-gf-provenance-codewords
Each synthetic corpus ships in full: docs.parquet (training documents),
train.parquet, qa.parquet (probe questions incl. phantom-fact controls),
generation intermediates (raw/), the sqlite sequence store (seqdb/), and
audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.south-african-monolingual-corpora-jsonl
South African Languages Pretraining Dataset
This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections.
The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity
Languages Included
Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.es_corpora_parliament_processeden_corpora_parliament_processedde_corpora_parliament_processedcouncil-training-corporatwngrams
Taiwanese Mandarin web n-grams
Word 1–4-gram counts for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW), computed over the
Taiwan slice of a large web crawl after variety filtering by
twfilter 0.1.0 with the published
twfilter-tables:
every sentence behind these counts passed the 教育部 character-inventory gate, the
simplified-character round-trip, the mainland-orthography, mainland-lexicon,
written-Cantonese, Hong Kong and Singapore detectors, and block-level evidence of
Taiwan-specific… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twngrams.twfilter-tables
twfilter reference tables
The tables that decide whether a span of traditional-Chinese text is Taiwanese Mandarin*
(臺灣華語, cmn-Hant-TW) rather than Hong Kong Cantonese, mainland text converted to traditional
characters, or literary Chinese. Plain text, one record per line, tab-separated where a
record has fields, LC_ALL=C sort order, UTF-8, LF.
Consumed by twfilter 0.1.0, where this
directory is vendored byte-for-byte and MANIFEST.json is verified by its test suite.
Usable… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twfilter-tables.hlwm-corpora
HLWM training corpora
The training data behind the Hierarchical Latent Workspace Model program — a twenty-day,
ten-experiment preregistered attempt to build a latent-workspace language model on a frozen
Qwen3-0.6B decoder.
All three proposed mechanisms failed their preregistered gates. Both papers are
negative-results reports. This dataset is published so the record is checkable, not because it
produced a working system.
Papers, code and full experimental record:… See the full description on the dataset page: https://huggingface.co/datasets/slashgg/hlwm-corpora.wmdp-mmlu-auxiliary-corpora
Dataset Card for WMDP Auxiliary Corpora
This dataset includes the auxiliary corpora used to perform unlearning on the MMLU Auxiliary Benchmark task, from the WMDP paper.
See our paper, website, and GitHub for more details!
The corpora are also available at the following mirrors with password wmdpauxiliarycorpora: 1, 2
physics-corpus
Corpus comprising textbooks in high school and college physics.
law-corpus
Corpus comprising textbooks in international and… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-mmlu-auxiliary-corpora.Parallel_corpora
[!NOTE]
Dataset origin: https://portulanclarin.net/repository/browse/parallel-corpora-finely-aligned-subsentencial-granularity/aa90dbbeb0ab11ea8dc202420a00040310be6f259e694e659c10c1d212b389f0/
Description
Text corpus for bilingual concordancing, single- and multi-word translation extraction, machine translation.
Languages: cs-pt, de-pt, en-pt, es-pt, fr-pt, it-pt, and pt-sk.
Domain: Law and Health.
