corpora
Datasets
All datasets matching “corpora”general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.c-sac-corpora
C-SAC LibriTTS-R training subset
Deterministically selected and resampled speech from
mythicinfinity/libritts_r
for the C-SAC causal speech-codec program. The package retains source revision,
Parquet shard, row, utterance, transcript, and content hashes. LibriTTS-R is
distributed under CC BY 4.0; downstream users remain responsible for attribution.
Only prefixes with a hash-bound _COMPLETE.json sentinel are admissible.
DeMix_Corpora
Dataset Card for DeMix Corpora
DeMix
📄 Paper: Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
🤗 Dataset: DeMix Corpora
🐱 Github: Demix
Dataset Details
Dataset Description
DeMix Corpora (15T original tokens and 22T mixture tokens) serves as a comprehensive, high-quality, large-scale, and carefully mixed resource that can be directly employed for pre-training.
(2026.2.7: This is an… See the full description on the dataset page: https://huggingface.co/datasets/lucius1022/DeMix_Corpora.kalamaki_corporawmdp-corpora
Dataset Card for WMDP Corpora
The Weapons of Mass Destruction Proxy (WMDP) Corpora includes all of the corpora used to perform unlearning on WMDP-Bio and WMDP-Cyber.
See our paper, website, and GitHub for more details!
The corpora are also available at the following mirrors with password wmdpcorpora: 1, 2
The bio forget corpus must be requested separately; please visit this form.
cyber-retain-corpus and cyber-forget-corpus
The forget and retain corpora consist of… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-corpora.legalbench_corporate_lobbying
LegalBenchCorporateLobbying
An MTEB dataset
Massive Text Embedding Benchmark
The dataset includes bill titles and bill summaries related to corporate lobbying.
Task category
t2t
Domains
Legal, Written
Reference
https://huggingface.co/datasets/nguha/legalbench/viewer/corporate_lobbying
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_corporate_lobbying.
