datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kalamaki_corporawmdp-corpora
Dataset Card for WMDP Corpora
The Weapons of Mass Destruction Proxy (WMDP) Corpora includes all of the corpora used to perform unlearning on WMDP-Bio and WMDP-Cyber.
See our paper, website, and GitHub for more details!
The corpora are also available at the following mirrors with password wmdpcorpora: 1, 2
The bio forget corpus must be requested separately; please visit this form.
cyber-retain-corpus and cyber-forget-corpus
The forget and retain corpora consist of… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-corpora.ALARM-Corpora
Dataset Card for ALARM-Corpora
Dataset Summary
This is the dataset used in the ALARM: Audio-Language Alignment for Reasoning Models paper.
It consists of Audio Captions and Reasoning Language Model responses rephrased to sound like they were provided by an audio-understanding model.
For more details regarding the dataset and the instructions for obtaining audio files, please refer to our GitHub.
Dataset Statistics
Audio Type
# Elements (M)
# Hours (K)… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/ALARM-Corpora.corporate-actions
US Corporate Actions — dividends and splits
391 639 dividends from 3 327 filers · 5 619 splits from 3 814 filers ·
2005 to 2026
Built to close a specific hole. A filing states shares and earnings per share
as of the day it was made; every price series is adjusted for splits since.
Multiply one by the other and the answer is wrong by the split factor — on
Deckers that turned a 6.9% earnings yield into 41.7%, a P/E of 1.8.
The pipeline lives in recipe/ at the same revision as the… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/corporate-actions.mats-gf-provenance-corpora
Provenance-codeword training corpora
All training corpora from the eight-experiment provenance codewords
program (per-source activation codewords in Qwen3 models). Code, paper, and
reproduction scripts:
https://github.com/Sid-MB/mats-gf-provenance-codewords
Each synthetic corpus ships in full: docs.parquet (training documents),
train.parquet, qa.parquet (probe questions incl. phantom-fact controls),
generation intermediates (raw/), the sqlite sequence store (seqdb/), and
audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.leipzig_corpora_collection
Leipzig Corpora Collection
The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs.
The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.es_corpora_parliament_processeden_corpora_parliament_processedThe_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden
Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden"
Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io
Dataset Summary
The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.wmdp-mmlu-auxiliary-corpora
Dataset Card for WMDP Auxiliary Corpora
This dataset includes the auxiliary corpora used to perform unlearning on the MMLU Auxiliary Benchmark task, from the WMDP paper.
See our paper, website, and GitHub for more details!
The corpora are also available at the following mirrors with password wmdpauxiliarycorpora: 1, 2
physics-corpus
Corpus comprising textbooks in high school and college physics.
law-corpus
Corpus comprising textbooks in international and… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-mmlu-auxiliary-corpora.de_corpora_parliament_processedgeneral-instruction-augmented-corporaThis is a reupload of general instruction-augmented corpora in a more accessible format. Please, cite the original repository.
tokbench-corpora
tokbench corpora
The input corpora for tokbench, a
benchmark that measures tokenizer implementations against each other on the
same bytes. Each config is one corpus: ~5 MB of real text chosen to stress a
different part of a tokenizer.
Nothing here is new text. It is a fixed, pinned, redistributable excerpt of
public datasets, packaged so a tokenizer benchmark is reproducible by anyone
without re-deriving the inputs. Provenance and licence for every config are in
the table below.… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/tokbench-corpora.sv_corpora_parliament_processedsv_corpora_parliament_processeddutch_corpora_parliament_processedfr_corpora_parliament_processednl_corpora_parliament_processedwmdp_deduped_corporaur_corpora_pibsv_corpora_parliament_processed_v0es_corpora_parliament_processedtask427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.sv_corpora_parliament_processedcorporate-emission-reports
Dataset Card for Dataset Name
A dataset of 100 corporate sustainability reports with manually extracted scope 1, 2 and 3 greenhouse gas emission values.
Dataset Details
Dataset Description
Data about corporate greenhouse gas emissions is usually published only as part of sustainability report PDF's, which is not a machine-readable format. Interested actors have to manually extract emission data from these reports, which is a tedious and time-consuming process.… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/corporate-emission-reports.N1K0_ka_en_corpora
N1K0 Georgian-English Corpora
A high-quality, heavily sanitized parallel translation dataset containing ~2 Million (1,988,071) parallel sentence pairs for Georgian (ka) and English (en).
This dataset is specifically engineered for neural machine translation tasks, sequence-to-sequence models (e.g., Transformers), and cross-attention decoder pretraining/fine-tuning.
🚀 Quick Start & Usage
You can load this dataset directly using the Hugging Face datasets library:… See the full description on the dataset page: https://huggingface.co/datasets/N1K0L0Z/N1K0_ka_en_corpora.leipzig-corpora-collectionsv_corpora_parliament_processedit_corpora_parliament_processedbashkir-russian-parallel-corpora
Dataset Card for "bashkir-russian-parallel-corpora"
How the dataset was assembled.
find the text in two languages. it can be a translated book or an internet page (wikipedia, news site)
our algorithm tries to match Bashkir sentences with their translation in Russian
We give these pairs to people to check
@inproceedings{
title={Bashkir-Russian parallel corpora},
author={Iskander Shakirov, Aigiz Kunafin},
year={2023}
}
