datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeMix_Corpora
Dataset Card for DeMix Corpora
DeMix
📄 Paper: Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
🤗 Dataset: DeMix Corpora
🐱 Github: Demix
Dataset Details
Dataset Description
DeMix Corpora (15T original tokens and 22T mixture tokens) serves as a comprehensive, high-quality, large-scale, and carefully mixed resource that can be directly employed for pre-training.
(2026.2.7: This is an… See the full description on the dataset page: https://huggingface.co/datasets/lucius1022/DeMix_Corpora.kalamaki_corporalma-individual-project-corporaentity-references
Entity References Database
A comprehensive entity database for organizations, people, roles, and locations with embedding-based semantic search. Built from authoritative sources (GLEIF, SEC, Companies House, Wikidata) for entity linking and named entity disambiguation.
Dataset Summary
This dataset provides fast lookup and qualification of named entities using vector similarity search. It stores records from authoritative global sources with embeddings generated by… See the full description on the dataset page: https://huggingface.co/datasets/Corp-o-Rate-Community/entity-references.twsyllables
twsyllables — Taiwanese Mandarin syllable acoustics
Per-syllable acoustic reference data for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW):
37,947 measured syllable tokens, position-sensitive acoustic templates for 1,491
syllable×tone types, voice-onset-time norms for all 17 obstruent initials, and
a between-speaker variability model estimated over 271 speakers.
Every number was measured from native Taiwanese recordings by one reproducible
pipeline; no figure in this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twsyllables.mats-gf-provenance-corpora
Provenance-codeword training corpora
All training corpora from the eight-experiment provenance codewords
program (per-source activation codewords in Qwen3 models). Code, paper, and
reproduction scripts:
https://github.com/Sid-MB/mats-gf-provenance-codewords
Each synthetic corpus ships in full: docs.parquet (training documents),
train.parquet, qa.parquet (probe questions incl. phantom-fact controls),
generation intermediates (raw/), the sqlite sequence store (seqdb/), and
audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.corporate-actions
US Corporate Actions — dividends and splits
391 639 dividends from 3 327 filers · 5 619 splits from 3 814 filers ·
2005 to 2026
Built to close a specific hole. A filing states shares and earnings per share
as of the day it was made; every price series is adjusted for splits since.
Multiply one by the other and the answer is wrong by the split factor — on
Deckers that turned a 6.9% earnings yield into 41.7%, a P/E of 1.8.
The pipeline lives in recipe/ at the same revision as the… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/corporate-actions.corporate-emission-reports
Dataset Card for Dataset Name
A dataset of 100 corporate sustainability reports with manually extracted scope 1, 2 and 3 greenhouse gas emission values.
Dataset Details
Dataset Description
Data about corporate greenhouse gas emissions is usually published only as part of sustainability report PDF's, which is not a machine-readable format. Interested actors have to manually extract emission data from these reports, which is a tedious and time-consuming process.… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/corporate-emission-reports.English-Zomi-OPUS_Tatoeba_v20230412
English–Zomi Parallel Corpus (1.78M)
This dataset contains 1.78 million English–Zomi sentence pairs, created to support
machine translation, linguistic research, and large‑scale language model training.
It is fully open and permissively licensed for commercial and non‑commercial use.
🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes
Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/English-Zomi-OPUS_Tatoeba_v20230412.mining-legal-arguments-us-corporate-case-law
Mining Legal Arguments in U.S. Corporate Case Law
This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.SA-Parallel-Corpora
SA-Parallel-Corpora
Sentence-aligned English to isiZulu, isiXhosa, Sesotho and Sepedi
bitext, drawn from South African government publications.
Produced for the doctoral thesis Injecting Commonsense Knowledge into
Pretrained Language Models for Low Resource Languages (University of Cape Town,
2026). Code at https://github.com/sello-ralethe/SA-knowledge
Structure
One configuration per language pair, each with train, validation
and test splits. Splits are assigned… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Parallel-Corpora.lilm1-tool-teacher-corpora
LiLM1 tool teacher corpora
This dataset contains synthetic tool-use records generated with Gemma and Qwen
teacher models.
Method
Each teacher received structured tool schemas and task templates. One
configuration preserves the records from each teacher and task set.
Configurations
Configuration
Content
gemma-26b-a4b-function
Gemma function-calling records
qwen-27b-function
Qwen function-calling records
qwen-35b-a3b-function
Qwen MoE… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-tool-teacher-corpora.crh-parallel-corpora-document-level-noisycorporate-crypto-treasuries
Corporate Crypto Treasury Holdings
Every company, government and ETF known to hold Bitcoin, Ethereum or Solana on
its balance sheet, with the primary source document behind each figure.
359 positions across 297 entities in 39 countries. 358 of the 359 carry a
link to the filing or disclosure the number came from.
Maintained by CorpStacking.
What makes this different from a price feed
Most crypto datasets are market data. This one is balance-sheet data read out
of… See the full description on the dataset page: https://huggingface.co/datasets/corpstacking/corporate-crypto-treasuries.low_resource_parallel_corpora
The Little Prince — multiparallel corpus (22 languages, RU pivot)
A sentence-level multiparallel corpus of Antoine de Saint-Exupéry's The Little Prince,
built around the classic Russian translation by Nora Gal as the pivot and covering
21 further editions, most of them in low-resource minority languages of Russia.
Every one of the 1,565 pivot sentences has exactly one aligned sentence in every
included language — a perfect N-way alignment (no gaps, no merges). The editions were… See the full description on the dataset page: https://huggingface.co/datasets/averoo/low_resource_parallel_corpora.quant-fidelity-corpora
Quantization fidelity corpora
Evaluation text for measuring how faithfully a quantized LLM reproduces its full-precision parent
(KL divergence of next-token distributions, top-1 agreement, perplexity ratio), as used by
Agention for the Signal and Qwen3.8-27B quantization campaigns.
mixedweb-v1
mixedweb-v1.txt (800,789 chars, 301 documents, md5 51e0045e8cabf37922aa82766a25b7b4) is a
seeded random slice of HuggingFaceFW/fineweb
sample-10BT: general English web text… See the full description on the dataset page: https://huggingface.co/datasets/agentionai/quant-fidelity-corpora.KYS-1.5B-Pretraining-Corpora
KYS-1.5B-Pretraining-Corpora
The six 10B-token pretraining mixtures from Know Your Sources: Data Selection Matters when
Rewriting for Data-Constrained Pretraining, stored without duplication.
The key idea: one anchor, six remainders
Every setting trains on the same 10B-token recipe:
10B mixture = 5B shared anchor + 5B strategy-specific tokens
(identical in all (this is the ONLY thing
six settings, that differs… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-1.5B-Pretraining-Corpora.evenki-rus-parallel-corpora-oralbitcoin-news-articles-text-corporaasia-owid-revenue-from-corporate-income-taxes-gdp
Revenue From Corporate Income Taxes Gdp | Asia (Our World in Data)
🌏 855 observations · 40 Asia countries · 1980–2017 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 855 observations of Revenue From Corporate Income Taxes Gdp data across 40 Asia countries, spanning 1980–2017.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Revenue From Corporate Income Taxes Gdp… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-revenue-from-corporate-income-taxes-gdp.Corporate_Governance_Risk_Leadership_Practical
Corporate Governance Risk Leadership — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Corporate_Governance_Risk_Leadership_Practical.asia-owid-taxes-on-incomes-of-individuals-and-corporations-gdp
Taxes On Incomes Of Individuals And Corporations Gdp | Asia (Our World in Data)
🌏 1,279 observations · 45 Asia countries · 1980–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 1,279 observations of Taxes On Incomes Of Individuals And Corporations Gdp data across 45 Asia countries, spanning 1980–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Taxes On Incomes Of… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-taxes-on-incomes-of-individuals-and-corporations-gdp.emotion-story-corpora
Emotion story corpora
Every story corpus behind How are emotions represented in large language
models?, in one place. These are the texts a model reads while its internal
state is recorded, each written to evoke a named emotion without naming it.
The project's most transferable finding is that who writes the stories decides
whether any of it works: holding everything else fixed, swapping the story
writer moved the number of working layers from 1 to 9 out of 20. That comparison… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-story-corpora.africa-uganda-depository-corporation-survey-billion-shillings-june-2019-310ed409
Depository Corporation Survey Billion Shillings June 2019 | Africa (Uganda Bureau of Statistics)
26 rows - 1 Africa country/area - 2019-2025 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 26 rows from Uganda Bureau of Statistics, covering Depository Corporation Survey Billion Shillings June 2019. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-depository-corporation-survey-billion-shillings-june-2019-310ed409.africa-uganda-depository-corporation-survey-billion-shillings-june-2014-f3c8d6c4
Depository Corporation Survey Billion Shillings June 2014 | Africa (Uganda Bureau of Statistics)
156 rows - 1 Africa country/area - 2014-2019 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 156 rows from Uganda Bureau of Statistics, covering Depository Corporation Survey Billion Shillings June 2014. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-depository-corporation-survey-billion-shillings-june-2014-f3c8d6c4.evenki-rus-parallel-corporaafrica-owid-revenue-from-corporate-income-taxes-gdp
Revenue From Corporate Income Taxes Gdp | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-revenue-from-corporate-income-taxes-gdp.europe-owid-revenue-from-corporate-income-taxes-gdp
Revenue From Corporate Income Taxes Gdp | Europe (Our World in Data)
🇪🇺 1,145 observations · 40 Europe countries · 1980–2017 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 1,145 observations of Revenue From Corporate Income Taxes Gdp data across 40 Europe countries, spanning 1980–2017.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Revenue From Corporate Income Taxes Gdp… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-revenue-from-corporate-income-taxes-gdp.corporate_esg_risk_analyticssdg-corpora-v0.3
SDG Ontology-Grounded Synthetic Corpus — v0.3 (first cut)
A verifiable, attribution-clean synthetic textbook corpus for relational data-governance
metadata, grounded in a BFO 2020 / CCO ontology. Unlike web-retrieval synthetic corpora, every
chapter is generated from deterministic ontology axioms (not scraped text), populates a
deterministic relational schema, and ships with the generator's reasoning trace.
Part of the Aegir project; the source artifacts live in… See the full description on the dataset page: https://huggingface.co/datasets/zndx/sdg-corpora-v0.3.
