CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucius1022 /DeMix_Corpora Dataset Card for DeMix Corpora DeMix 📄 Paper: Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training 🤗 Dataset: DeMix Corpora 🐱 Github: Demix Dataset Details Dataset Description DeMix Corpora (15T original tokens and 22T mixture tokens) serves as a comprehensive, high-quality, large-scale, and carefully mixed resource that can be directly employed for pre-training. (2026.2.7: This is an… See the full description on the dataset page: https://huggingface.co/datasets/lucius1022/DeMix_Corpora.tabularn<1K3 likes3.4k downloads7mo agoHugging Face02SotirisLegkas /kalamaki_corporatabular100M<n<1B0 likes2.6k downloads1y agoHugging Face03Mohith202 /lma-individual-project-corporatabular10K<n<100K0 likes384 downloads6d agoHugging Face04Corp-o-Rate-Community /entity-references Entity References Database A comprehensive entity database for organizations, people, roles, and locations with embedding-based semantic search. Built from authoritative sources (GLEIF, SEC, Companies House, Wikidata) for entity linking and named entity disambiguation. Dataset Summary This dataset provides fast lookup and qualification of named entities using vector similarity search. It stores records from authoritative global sources with embeddings generated by… See the full description on the dataset page: https://huggingface.co/datasets/Corp-o-Rate-Community/entity-references.tabulartext-classificationn<1K0 likes264 downloads5mo agoHugging Face05taiwan-corpora /twsyllables twsyllables — Taiwanese Mandarin syllable acoustics Per-syllable acoustic reference data for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW): 37,947 measured syllable tokens, position-sensitive acoustic templates for 1,491 syllable×tone types, voice-onset-time norms for all 17 obstruent initials, and a between-speaker variability model estimated over 271 speakers. Every number was measured from native Taiwanese recordings by one reproducible pipeline; no figure in this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twsyllables.tabularother1K<n<10K0 likes247 downloads28d agoHugging Face06siddharthmb /mats-gf-provenance-corpora Provenance-codeword training corpora All training corpora from the eight-experiment provenance codewords program (per-source activation codewords in Qwen3 models). Code, paper, and reproduction scripts: https://github.com/Sid-MB/mats-gf-provenance-codewords Each synthetic corpus ships in full: docs.parquet (training documents), train.parquet, qa.parquet (probe questions incl. phantom-fact controls), generation intermediates (raw/), the sqlite sequence store (seqdb/), and audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.tabular1K<n<10K0 likes227 downloads2mo agoHugging Face07ZipLime /corporate-actions US Corporate Actions — dividends and splits 391 639 dividends from 3 327 filers · 5 619 splits from 3 814 filers · 2005 to 2026 Built to close a specific hole. A filing states shares and earnings per share as of the day it was made; every price series is adjusted for splits since. Multiply one by the other and the answer is wrong by the split factor — on Deckers that turned a 6.9% earnings yield into 41.7%, a P/E of 1.8. The pipeline lives in recipe/ at the same revision as the… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/corporate-actions.tabulartabular-regression100K<n<1M0 likes157 downloads1d agoHugging Face08nopperl /corporate-emission-reports Dataset Card for Dataset Name A dataset of 100 corporate sustainability reports with manually extracted scope 1, 2 and 3 greenhouse gas emission values. Dataset Details Dataset Description Data about corporate greenhouse gas emissions is usually published only as part of sustainability report PDF's, which is not a machine-readable format. Interested actors have to manually extract emission data from these reports, which is a tedious and time-consuming process.… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/corporate-emission-reports.tabularn<1K0 likes79 downloads3y agoHugging Face09zomi-language-corpora /English-Zomi-OPUS_Tatoeba_v20230412 English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/English-Zomi-OPUS_Tatoeba_v20230412.tabulartranslation1M<n<10M1 likes74 downloads5mo agoHugging Face10lbrenap1 /mining-legal-arguments-us-corporate-case-law Mining Legal Arguments in U.S. Corporate Case Law This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.tabulartext-classification10K<n<100K0 likes72 downloads23d agoHugging Face11sello-ralethe /SA-Parallel-Corpora SA-Parallel-Corpora Sentence-aligned English to isiZulu, isiXhosa, Sesotho and Sepedi bitext, drawn from South African government publications. Produced for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Code at https://github.com/sello-ralethe/SA-knowledge Structure One configuration per language pair, each with train, validation and test splits. Splits are assigned… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Parallel-Corpora.tabular10K<n<100K0 likes68 downloads14d agoHugging Face12glouriousgautam /lilm1-tool-teacher-corpora LiLM1 tool teacher corpora This dataset contains synthetic tool-use records generated with Gemma and Qwen teacher models. Method Each teacher received structured tool schemas and task templates. One configuration preserves the records from each teacher and task set. Configurations Configuration Content gemma-26b-a4b-function Gemma function-calling records qwen-27b-function Qwen function-calling records qwen-35b-a3b-function Qwen MoE… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-tool-teacher-corpora.tabulartext-generation10K<n<100K0 likes56 downloads21d agoHugging Face13QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes52 downloads2y agoHugging Face14corpstacking /corporate-crypto-treasuries Corporate Crypto Treasury Holdings Every company, government and ETF known to hold Bitcoin, Ethereum or Solana on its balance sheet, with the primary source document behind each figure. 359 positions across 297 entities in 39 countries. 358 of the 359 carry a link to the filing or disclosure the number came from. Maintained by CorpStacking. What makes this different from a price feed Most crypto datasets are market data. This one is balance-sheet data read out of… See the full description on the dataset page: https://huggingface.co/datasets/corpstacking/corporate-crypto-treasuries.tabulartabular-regressionn<1K0 likes48 downloads6d agoHugging Face15averoo /low_resource_parallel_corpora The Little Prince — multiparallel corpus (22 languages, RU pivot) A sentence-level multiparallel corpus of Antoine de Saint-Exupéry's The Little Prince, built around the classic Russian translation by Nora Gal as the pivot and covering 21 further editions, most of them in low-resource minority languages of Russia. Every one of the 1,565 pivot sentences has exactly one aligned sentence in every included language — a perfect N-way alignment (no gaps, no merges). The editions were… See the full description on the dataset page: https://huggingface.co/datasets/averoo/low_resource_parallel_corpora.tabulartranslation1K<n<10K9 likes34 downloads2mo agoHugging Face16agentionai /quant-fidelity-corpora Quantization fidelity corpora Evaluation text for measuring how faithfully a quantized LLM reproduces its full-precision parent (KL divergence of next-token distributions, top-1 agreement, perplexity ratio), as used by Agention for the Signal and Qwen3.8-27B quantization campaigns. mixedweb-v1 mixedweb-v1.txt (800,789 chars, 301 documents, md5 51e0045e8cabf37922aa82766a25b7b4) is a seeded random slice of HuggingFaceFW/fineweb sample-10BT: general English web text… See the full description on the dataset page: https://huggingface.co/datasets/agentionai/quant-fidelity-corpora.tabularn<1K0 likes28 downloads3d agoHugging Face17blab-jhu /KYS-1.5B-Pretraining-Corporagated KYS-1.5B-Pretraining-Corpora The six 10B-token pretraining mixtures from Know Your Sources: Data Selection Matters when Rewriting for Data-Constrained Pretraining, stored without duplication. The key idea: one anchor, six remainders Every setting trains on the same 10B-token recipe: 10B mixture = 5B shared anchor + 5B strategy-specific tokens (identical in all (this is the ONLY thing six settings, that differs… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-1.5B-Pretraining-Corpora.tabulartext-generation10M<n<100M0 likes23 downloads29d agoHugging Face18siberian-lang-lab /evenki-rus-parallel-corpora-oraltabular10K<n<100K0 likes19 downloads10mo agoHugging Face19khaihernlow /bitcoin-news-articles-text-corporaimage1K<n<10K0 likes18 downloads2y agoHugging Face20electricsheepasia /asia-owid-revenue-from-corporate-income-taxes-gdp Revenue From Corporate Income Taxes Gdp | Asia (Our World in Data) 🌏 855 observations · 40 Asia countries · 1980–2017 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 855 observations of Revenue From Corporate Income Taxes Gdp data across 40 Asia countries, spanning 1980–2017. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Revenue From Corporate Income Taxes Gdp… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-revenue-from-corporate-income-taxes-gdp.tabulartabular-classificationn<1K0 likes17 downloads4mo agoHugging Face21PhillyMac /Corporate_Governance_Risk_Leadership_Practical Corporate Governance Risk Leadership — Practical This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Corporate_Governance_Risk_Leadership_Practical.tabulartext-generationn<1K0 likes15 downloads5mo agoHugging Face22electricsheepasia /asia-owid-taxes-on-incomes-of-individuals-and-corporations-gdp Taxes On Incomes Of Individuals And Corporations Gdp | Asia (Our World in Data) 🌏 1,279 observations · 45 Asia countries · 1980–2023 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 1,279 observations of Taxes On Incomes Of Individuals And Corporations Gdp data across 45 Asia countries, spanning 1980–2023. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Taxes On Incomes Of… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-taxes-on-incomes-of-individuals-and-corporations-gdp.tabulartabular-classification1K<n<10K0 likes15 downloads3mo agoHugging Face23abotresol /emotion-story-corpora Emotion story corpora Every story corpus behind How are emotions represented in large language models?, in one place. These are the texts a model reads while its internal state is recorded, each written to evoke a named emotion without naming it. The project's most transferable finding is that who writes the stories decides whether any of it works: holding everything else fixed, swapping the story writer moved the number of working layers from 1 to 9 out of 20. That comparison… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-story-corpora.tabular10K<n<100K0 likes15 downloads2mo agoHugging Face24electricsheepafrica /africa-uganda-depository-corporation-survey-billion-shillings-june-2019-310ed409 Depository Corporation Survey Billion Shillings June 2019 | Africa (Uganda Bureau of Statistics) 26 rows - 1 Africa country/area - 2019-2025 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 26 rows from Uganda Bureau of Statistics, covering Depository Corporation Survey Billion Shillings June 2019. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-depository-corporation-survey-billion-shillings-june-2019-310ed409.tabulartabular-classificationn<1K0 likes14 downloads1mo agoHugging Face25electricsheepafrica /africa-uganda-depository-corporation-survey-billion-shillings-june-2014-f3c8d6c4 Depository Corporation Survey Billion Shillings June 2014 | Africa (Uganda Bureau of Statistics) 156 rows - 1 Africa country/area - 2014-2019 - 1 indicator - Engineered by Electric Sheep Africa TL;DR This dataset contains 156 rows from Uganda Bureau of Statistics, covering Depository Corporation Survey Billion Shillings June 2014. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-depository-corporation-survey-billion-shillings-june-2014-f3c8d6c4.tabulartabular-regressionn<1K0 likes13 downloads1mo agoHugging Face26siberian-lang-lab /evenki-rus-parallel-corporatabular1K<n<10K0 likes12 downloads10mo agoHugging Face27electricsheepafrica /africa-owid-revenue-from-corporate-income-taxes-gdp Revenue From Corporate Income Taxes Gdp | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory) Size category: n<1K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-revenue-from-corporate-income-taxes-gdp.tabulartabular-classificationn<1K0 likes12 downloads1mo agoHugging Face28electricsheepeurope /europe-owid-revenue-from-corporate-income-taxes-gdp Revenue From Corporate Income Taxes Gdp | Europe (Our World in Data) 🇪🇺 1,145 observations · 40 Europe countries · 1980–2017 · Repackaged by Electric Sheep Europe TL;DR This dataset contains 1,145 observations of Revenue From Corporate Income Taxes Gdp data across 40 Europe countries, spanning 1980–2017. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Revenue From Corporate Income Taxes Gdp… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-revenue-from-corporate-income-taxes-gdp.tabulartabular-classification1K<n<10K0 likes9 downloads4mo agoHugging Face29Shoriful025 /corporate_esg_risk_analyticstabularn<1K0 likes8 downloads9mo agoHugging Face30zndx /sdg-corpora-v0.3 SDG Ontology-Grounded Synthetic Corpus — v0.3 (first cut) A verifiable, attribution-clean synthetic textbook corpus for relational data-governance metadata, grounded in a BFO 2020 / CCO ontology. Unlike web-retrieval synthetic corpora, every chapter is generated from deterministic ontology axioms (not scraped text), populates a deterministic relational schema, and ships with the generator's reasoning trace. Part of the Aegir project; the source artifacts live in… See the full description on the dataset page: https://huggingface.co/datasets/zndx/sdg-corpora-v0.3.tabulartext-classification10K<n<100K0 likes8 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.