CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01instruction-pretrain /general-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.texttext-classification24 likes38k downloads7mo agoHugging Face02GuyHam /c-sac-corpora C-SAC LibriTTS-R training subset Deterministically selected and resampled speech from mythicinfinity/libritts_r for the C-SAC causal speech-codec program. The package retains source revision, Parquet shard, row, utterance, transcript, and content hashes. LibriTTS-R is distributed under CC BY 4.0; downstream users remain responsible for attribution. Only prefixes with a hash-bound _COMPLETE.json sentinel are admissible. audiotext-to-speech0 likes14k downloads2d agoHugging Face03lucius1022 /DeMix_Corpora Dataset Card for DeMix Corpora DeMix 📄 Paper: Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training 🤗 Dataset: DeMix Corpora 🐱 Github: Demix Dataset Details Dataset Description DeMix Corpora (15T original tokens and 22T mixture tokens) serves as a comprehensive, high-quality, large-scale, and carefully mixed resource that can be directly employed for pre-training. (2026.2.7: This is an… See the full description on the dataset page: https://huggingface.co/datasets/lucius1022/DeMix_Corpora.tabularn<1K3 likes3.4k downloads7mo agoHugging Face04SotirisLegkas /kalamaki_corporatabular100M<n<1B0 likes2.4k downloads1y agoHugging Face05cais /wmdp-corpora Dataset Card for WMDP Corpora The Weapons of Mass Destruction Proxy (WMDP) Corpora includes all of the corpora used to perform unlearning on WMDP-Bio and WMDP-Cyber. See our paper, website, and GitHub for more details! The corpora are also available at the following mirrors with password wmdpcorpora: 1, 2 The bio forget corpus must be requested separately; please visit this form. cyber-retain-corpus and cyber-forget-corpus The forget and retain corpora consist of… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-corpora.texttext-generation10K<n<100K5 likes2.3k downloads2y agoHugging Face06mteb /legalbench_corporate_lobbying LegalBenchCorporateLobbying An MTEB dataset Massive Text Embedding Benchmark The dataset includes bill titles and bill summaries related to corporate lobbying. Task category t2t Domains Legal, Written Reference https://huggingface.co/datasets/nguha/legalbench/viewer/corporate_lobbying How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_corporate_lobbying.texttext-retrievaln<1K0 likes889 downloads7mo agoHugging Face07SVECTOR-CORPORATION /ThinkChain-20M We are excited to announce the release of SVECTOR-CORPORATION/ThinkChain-20M, a synthetic reasoning dataset containing over 22 million general reasoning questions and responses generated using Spec-T1. While multiple efforts exist to build open reasoning datasets for math and code tasks, there has been a gap in large datasets covering diverse non code/math topics such as social and natural sciences, education, creative writing, and general conversations. This dataset fills that gap. Note: The… See the full description on the dataset page: https://huggingface.co/datasets/SVECTOR-CORPORATION/ThinkChain-20M.text-generation10M<n<100M4 likes640 downloads1y agoHugging Face08zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes633 downloads4mo agoHugging Face09instruction-pretrain /medicine-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the instruction-augmented corpora in biomedicine domain used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train language models. The instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/medicine-instruction-augmented-corpora.text-classification13 likes603 downloads7mo agoHugging Face10AtomicChat /calib-corpora calib-corpora A pool of calibration material, the recipes that turn it into a calibration set for one specific model, and the measurement corpora those quants are scored against. This repository is not a corpus. Nothing here is meant to be fed to llama-imatrix as-is except the files under builds/, and each of those was made for one named model and is close to useless for any other. Why it is built this way The first version of this repository was a single… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/calib-corpora.texttext-generation7 likes597 downloads25d agoHugging Face11Blinorot /ALARM-Corpora Dataset Card for ALARM-Corpora Dataset Summary This is the dataset used in the ALARM: Audio-Language Alignment for Reasoning Models paper. It consists of Audio Captions and Reasoning Language Model responses rephrased to sound like they were provided by an audio-understanding model. For more details regarding the dataset and the instructions for obtaining audio files, please refer to our GitHub. Dataset Statistics Audio Type # Elements (M) # Hours (K)… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/ALARM-Corpora.text10M<n<100M0 likes404 downloads6mo agoHugging Face12Mohith202 /lma-individual-project-corporatabular10K<n<100K0 likes384 downloads6d agoHugging Face13corporationgoorac /marketingVoiceaudion<1K0 likes382 downloads3mo agoHugging Face14castorini /odqa-wiki-corpora Dataset Card for Open-Domain Question Answering Wikipedia Corpora Dataset Description Dataset Summary The Wikipedia corpus variants provided can serve as knowledge sources for question-answering systems based on a retriever–reader pipeline. These corpus variants and their corresponding experiments are described further in the paper entitled: Pre-Processing Matters! Improved Wikipedia Corpora for Open-Domain Question Answering. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/castorini/odqa-wiki-corpora.textquestion-answering10M<n<100M0 likes381 downloads4y agoHugging Face15taiwan-corpora /twsyllables twsyllables — Taiwanese Mandarin syllable acoustics Per-syllable acoustic reference data for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW): 37,947 measured syllable tokens, position-sensitive acoustic templates for 1,491 syllable×tone types, voice-onset-time norms for all 17 obstruent initials, and a between-speaker variability model estimated over 271 speakers. Every number was measured from native Taiwanese recordings by one reproducible pipeline; no figure in this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twsyllables.tabularother1K<n<10K0 likes277 downloads28d agoHugging Face16imvladikon /leipzig_corpora_collection Leipzig Corpora Collection The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs. The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.texttext-generation1K<n<10K4 likes270 downloads5mo agoHugging Face17Corp-o-Rate-Community /entity-references Entity References Database A comprehensive entity database for organizations, people, roles, and locations with embedding-based semantic search. Built from authoritative sources (GLEIF, SEC, Companies House, Wikidata) for entity linking and named entity disambiguation. Dataset Summary This dataset provides fast lookup and qualification of named entities using vector similarity search. It stores records from authoritative global sources with embeddings generated by… See the full description on the dataset page: https://huggingface.co/datasets/Corp-o-Rate-Community/entity-references.tabulartext-classificationn<1K0 likes265 downloads5mo agoHugging Face18epiq-ai-labs /epiq-laer-corporate-benchmark Epiq LAER CorporateBench Harbor release of Epiq LAER CorporateBench: Enterprise Knowledge. Epiq LAER CorporateBench is CorporateBench in its Harbor configuration. It packages the five CorporateBench capabilities as 128 scored Harbor tasks with 1,132 graded cases over the four synthetic companies of the original paper, with five development tasks alongside. What is this? LLMs are increasingly able to answer complex questions about enterprise-scale document… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/epiq-laer-corporate-benchmark.textquestion-answering1K<n<10K0 likes249 downloads6d agoHugging Face19epiq-ai-labs /corporatebench CorporateBench Dataset release for CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases. CorporateBench evaluates information extraction, retrieval, and question answering over four synthetic corporate corpora ranging from 353 to 232,692 released documents. The corpora are generated from temporally evolving knowledge bases, providing deterministic ground truth across related documents. Dataset Viewer https://corporatebench.epiqai.com/… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/corporatebench.question-answering100K<n<1M0 likes238 downloads12d agoHugging Face20siddharthmb /mats-gf-provenance-corpora Provenance-codeword training corpora All training corpora from the eight-experiment provenance codewords program (per-source activation codewords in Qwen3 models). Code, paper, and reproduction scripts: https://github.com/Sid-MB/mats-gf-provenance-codewords Each synthetic corpus ships in full: docs.parquet (training documents), train.parquet, qa.parquet (probe questions incl. phantom-fact controls), generation intermediates (raw/), the sqlite sequence store (seqdb/), and audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.tabular1K<n<10K0 likes231 downloads2mo agoHugging Face21SimbaMaw1547 /south-african-monolingual-corpora-jsonl South African Languages Pretraining Dataset This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections. The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity Languages Included Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.text1M<n<10M0 likes219 downloads1y agoHugging Face22LuisG07 /es_corpora_parliament_processedtext1M<n<10M0 likes208 downloads5y agoHugging Face23JonathanSum /en_corpora_parliament_processedtext1M<n<10M0 likes173 downloads5y agoHugging Face24AndrewMcDowell /de_corpora_parliament_processedtext100K<n<1M2 likes171 downloads5y agoHugging Face25mxguru1 /council-training-corpora0 likes168 downloads3mo agoHugging Face26taiwan-corpora /twngrams Taiwanese Mandarin web n-grams Word 1–4-gram counts for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW), computed over the Taiwan slice of a large web crawl after variety filtering by twfilter 0.1.0 with the published twfilter-tables: every sentence behind these counts passed the 教育部 character-inventory gate, the simplified-character round-trip, the mainland-orthography, mainland-lexicon, written-Cantonese, Hong Kong and Singapore detectors, and block-level evidence of Taiwan-specific… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twngrams.texttext-generation1M<n<10M0 likes161 downloads28d agoHugging Face27taiwan-corpora /twfilter-tables twfilter reference tables The tables that decide whether a span of traditional-Chinese text is Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW) rather than Hong Kong Cantonese, mainland text converted to traditional characters, or literary Chinese. Plain text, one record per line, tab-separated where a record has fields, LC_ALL=C sort order, UTF-8, LF. Consumed by twfilter 0.1.0, where this directory is vendored byte-for-byte and MANIFEST.json is verified by its test suite. Usable… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twfilter-tables.texttext-classification10K<n<100K0 likes160 downloads28d agoHugging Face28slashgg /hlwm-corpora HLWM training corpora The training data behind the Hierarchical Latent Workspace Model program — a twenty-day, ten-experiment preregistered attempt to build a latent-workspace language model on a frozen Qwen3-0.6B decoder. All three proposed mechanisms failed their preregistered gates. Both papers are negative-results reports. This dataset is published so the record is checkable, not because it produced a working system. Papers, code and full experimental record:… See the full description on the dataset page: https://huggingface.co/datasets/slashgg/hlwm-corpora.text-generation1K<n<10K0 likes154 downloads16d agoHugging Face29cais /wmdp-mmlu-auxiliary-corpora Dataset Card for WMDP Auxiliary Corpora This dataset includes the auxiliary corpora used to perform unlearning on the MMLU Auxiliary Benchmark task, from the WMDP paper. See our paper, website, and GitHub for more details! The corpora are also available at the following mirrors with password wmdpauxiliarycorpora: 1, 2 physics-corpus Corpus comprising textbooks in high school and college physics. law-corpus Corpus comprising textbooks in international and… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-mmlu-auxiliary-corpora.text1K<n<10K5 likes150 downloads2y agoHugging Face30FrancophonIA /Parallel_corpora [!NOTE] Dataset origin: https://portulanclarin.net/repository/browse/parallel-corpora-finely-aligned-subsentencial-granularity/aa90dbbeb0ab11ea8dc202420a00040310be6f259e694e659c10c1d212b389f0/ Description Text corpus for bilingual concordancing, single- and multi-word translation extraction, machine translation. Languages: cs-pt, de-pt, en-pt, es-pt, fr-pt, it-pt, and pt-sk. Domain: Law and Health. 0 likes147 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.