CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01instruction-pretrain /general-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.texttext-classification24 likes47k downloads7mo agoHugging Face02SotirisLegkas /kalamaki_corporatabular100M<n<1B0 likes2.6k downloads1y agoHugging Face03cais /wmdp-corpora Dataset Card for WMDP Corpora The Weapons of Mass Destruction Proxy (WMDP) Corpora includes all of the corpora used to perform unlearning on WMDP-Bio and WMDP-Cyber. See our paper, website, and GitHub for more details! The corpora are also available at the following mirrors with password wmdpcorpora: 1, 2 The bio forget corpus must be requested separately; please visit this form. cyber-retain-corpus and cyber-forget-corpus The forget and retain corpora consist of… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-corpora.texttext-generation10K<n<100K5 likes2.5k downloads2y agoHugging Face04mteb /legalbench_corporate_lobbying LegalBenchCorporateLobbying An MTEB dataset Massive Text Embedding Benchmark The dataset includes bill titles and bill summaries related to corporate lobbying. Task category t2t Domains Legal, Written Reference https://huggingface.co/datasets/nguha/legalbench/viewer/corporate_lobbying How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_corporate_lobbying.texttext-retrievaln<1K0 likes905 downloads7mo agoHugging Face05zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes647 downloads4mo agoHugging Face06AtomicChat /calib-corpora calib-corpora A pool of calibration material, the recipes that turn it into a calibration set for one specific model, and the measurement corpora those quants are scored against. This repository is not a corpus. Nothing here is meant to be fed to llama-imatrix as-is except the files under builds/, and each of those was made for one named model and is close to useless for any other. Why it is built this way The first version of this repository was a single… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/calib-corpora.texttext-generation7 likes573 downloads26d agoHugging Face07Blinorot /ALARM-Corpora Dataset Card for ALARM-Corpora Dataset Summary This is the dataset used in the ALARM: Audio-Language Alignment for Reasoning Models paper. It consists of Audio Captions and Reasoning Language Model responses rephrased to sound like they were provided by an audio-understanding model. For more details regarding the dataset and the instructions for obtaining audio files, please refer to our GitHub. Dataset Statistics Audio Type # Elements (M) # Hours (K)… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/ALARM-Corpora.text10M<n<100M0 likes408 downloads7mo agoHugging Face08Mohith202 /lma-individual-project-corporatabular10K<n<100K0 likes384 downloads6d agoHugging Face09castorini /odqa-wiki-corpora Dataset Card for Open-Domain Question Answering Wikipedia Corpora Dataset Description Dataset Summary The Wikipedia corpus variants provided can serve as knowledge sources for question-answering systems based on a retriever–reader pipeline. These corpus variants and their corresponding experiments are described further in the paper entitled: Pre-Processing Matters! Improved Wikipedia Corpora for Open-Domain Question Answering. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/castorini/odqa-wiki-corpora.textquestion-answering10M<n<100M0 likes383 downloads4y agoHugging Face10imvladikon /leipzig_corpora_collection Leipzig Corpora Collection The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs. The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.texttext-generation1K<n<10K4 likes271 downloads5mo agoHugging Face11epiq-ai-labs /epiq-laer-corporate-benchmark Epiq LAER CorporateBench Harbor release of Epiq LAER CorporateBench: Enterprise Knowledge. Epiq LAER CorporateBench is CorporateBench in its Harbor configuration. It packages the five CorporateBench capabilities as 128 scored Harbor tasks with 1,132 graded cases over the four synthetic companies of the original paper, with five development tasks alongside. What is this? LLMs are increasingly able to answer complex questions about enterprise-scale document… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/epiq-laer-corporate-benchmark.textquestion-answering1K<n<10K0 likes270 downloads6d agoHugging Face12Corp-o-Rate-Community /entity-references Entity References Database A comprehensive entity database for organizations, people, roles, and locations with embedding-based semantic search. Built from authoritative sources (GLEIF, SEC, Companies House, Wikidata) for entity linking and named entity disambiguation. Dataset Summary This dataset provides fast lookup and qualification of named entities using vector similarity search. It stores records from authoritative global sources with embeddings generated by… See the full description on the dataset page: https://huggingface.co/datasets/Corp-o-Rate-Community/entity-references.tabulartext-classificationn<1K0 likes264 downloads5mo agoHugging Face13taiwan-corpora /twsyllables twsyllables — Taiwanese Mandarin syllable acoustics Per-syllable acoustic reference data for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW): 37,947 measured syllable tokens, position-sensitive acoustic templates for 1,491 syllable×tone types, voice-onset-time norms for all 17 obstruent initials, and a between-speaker variability model estimated over 271 speakers. Every number was measured from native Taiwanese recordings by one reproducible pipeline; no figure in this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twsyllables.tabularother1K<n<10K0 likes247 downloads29d agoHugging Face14siddharthmb /mats-gf-provenance-corpora Provenance-codeword training corpora All training corpora from the eight-experiment provenance codewords program (per-source activation codewords in Qwen3 models). Code, paper, and reproduction scripts: https://github.com/Sid-MB/mats-gf-provenance-codewords Each synthetic corpus ships in full: docs.parquet (training documents), train.parquet, qa.parquet (probe questions incl. phantom-fact controls), generation intermediates (raw/), the sqlite sequence store (seqdb/), and audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.tabular1K<n<10K0 likes227 downloads2mo agoHugging Face15LuisG07 /es_corpora_parliament_processedtext1M<n<10M0 likes210 downloads5y agoHugging Face16SimbaMaw1547 /south-african-monolingual-corpora-jsonl South African Languages Pretraining Dataset This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections. The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity Languages Included Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.text1M<n<10M0 likes209 downloads1y agoHugging Face17JonathanSum /en_corpora_parliament_processedtext1M<n<10M0 likes175 downloads5y agoHugging Face18AndrewMcDowell /de_corpora_parliament_processedtext100K<n<1M2 likes174 downloads5y agoHugging Face19taiwan-corpora /twngrams Taiwanese Mandarin web n-grams Word 1–4-gram counts for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW), computed over the Taiwan slice of a large web crawl after variety filtering by twfilter 0.1.0 with the published twfilter-tables: every sentence behind these counts passed the 教育部 character-inventory gate, the simplified-character round-trip, the mainland-orthography, mainland-lexicon, written-Cantonese, Hong Kong and Singapore detectors, and block-level evidence of Taiwan-specific… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twngrams.texttext-generation1M<n<10M0 likes158 downloads29d agoHugging Face20docketx /docketrouter-legal-corpora DocketRouter Legal Corpora Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Verbatim, provenance-carrying legal text published by DocketRouter, the legal-grounding API from DocketX, so anyone can build on it. Every row carries its official source URL and retrieval date. The… See the full description on the dataset page: https://huggingface.co/datasets/docketx/docketrouter-legal-corpora.texttext-retrieval1K<n<10K0 likes158 downloads2d agoHugging Face21ZipLime /corporate-actions US Corporate Actions — dividends and splits 391 639 dividends from 3 327 filers · 5 619 splits from 3 814 filers · 2005 to 2026 Built to close a specific hole. A filing states shares and earnings per share as of the day it was made; every price series is adjusted for splits since. Multiply one by the other and the answer is wrong by the split factor — on Deckers that turned a 6.9% earnings yield into 41.7%, a P/E of 1.8. The pipeline lives in recipe/ at the same revision as the… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/corporate-actions.tabulartabular-regression100K<n<1M0 likes157 downloads2d agoHugging Face22taiwan-corpora /twfilter-tables twfilter reference tables The tables that decide whether a span of traditional-Chinese text is Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW) rather than Hong Kong Cantonese, mainland text converted to traditional characters, or literary Chinese. Plain text, one record per line, tab-separated where a record has fields, LC_ALL=C sort order, UTF-8, LF. Consumed by twfilter 0.1.0, where this directory is vendored byte-for-byte and MANIFEST.json is verified by its test suite. Usable… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twfilter-tables.texttext-classification10K<n<100K0 likes151 downloads29d agoHugging Face23cais /wmdp-mmlu-auxiliary-corpora Dataset Card for WMDP Auxiliary Corpora This dataset includes the auxiliary corpora used to perform unlearning on the MMLU Auxiliary Benchmark task, from the WMDP paper. See our paper, website, and GitHub for more details! The corpora are also available at the following mirrors with password wmdpauxiliarycorpora: 1, 2 physics-corpus Corpus comprising textbooks in high school and college physics. law-corpus Corpus comprising textbooks in international and… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-mmlu-auxiliary-corpora.text1K<n<10K5 likes148 downloads2y agoHugging Face24Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes138 downloads2y agoHugging Face25adorkin /general-instruction-augmented-corporaThis is a reupload of general instruction-augmented corpora in a more accessible format. Please, cite the original repository. text10M<n<100M3 likes137 downloads2y agoHugging Face26Moo /korean-parallel-corporatexttranslation10K<n<100K21 likes136 downloads4y agoHugging Face27Jaspernl /The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden" Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io Dataset Summary The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.audioautomatic-speech-recognition10K<n<100K1 likes136 downloads2y agoHugging Face28HarrisDePerceptron /sv_corpora_parliament_processedtext1M<n<10M0 likes132 downloads5y agoHugging Face29Iskaj /dutch_corpora_parliament_processedtext1M<n<10M1 likes132 downloads5y agoHugging Face30JonathanSum /sv_corpora_parliament_processedtext1M<n<10M0 likes130 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.