CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SotirisLegkas /kalamaki_corporatabular100M<n<1B0 likes3.7k downloads1y agoHugging Face02cais /wmdp-corpora Dataset Card for WMDP Corpora The Weapons of Mass Destruction Proxy (WMDP) Corpora includes all of the corpora used to perform unlearning on WMDP-Bio and WMDP-Cyber. See our paper, website, and GitHub for more details! The corpora are also available at the following mirrors with password wmdpcorpora: 1, 2 The bio forget corpus must be requested separately; please visit this form. cyber-retain-corpus and cyber-forget-corpus The forget and retain corpora consist of… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-corpora.texttext-generation10K<n<100K5 likes2.5k downloads2y agoHugging Face03Blinorot /ALARM-Corpora Dataset Card for ALARM-Corpora Dataset Summary This is the dataset used in the ALARM: Audio-Language Alignment for Reasoning Models paper. It consists of Audio Captions and Reasoning Language Model responses rephrased to sound like they were provided by an audio-understanding model. For more details regarding the dataset and the instructions for obtaining audio files, please refer to our GitHub. Dataset Statistics Audio Type # Elements (M) # Hours (K)… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/ALARM-Corpora.text10M<n<100M0 likes458 downloads7mo agoHugging Face04ZipLime /corporate-actions US Corporate Actions — dividends and splits 391 639 dividends from 3 327 filers · 5 619 splits from 3 814 filers · 2005 to 2026 Built to close a specific hole. A filing states shares and earnings per share as of the day it was made; every price series is adjusted for splits since. Multiply one by the other and the answer is wrong by the split factor — on Deckers that turned a 6.9% earnings yield into 41.7%, a P/E of 1.8. The pipeline lives in recipe/ at the same revision as the… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/corporate-actions.tabulartabular-regression100K<n<1M0 likes250 downloads17h agoHugging Face05siddharthmb /mats-gf-provenance-corpora Provenance-codeword training corpora All training corpora from the eight-experiment provenance codewords program (per-source activation codewords in Qwen3 models). Code, paper, and reproduction scripts: https://github.com/Sid-MB/mats-gf-provenance-codewords Each synthetic corpus ships in full: docs.parquet (training documents), train.parquet, qa.parquet (probe questions incl. phantom-fact controls), generation intermediates (raw/), the sqlite sequence store (seqdb/), and audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.tabular1K<n<10K0 likes243 downloads2mo agoHugging Face06imvladikon /leipzig_corpora_collection Leipzig Corpora Collection The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs. The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.texttext-generation1K<n<10K4 likes227 downloads5mo agoHugging Face07LuisG07 /es_corpora_parliament_processedtext1M<n<10M0 likes206 downloads5y agoHugging Face08JonathanSum /en_corpora_parliament_processedtext1M<n<10M0 likes174 downloads5y agoHugging Face09Jaspernl /The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden" Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io Dataset Summary The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.audioautomatic-speech-recognition10K<n<100K1 likes164 downloads2y agoHugging Face10cais /wmdp-mmlu-auxiliary-corpora Dataset Card for WMDP Auxiliary Corpora This dataset includes the auxiliary corpora used to perform unlearning on the MMLU Auxiliary Benchmark task, from the WMDP paper. See our paper, website, and GitHub for more details! The corpora are also available at the following mirrors with password wmdpauxiliarycorpora: 1, 2 physics-corpus Corpus comprising textbooks in high school and college physics. law-corpus Corpus comprising textbooks in international and… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-mmlu-auxiliary-corpora.text1K<n<10K5 likes146 downloads2y agoHugging Face11AndrewMcDowell /de_corpora_parliament_processedtext100K<n<1M2 likes145 downloads5y agoHugging Face12adorkin /general-instruction-augmented-corporaThis is a reupload of general instruction-augmented corpora in a more accessible format. Please, cite the original repository. text10M<n<100M3 likes145 downloads2y agoHugging Face13huggingface /tokbench-corpora tokbench corpora The input corpora for tokbench, a benchmark that measures tokenizer implementations against each other on the same bytes. Each config is one corpus: ~5 MB of real text chosen to stress a different part of a tokenizer. Nothing here is new text. It is a fixed, pinned, redistributable excerpt of public datasets, packaged so a tokenizer benchmark is reproducible by anyone without re-deriving the inputs. Provenance and licence for every config are in the table below.… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/tokbench-corpora.texttext-generation100K<n<1M0 likes139 downloads5d agoHugging Face14HarrisDePerceptron /sv_corpora_parliament_processedtext1M<n<10M0 likes134 downloads5y agoHugging Face15JonathanSum /sv_corpora_parliament_processedtext1M<n<10M0 likes132 downloads5y agoHugging Face16Iskaj /dutch_corpora_parliament_processedtext1M<n<10M1 likes131 downloads5y agoHugging Face17Plim /fr_corpora_parliament_processedtext1M<n<10M0 likes131 downloads4y agoHugging Face18RuudVelo /nl_corpora_parliament_processedtext1M<n<10M1 likes131 downloads5y agoHugging Face19aptl26 /wmdp_deduped_corporatext1K<n<10K0 likes123 downloads2y agoHugging Face20HarrisDePerceptron /ur_corpora_pibtext10K<n<100K0 likes120 downloads5y agoHugging Face21anushakamath /sv_corpora_parliament_processed_v0text1M<n<10M0 likes118 downloads5y agoHugging Face22azuur /es_corpora_parliament_processedtext1M<n<10M0 likes115 downloads5y agoHugging Face23Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes106 downloads2y agoHugging Face24zzySJ /sv_corpora_parliament_processedtext1M<n<10M0 likes86 downloads4y agoHugging Face25nopperl /corporate-emission-reports Dataset Card for Dataset Name A dataset of 100 corporate sustainability reports with manually extracted scope 1, 2 and 3 greenhouse gas emission values. Dataset Details Dataset Description Data about corporate greenhouse gas emissions is usually published only as part of sustainability report PDF's, which is not a machine-readable format. Interested actors have to manually extract emission data from these reports, which is a tedious and time-consuming process.… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/corporate-emission-reports.tabularn<1K0 likes80 downloads3y agoHugging Face26N1K0L0Z /N1K0_ka_en_corpora N1K0 Georgian-English Corpora A high-quality, heavily sanitized parallel translation dataset containing ~2 Million (1,988,071) parallel sentence pairs for Georgian (ka) and English (en). This dataset is specifically engineered for neural machine translation tasks, sequence-to-sequence models (e.g., Transformers), and cross-attention decoder pretraining/fine-tuning. 🚀 Quick Start & Usage You can load this dataset directly using the Hugging Face datasets library:… See the full description on the dataset page: https://huggingface.co/datasets/N1K0L0Z/N1K0_ka_en_corpora.texttranslation1M<n<10M3 likes80 downloads1mo agoHugging Face27Texttechnologylab /leipzig-corpora-collectiontext1M<n<10M0 likes75 downloads1y agoHugging Face28wrice /sv_corpora_parliament_processedtext1M<n<10M0 likes74 downloads4y agoHugging Face29zakria /it_corpora_parliament_processedtext1M<n<10M0 likes72 downloads4y agoHugging Face30AigizK /bashkir-russian-parallel-corpora Dataset Card for "bashkir-russian-parallel-corpora" How the dataset was assembled. find the text in two languages. it can be a translated book or an internet page (wikipedia, news site) our algorithm tries to match Bashkir sentences with their translation in Russian We give these pairs to people to check @inproceedings{ title={Bashkir-Russian parallel corpora}, author={Iskander Shakirov, Aigiz Kunafin}, year={2023} } texttranslation1M<n<10M16 likes69 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.