CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mrinaldi /UsenetArchiveIT Usenet Archive IT Dataset 🇮🇹 Description Dataset Content This dataset contains Usenet posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing. The only preprocessing conducted on the text was the removal of two conversations in which the VBS source code of the malicious script "ILOVEYOU" was present as it was shared by two users for didactical… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/UsenetArchiveIT.tabulartext-generation10M<n<100M11 likes1.3k downloads2y agoHugging Face02MRiabov /cadqa-rl-2000 CAD-QA RL Training Set (2000 samples) Verifiable-reward RL training rows for CAD geometry reasoning, derived from the CAD-QA benchmark (release_10k train split, cad_browsecomp difficulty). Each row is a closed-book question: a CadQuery script + a question about the resulting geometry. The gold answer is exact-match verifiable. Files data/train/ — 2000 rows data/validation/ — 98 rows (the benchmark's eval-slice rows; excluded from training — do not train on this… See the full description on the dataset page: https://huggingface.co/datasets/MRiabov/cadqa-rl-2000.tabularquestion-answering10K<n<100K0 likes96 downloads2mo agoHugging Face03m-ric /TRM-modified-datamix-tokenized TRM modified datamix (tokenized) Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model) training, built by running data_io — the HRM-Text data pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three deliberate, documented deviations (below). It is emitted in the V1 tokenized dataset format (a single concatenated token pool + per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.tabulartext-generationn<1K0 likes53 downloads3mo agoHugging Face04mrinaldi /TestiMolegated Dataset Card for TestiMole -- A multi-billion tokens Italian text corpus Testimole is a large linguistic resource for Italian obtained through a massive web scraping effort. As of June 2024, it is one of the largest datasets for the Italian language, if not the largest, publicly available, consisting of almost 100B tokens counted with the Tiktoken cl100k BPE tokenizer. It consists mainly of conversational data (Italian Usenet hierarchies, Italian message boards, Italian… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/TestiMole.tabulartext-classification100M<n<1B12 likes41 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.