datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UsenetArchiveIT
Usenet Archive IT Dataset 🇮🇹
Description
Dataset Content
This dataset contains Usenet posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing.
The only preprocessing conducted on the text was the removal of two conversations in which the VBS source code of the malicious script "ILOVEYOU" was present as it was shared by two users for didactical… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/UsenetArchiveIT.cadqa-rl-2000
CAD-QA RL Training Set (2000 samples)
Verifiable-reward RL training rows for CAD geometry reasoning, derived from the
CAD-QA benchmark (release_10k train split, cad_browsecomp difficulty).
Each row is a closed-book question: a CadQuery script + a question about the
resulting geometry. The gold answer is exact-match verifiable.
Files
data/train/ — 2000 rows
data/validation/ — 98 rows (the benchmark's eval-slice rows; excluded
from training — do not train on this… See the full description on the dataset page: https://huggingface.co/datasets/MRiabov/cadqa-rl-2000.TRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.TestiMole
Dataset Card for TestiMole -- A multi-billion tokens Italian text corpus
Testimole is a large linguistic resource for Italian obtained through a massive web scraping effort. As of June 2024, it is one of the largest datasets for the Italian language, if not the largest, publicly available, consisting of almost 100B tokens counted with the Tiktoken cl100k BPE tokenizer. It consists mainly of conversational data (Italian Usenet hierarchies, Italian message boards, Italian… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/TestiMole.
