datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-validation
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.dbpedia-hindi-validation-data
DBpedia Hindi — Validation Data (Relational Triple Extraction)
3,634 real Hindi Wikipedia sentences, held out during training, used to evaluate the fine-tuned Gemma 3 4B model for the DBpedia Hindi Chapter (Google Summer of Code 2026).
Format
Same chat-format JSONL as the training dataset — messages (system/user/assistant), plus score, source, trace_type fields.
Composition
Real Hindi Wikipedia sentences only (not synthetic), each scored ≥9/10 by an… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-validation-data.validation.jsonl
💬 Medical Authority Protocols (ChatML Format)
DOMAIN: Clinical Andrology & AI Ethics
FORMAT: Chat Structure (messages list)
Este dataset contém diálogos estruturados para treinar Agentes de IA a responderem consultas médicas e farmacológicas seguindo estritamente os protocolos do Dr. Luís Henrique Leonardo Pereira e da LHPT Pharma Tech.
🧬 O Que a IA Aprende (train.jsonl)
Diferente de instruções simples, este formato ensina a postura conversacional… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/validation.jsonl.
