datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-bert-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.TinyPython
TinyPython Tasks
TinyPython is a synthetic Python dataset inspired by the idea behind TinyStories: if the data distribution is narrow, clean, and high quality, even very small language models can learn useful structure.
Instead of broad repository code or competitive-programming solutions, TinyPython focuses on short natural-language programming tasks paired with complete, typed, standalone Python functions.
The goal is to provide a compact instruction-to-code corpus for… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/TinyPython.voice-light-tool-use-synthetic
Voice Light Teacher-Led Tool-Use Synthetic
This repository contains the current canonical synthetic source dataset for Voice Light's
conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English
conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation
has four user turns so follow-up requests can depend naturally on prior turns and tool results.
The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.Turkish-Municipality-Instruction-Tuning-Dataset
