HuggingFaceFW/fineweb-edu (sample-10BT) | odc-by | quality-filtered educational web text; the strongest general English base for a small model |
HuggingFaceTB/smollm-corpus (cosmopedia-v2) | odc-by | synthetic textbook-style articles; dense knowledge per token, ideal at sub-100M scale |
roneneldan/TinyStories | cdla-sharing-1.0 | simple narratives that give tiny models early grammatical coherence |
HuggingFaceTB/smoltalk (all) | not stated on card (HuggingFaceTB release) | 1M multi-turn conversations curated specifically for sub-1B models |
teknium/OpenHermes-2.5 | not stated on card (compilation of variously-licensed sets) | 1M diverse instruction conversations; breadth of tasks and styles |
HuggingFaceH4/ultrachat_200k | mit | 200k long multi-turn dialogues; strengthens conversational depth |
NousResearch/hermes-function-calling-v1 (func_calling) | apache-2.0 | tool-calling conversations already using <toolcall>/<toolresponse> tags |
NousResearch/hermes-function-calling-v1 (funccallingsingleturn) | apache-2.0 | single-turn tool-calling; clean minimal examples of the call format |
NousResearch/hermes-function-calling-v1 (glaivefunccalling) | apache-2.0 | curated glaive subset in hermes tagging; adds tool diversity |
glaiveai/glaive-function-calling-v2 | apache-2.0 | 113k function-calling dialogues; the volume backbone of tool-use training |
open-thoughts/OpenThoughts-114k | apache-2.0 | reasoning traces split into reasoning_content + answer; only short traces fit a small context |
HuggingFaceH4/ultrafeedback_binarized | mit | canonical chosen/rejected preference pairs; schema-identical to our DPO contract |
lavita/ChatDoctor-HealthCareMagic-100k | not stated on card (mirror of ChatDoctor data) | real patient-question medical QA for the domain-adapter LoRA demo |
openai/gsm8k (main) | mit | grade-school word problems with verifiable numeric answers; ground truth for agentic RL |
cais/mmlu (all) | mit | exam-format alignment; validation split only, never test |
cais/mmlu (all) | mit | exam-format alignment; dev split only, never test |
allenai/ai2_arc (ARC-Easy) | cc-by-sa-4.0 | science multiple choice; train split only |
allenai/ai2_arc (ARC-Challenge) | cc-by-sa-4.0 | harder science multiple choice; train split only |
allenai/openbookqa (main) | unknown (per dataset card) | open-book science questions; train split only |