datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Domofon-Cot-Conversations-700k
Domofon-Cot-Conversations-700k
Synthetic XML conversation data for training small language models on reasoning,
instruction following, XML formatting, and tool-use traces.
Repository: domofon/Domofon-Cot-Conversations-700k
What is inside
The dataset contains cleaned generated XML conversations from six families:
conv: multi-turn factual conversations with tool-use traces.
instruct: text-processing instructions, including deterministic count tool calls.
ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.processed-pretrainbritannica-pretrain
britannica-pretrain
Private pretrain documents from Encyclopaedia Britannica scans (Internet Archive OCR *_djvu.txt), indexed via biglam/britannica-illustrated-pages manifest. One Arrow row is one article. No chunking. Trainer reads text.
<encyclopedia>
<meta source="britannica" edition="11th" year="1911" />
<title>Elasticity</title>
<body>
...
</body>
</encyclopedia>
<|endoftext|>
<|endoftext|> is SmolLM2-1.7B eos. OCR is noisy. Duplicate library copies were collapsed by… See the full description on the dataset page: https://huggingface.co/datasets/domofon/britannica-pretrain.
