smoldataenvs
SmolDataEnvs
📈 SmolDataEnvs
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
Data-analysis tasks as a plain, load-and-go dataset: no runtime, no framework required. Each
row is one self-contained task: a real tabular dataset, a question about it, and a gold answer… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs.SmolDataEnvs-sft
🛠️ SmolDataEnvs: SFT
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
4,677 worked examples of an agent doing data science the right way. Each row is a complete,
verified-correct trajectory: read the question, poke at the data with a shell tool… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-sft.SmolDataEnvs-harbor-train
📊 SmolDataEnvs: Harbor (train)
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
The training suite: 5,000 hands-on data-analysis tasks. Each one drops an agent into a sandbox with a
real dataset and a question, and asks it to explore the data… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-harbor-train.SmolDataEnvs-harbor-test
📊 SmolDataEnvs: Harbor (test)
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
The held-out benchmark: 250 tasks the model never trains on, weighted towards the hard end on
purpose. This is the suite to quote a number from.
Packaged in Harbor format… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-harbor-test.SmolDataEnvs-harbor-eval
📊 SmolDataEnvs: Harbor (eval)
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
The validation suite: 144 tasks, small enough to run every few hundred training steps without
the eval becoming the expensive part of the loop.
Packaged in Harbor format… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-harbor-eval.
