datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.RLVE-Qwen3-1.7B-Pass1-Rollouts
RLVE teacher rollouts — Qwen3-1.7B (pass@1)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-1.7B
Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym
environments (counting / combinatorics / optimization tasks)
Sampling: 1 sample/question (pass@1) = 9000 records,
temperature 0.7, max 4096 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.Qwen-3-1.7B-with-Reasoning-x500
Qwen-3-1.7B-with-Reasoning-x500
This is version v1 - we continue updating and upscaling this dataset!
Overview
This is a high-quality synthetic dataset consisting of 500 diverse samples generated by Qwen 3 1.7B.
The goal of this dataset is to provide clean, direct, and logical reasoning traces for distilling larger model capabilities into Small Language Models (SLMs) like my Apex models or those of CompactAI.
Dataset Structure
The data is provided… See the full description on the dataset page: https://huggingface.co/datasets/LH-Tech-AI/Qwen-3-1.7B-with-Reasoning-x500.Qwen-3-1.7B-with-Reasoning-x100
Qwen-3-1.7B-with-Reasoning-x100
This is version v1 - we continue updating and upscaling this dataset!
Overview
This is a high-quality synthetic dataset consisting of 100 diverse samples generated by Qwen 3 1.7B.
The goal of this dataset is to provide clean, direct, and logical reasoning traces for distilling larger model capabilities into Small Language Models (SLMs) like my Apex models or those of CompactAI.
Dataset Structure
The data is provided… See the full description on the dataset page: https://huggingface.co/datasets/LH-Tech-AI/Qwen-3-1.7B-with-Reasoning-x100.rose_code-Qwen3-1.7B-Pass8-Rollouts
rose_code rollouts — Qwen3-1.7B (pass@8)
Model rollouts on the rose_code test split, for the OPD coding pipeline.
Model / sampler: Qwen3-1.7B
Source prompts: CL-From-Nothing/rose_code test split — 408 competitive-programming questions (codeforces-style)
Sampling: 8 samples/question (pass@8) = 3264 records, temperature 0.7, max 16384 new tokens, max_model_len 32000
Rewards: DeepCoder code verifier (deepcoder_reward_fn.py) — 1.0 if the generated program passes all unit tests… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code-Qwen3-1.7B-Pass8-Rollouts.whetstone-Qwen-1.7B-generations
whetstone-Qwen-1.7B-generations
Paired verbose and compact-register reasoning traces for 2,414 maths
problems, with per-trace follow-ability scores.
Each row holds a problem, the long chain-of-thought Qwen3-1.7B produced for it,
a compact-notation rewrite of that same reasoning, token counts for both, and
the scores used to measure how followable the compact version is to Qwen3-1.7B.
11,174,460 original think tokens → 750,087 compressed (14.9×).
Selection
Every… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/whetstone-Qwen-1.7B-generations.rlve_offline_20K_POPE_prefix_pass1_qwen3-1.7b
RLVE offline-20K POPE-prefix completions — Qwen3-1.7B (pass1)
Prefix-conditioned completions generated by Qwen3-1.7B over the
rlve_offline_20K_POPE_prefix prompt set (20000 records, 1 sample/prompt).
Produced by SLURM job 6580578 (vLLM, tp=2), 2026-06-15.
Fields
index, sample_id, prompt, prefix, response, answer, rewards
⚠️ Caveat on rewards
The inline rewards field is all 0.0 — this is the known inline-Gym-verifier
artifact (same as the old… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rlve_offline_20K_POPE_prefix_pass1_qwen3-1.7b.RLVE-Test-Qwen3-1.7B-GRPO-step70-Pass8
RLVE test eval — GRPO step70 (pass@8)
Evaluation rollouts on the RLVE test split.
Model: grpo_train_Qwen3-1.7B-SFT-rlve-20K-1epoch (GRPO, step 70)
Source prompts: RLVE test split — 180 questions (RLVE-Eval Gym environments)
Sampling: 8 samples/question (pass@8) = 1440 records, temperature 0.7,
max 16384 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Record-level accuracy (reward>0): 216 / 1440 = 15.0%, mean reward -0.677
pass@8 (>=1 of 8… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Test-Qwen3-1.7B-GRPO-step70-Pass8.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.RLVE-Test-Qwen3-1.7B-SFT-warmup-Pass8
RLVE test eval — SFT warmup (pass@8)
Evaluation rollouts on the RLVE test split.
Model: Qwen3-1.7B-SFT-rlve-20K-1epoch (RL-warmup baseline, pre-RL)
Source prompts: RLVE test split — 180 questions (RLVE-Eval Gym environments)
Sampling: 8 samples/question (pass@8) = 1440 records, temperature 0.7,
max 16384 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Record-level accuracy (reward>0): 176 / 1440 = 12.2%, mean reward -0.724
pass@8 (>=1 of 8… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Test-Qwen3-1.7B-SFT-warmup-Pass8.
