datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3_4b_instruct_lcbv6_rsa_pop_32_k_4_steps_10_s_65_e_1310_25_rsa_pop_32_k_4_steps_10_v2Math-steptok-steps-mcvalue-train-part4-of-525_50_rsa_pop_32_k_4_steps_10_v275_100_rsa_pop_32_k_4_steps_10_v2QwQ-Long-CoT-10k-subset-llama3.1-8b-Inst-GPT4-Step-Perturbation-8-rejectsbigmath-custom-checkpoint-step-by-step-confidence-ckpt-8192-v450_75_rsa_pop_32_k_4_steps_10_v2qwen3_4b_instruct_start_325_end_350_rsa_pop_32_k_4_steps_10_timeout_53k_forcing_400_mask25_step4_022525qwen3_4b_instruct_start_25_end_50_rsa_pop_32_k_4_steps_10_5_timeout_5qwen3_4b_instruct_start_400_end_425_rsa_pop_32_k_4_steps_10_timeout_53k-forced-p301-final-022825-step4-collatedqwen3_4b_instruct_start_250_end_275_rsa_pop_32_k_4_steps_10_timeout_5bilevel-grpo-4b-step110-bestofn-o4_mini
CodeRM LLM Judge Trajectories (Multi-Model)
Dataset Description
This dataset contains selection trajectories from multiple model solutions evaluated by an LLM judge.
Available Splits
Split
Trajectories
o4_mini
786
Total: 786 trajectories across 1 models
Usage
from datasets import load_dataset
# Load all splits
ds = load_dataset("t2ance/bilevel-grpo-4b-step110-bestofn-o4_mini")
# Load specific split
ds =… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/bilevel-grpo-4b-step110-bestofn-o4_mini.deepcoder-train-qwen4b-instr-tok-mean_bs64_roll4-run3_step50-codeonly_truncation_inferencepromptres_gptoss20b_verbalized_1_step_None_0.7_4096_gpt-4_1-mini-2025-04-1425_50_rsa_pop_32_k_4_steps_100__rsa_pop_32_k_4_steps_10_testqwen3_4b_instruct_start_75_end_100_rsa_pop_32_k_4_steps_10_timeout_5qwen3_4b_instruct_start_0_end_25_rsa_pop_32_k_4_steps_10_timeout_5qwen3_4b_instruct_start_125_end_150_rsa_pop_32_k_4_steps_10_timeout_5qwen3_4b_instruct_start_100_end_125_rsa_pop_32_k_4_steps_10_timeout_53k-forcing-clipped-022225-step4-collatedeval_smolvla_sorting_full_ft_121ep_step160000_B_IID_dir4_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1322,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Harumo/eval_smolvla_sorting_full_ft_121ep_step160000_B_IID_dir4_v1.3k-forcing-400-022325-mask25-ttt8-seed8-step43k-unsolved-priority-022525-step4-collatedSTEPS__r1_8d_eval__v4
Dataset card for STEPS__r1_8d_eval__v4
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"question": "What is the solution to the long multiplication equation below?\n\n97511989 x 29677209\n\nThink step by step.",
"solution": "2893883677558701",
"eval_prompt": [
{
"content": "You are a reasoning assistant. Given a chain-of-thought reasoning trace, output a JSON object with a single key 'steps' whose value… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/STEPS__r1_8d_eval__v4.res_gptoss20b_original_1_step_None_0.7_4096_gpt-4_1-mini-2025-04-145pc-short-aug-4-step-1-1-batch1
