datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sera-4.5A-Full-T1-v3
laion/Sera-4.5A-Full-T1-v3
Subset of allenai/Sera-4.5A-Full-T1.
Size: 72,118 rows (full dataset: 72,118 rows).
Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages
field (as JSON string), instance_id, rollout_patch, func_name, func_path,
problem_statement, target_patch, docker_image. Adds a source field pointing
back to the parent dataset.
Each assistant message carries a native tool_calls array (OpenAI tool-calling format)
and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3.sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
Nemotron Terminal SFT reproduction evaluation artifacts
This repository contains the complete Harbor artifact tree for the 300-trial
OpenThoughts-TBLite evaluation of
laion/sft-repro-thinking-step630-nemotron-terminal-step1888.
The checkpoint was trained from the Grug stage-2 thinking checkpoint on the
Nemotron Terminal corpus for 1,888 steps.
Result
Measure
Value
Attempted / completed
300 / 300
Verifier-scoreable
259 (86.33%)
Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.Sera-4.5A-Full-T1-v3-1000
laion/Sera-4.5A-Full-T1-v3-1000
Subset of allenai/Sera-4.5A-Full-T1.
Size: 1,000 rows (full dataset: 72,118 rows).
Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages
field (as JSON string), instance_id, rollout_patch, func_name, func_path,
problem_statement, target_patch, docker_image. Adds a source field pointing
back to the parent dataset.
Each assistant message carries a native tool_calls array (OpenAI tool-calling format)
and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-1000.Sera-4.5A-Full-T1-v3-3160
laion/Sera-4.5A-Full-T1-v3-3160
Subset of allenai/Sera-4.5A-Full-T1.
Size: 3,160 rows (full dataset: 72,118 rows).
Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages
field (as JSON string), instance_id, rollout_patch, func_name, func_path,
problem_statement, target_patch, docker_image. Adds a source field pointing
back to the parent dataset.
Each assistant message carries a native tool_calls array (OpenAI tool-calling format)
and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-3160.Sera-4.6-Lite-T2-v4-1000
laion/Sera-4.6-Lite-T2-v4-1000
Row-subset of allenai/Sera-4.6-Lite-T2
(the dataset upstream SERA-8B was trained on), with OpenAI tool_calls pre-rendered
into the content string as Hermes/Qwen3-style <tool_call>...</tool_call> wire tokens
and tool responses wrapped as <tool_response>...</tool_response>.
This mirrors Ai2's sera/datagen/data/postprocess/utils.py::transform_traj_hermes
(default tool_call_format: "hermes") which is the missing step between the public
Sera-4.6-Lite-T2… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.6-Lite-T2-v4-1000.Sera-4.5A-Full-T1-v3-10000
laion/Sera-4.5A-Full-T1-v3-10000
Subset of allenai/Sera-4.5A-Full-T1.
Size: 10,000 rows (full dataset: 72,118 rows).
Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages
field (as JSON string), instance_id, rollout_patch, func_name, func_path,
problem_statement, target_patch, docker_image. Adds a source field pointing
back to the parent dataset.
Each assistant message carries a native tool_calls array (OpenAI tool-calling format)
and a train: bool flag… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-10000.CoderForge-Preview-v6-1000
laion/CoderForge-Preview-v6-1000
Row-subset of togethercomputer/CoderForge-Preview
(trajectories split, filtered_reward1), rendered into Qwen3-compatible
think-first OpenHands-XML wire format.
Why v6?
v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced
garbage at eval time (8888..., 0.0.0.0...) despite clean training losses.
Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first
token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-1000.Sera-4.5A-Full-T1-v3-316
laion/Sera-4.5A-Full-T1-v3-316
Subset of allenai/Sera-4.5A-Full-T1.
Size: 316 rows (full dataset: 72,118 rows).
Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages
field (as JSON string), instance_id, rollout_patch, func_name, func_path,
problem_statement, target_patch, docker_image. Adds a source field pointing
back to the parent dataset.
Each assistant message carries a native tool_calls array (OpenAI tool-calling format)
and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-316.Sera-4.6-Lite-T2-v4-316
laion/Sera-4.6-Lite-T2-v4-316
Row-subset of allenai/Sera-4.6-Lite-T2
(the dataset upstream SERA-8B was trained on), with OpenAI tool_calls pre-rendered
into the content string as Hermes/Qwen3-style <tool_call>...</tool_call> wire tokens
and tool responses wrapped as <tool_response>...</tool_response>.
This mirrors Ai2's sera/datagen/data/postprocess/utils.py::transform_traj_hermes
(default tool_call_format: "hermes") which is the missing step between the public
Sera-4.6-Lite-T2… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.6-Lite-T2-v4-316.CoderForge-Preview-v6-316
laion/CoderForge-Preview-v6-316
Row-subset of togethercomputer/CoderForge-Preview
(trajectories split, filtered_reward1), rendered into Qwen3-compatible
think-first OpenHands-XML wire format.
Why v6?
v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced
garbage at eval time (8888..., 0.0.0.0...) despite clean training losses.
Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first
token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-316.sera-subset-mixed-316
sera-subset-mixed-316
Random subset of 316 rows drawn from ethanlshen/sera-subset, mixed across the two
upstream stages (stage1 unresolved + stage2 resolved) and shuffled deterministically.
Source
Upstream: ethanlshen/sera-subset.
Two upstream JSONLs are concatenated:
22972_0.88_stage1_scaling_final_glm46_e2e_1ipf_swesmith_unresolved_ipf_1_atk_rft-think_SYSTEM_SIMPLE.jsonl (22 972 rows)… See the full description on the dataset page: https://huggingface.co/datasets/laion/sera-subset-mixed-316.
