datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nopm_claude_writing_fixedThis is Nopm/Opus_WritingStruct, reuploaded and properly converted to ShareGPT format.
wsc_fixed
Glue WSC Fixed
This dataset is a port of the official wsc.fixed dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
mistral_tokenized_2048_fixed_shardsfixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
FIXED-Cleaned-Claude-Sonnet-5-Grok-4.5-ChatGPT-5.6-Luna-Qwen-3.8-MAX
Ultra-Clean HF Dataset
492 pairs. Zero residual JSON garbage. High-tier technical SFT data.
gol-rl-fixed-validation-37156495
GoL World Model — Genie at Tiny Scale
Current implementation status is tracked in STATUS_2026-05-19.md.
The repo currently has three executable tracks: the recursive world-model
demos/training path, agentic trajectory collection, and online GRPO training
via train_grpo.py.
A miniature implementation of the Genie world model
architecture using Conway's Game of Life as the substrate.
Goal: Show that resource-constrained researchers can experiment with world model ideas
using a… See the full description on the dataset page: https://huggingface.co/datasets/brysgo/gol-rl-fixed-validation-37156495.DrafterBench-fixed-trajectories
AgentSuite/DrafterBench-fixed-trajectories
Per-model agent trajectory data for DrafterBench-fixed (public release).
Models: 30
Tasks per model: 1,920
One file per model: {model}.jsonl, one JSON object per line.
Fields: model_path, user_model_path, benchmark_name, task_name, sampling_params, user_sampling_params, messages, eval_result, meta.
sampling_params reflect each benchmark's own implementation; values the benchmark leaves unset are recorded as null (provider default).… See the full description on the dataset page: https://huggingface.co/datasets/AgentSuite/DrafterBench-fixed-trajectories.instruct-set-longer-fixedfixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts
fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
aurel_tensors_EOS_fixedrobomme_1cuben_fixedcup_split
robomme_1cuben_fixedcup_split (VideoUnmaskSwap1CubeN — fixed cups, disjoint split)
A single red cube is hidden under one of three cups; the cups are shuffled a variable
number of times (0..3) and the robot must pick the cup now hiding the cube. Prompt is
color-free: "watch the video carefully, then pick up the container hiding the cube".
Derived from robomme_1cuben_allcases,
with two changes for a clean generalization study:
1. Fixed cup locations. Cup-pose perturbation is… See the full description on the dataset page: https://huggingface.co/datasets/alfayoung/robomme_1cuben_fixedcup_split.aurel_tensors_ch0_fixedrepro-fixed-budget-no-harder-than-fixed-confidence-bai-traces
Agent traces
Agent sessions published from a Trackio Logbook.
fixed-n-rb-offset256-qwen3-1.7b-base-math12k-754f8ca2-rollouts
fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset256_token_mean rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
fixed-kkc-dataset
Fixed KKC Dataset
日本語Wikipedia入力誤りデータセット (v2) から生成した、かな漢字変換(KKC)タスク用の選好ペアデータセットです。
データセットの概要
Wikipediaの編集差分のうち kanji-conversion_a カテゴリ(誤変換の修正)に該当するものを抽出しています。
各レコードは、カタカナの読みに対して「正しい漢字表記(chosen)」と「誤った表記(rejected)」のペアを持ちます。
かな漢字変換モデルの学習・評価や、選好学習(RLHF / DPO)に利用できます。
データ形式
各レコードは以下のフィールドを持つ JSON Lines 形式です。
フィールド
型
説明
left_context
string
変換箇所より前の文脈テキスト
prompt
string
変換対象語のカタカナ読み
chosen
string
正しい漢字表記(Wikipedia編集後)
rejected
string… See the full description on the dataset page: https://huggingface.co/datasets/yuuki14202028/fixed-kkc-dataset.Qwen3.7_5k_fr60_fixed
Qwen 3.7 Max Thinking — Distilled Reasoning Dataset (FR60, cleaned)
5,000 chain-of-thought (CoT) reasoning traces, ~60% machine-translated to French, derived from the original dataset WithinUsAI/Qwen3.7_Max_Thinking_dataset_5K.
Each example contains a problem, a detailed step-by-step reasoning trace (in the Qwen 3.7 Max Thinking style), and a concise final answer.
Source and translation
This dataset is a partial translation of the original English dataset… See the full description on the dataset page: https://huggingface.co/datasets/Tivaphraen/Qwen3.7_5k_fr60_fixed.RealMythosReasoning-unsloth-studio-fixedThe exactly same dataset as RealMythosReasoning(https://huggingface.co/datasets/RealMythos/RealMythosReasoning) but fixed for unsloth studio. Tested on cli and gui, it does work perfectly fine.
OPUS-MT-EN-Fixed
OPUS-100-Fixed: Tokenisation-Improved English-Maltese Dataset
Overview
OPUS-100-Fixed is an updated version of the OPUS-100 parallel English-Maltese dataset.
This version addresses tokenisation inconsistencies in the Maltese text using the MLRS tokeniser, aiming to improve machine translation quality.
The "en" column is the same as in the original OPUS-100 data, while the "mt" column has been corrected with the MLRS detokeniser.
Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/OPUS-MT-EN-Fixed.fixed-n-rb-offset512-qwen3-1.7b-base-math12k-8be8236f-rollouts
fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset512_token_mean rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
fixedimstupidfixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts
fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
fixeddatatinyperson-yolov8n-p2p3p4-oacp-fixedsplit42-0a2ca54-seed43fixedqwen3_4b_chimera_fixedtopics_questions_nofilterfixed-n-rb-offset1024-qwen3-1.7b-base-math12k-d8311fda-rollouts
fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset1024_token_mean_rerun rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
tinyperson-yolov8n-p2p3p4-oacp-fixedsplit42-0a2ca54-seed42HEC3R-ckpt-fixed_view_axis_freeze_dec
