CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anthracite-org /nopm_claude_writing_fixedThis is Nopm/Opus_WritingStruct, reuploaded and properly converted to ShareGPT format. text1K<n<10K19 likes909 downloads2y agoHugging Face02SetFit /wsc_fixed Glue WSC Fixed This dataset is a port of the official wsc.fixed dataset on the Hub. Also, the test split is not labeled; the label column values are always -1. tabularn<1K1 likes311 downloads4y agoHugging Face03rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes297 downloads1y agoHugging Face04zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes239 downloads1mo agoHugging Face05AgentSuite /DrafterBench-fixed-trajectories AgentSuite/DrafterBench-fixed-trajectories Per-model agent trajectory data for DrafterBench-fixed (public release). Models: 30 Tasks per model: 1,920 One file per model: {model}.jsonl, one JSON object per line. Fields: model_path, user_model_path, benchmark_name, task_name, sampling_params, user_sampling_params, messages, eval_result, meta. sampling_params reflect each benchmark's own implementation; values the benchmark leaves unset are recorded as null (provider default).… See the full description on the dataset page: https://huggingface.co/datasets/AgentSuite/DrafterBench-fixed-trajectories.text10K<n<100K0 likes195 downloads4mo agoHugging Face06souvik18 /mistral_tokenized_2048_fixed_shards1M<n<10M0 likes192 downloads10mo agoHugging Face07brysgo /gol-rl-fixed-validation-37156495 GoL World Model — Genie at Tiny Scale Current implementation status is tracked in STATUS_2026-05-19.md. The repo currently has three executable tracks: the recursive world-model demos/training path, agentic trajectory collection, and online GRPO training via train_grpo.py. A miniature implementation of the Genie world model architecture using Conway's Game of Life as the substrate. Goal: Show that resource-constrained researchers can experiment with world model ideas using a… See the full description on the dataset page: https://huggingface.co/datasets/brysgo/gol-rl-fixed-validation-37156495.tabularn<1K0 likes144 downloads4mo agoHugging Face08CodeFlame /FIXED-Cleaned-Claude-Sonnet-5-Grok-4.5-ChatGPT-5.6-Luna-Qwen-3.8-MAX Ultra-Clean HF Dataset 492 pairs. Zero residual JSON garbage. High-tier technical SFT data. textn<1K3 likes144 downloads1mo agoHugging Face09G-reen /instruct-set-longer-fixedtext100K<n<1M0 likes126 downloads9mo agoHugging Face10zjhhhh /fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes90 downloads19d agoHugging Face11zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes83 downloads18d agoHugging Face12healthyfat /aurel_tensors_EOS_fixedtextn<1K0 likes77 downloads4mo agoHugging Face13alfayoung /robomme_1cuben_fixedcup_split robomme_1cuben_fixedcup_split (VideoUnmaskSwap1CubeN — fixed cups, disjoint split) A single red cube is hidden under one of three cups; the cups are shuffled a variable number of times (0..3) and the robot must pick the cup now hiding the cube. Prompt is color-free: "watch the video carefully, then pick up the container hiding the cube". Derived from robomme_1cuben_allcases, with two changes for a clean generalization study: 1. Fixed cup locations. Cup-pose perturbation is… See the full description on the dataset page: https://huggingface.co/datasets/alfayoung/robomme_1cuben_fixedcup_split.tabularroboticsn<1K0 likes73 downloads2mo agoHugging Face14Eishaan /repro-fixed-budget-no-harder-than-fixed-confidence-bai-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes71 downloads2mo agoHugging Face15healthyfat /aurel_tensors_ch0_fixedtextn<1K0 likes69 downloads4mo agoHugging Face16zjhhhh /fixed-n-rb-offset256-qwen3-1.7b-base-math12k-754f8ca2-rollouts fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset256_token_mean rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes68 downloads12d agoHugging Face17yuuki14202028 /fixed-kkc-dataset Fixed KKC Dataset 日本語Wikipedia入力誤りデータセット (v2) から生成した、かな漢字変換(KKC)タスク用の選好ペアデータセットです。 データセットの概要 Wikipediaの編集差分のうち kanji-conversion_a カテゴリ(誤変換の修正)に該当するものを抽出しています。 各レコードは、カタカナの読みに対して「正しい漢字表記(chosen)」と「誤った表記(rejected)」のペアを持ちます。 かな漢字変換モデルの学習・評価や、選好学習(RLHF / DPO)に利用できます。 データ形式 各レコードは以下のフィールドを持つ JSON Lines 形式です。 フィールド 型 説明 left_context string 変換箇所より前の文脈テキスト prompt string 変換対象語のカタカナ読み chosen string 正しい漢字表記(Wikipedia編集後) rejected string… See the full description on the dataset page: https://huggingface.co/datasets/yuuki14202028/fixed-kkc-dataset.texttext-generation100K<n<1M1 likes64 downloads7mo agoHugging Face18Tivaphraen /Qwen3.7_5k_fr60_fixed Qwen 3.7 Max Thinking — Distilled Reasoning Dataset (FR60, cleaned) 5,000 chain-of-thought (CoT) reasoning traces, ~60% machine-translated to French, derived from the original dataset WithinUsAI/Qwen3.7_Max_Thinking_dataset_5K. Each example contains a problem, a detailed step-by-step reasoning trace (in the Qwen 3.7 Max Thinking style), and a concise final answer. Source and translation This dataset is a partial translation of the original English dataset… See the full description on the dataset page: https://huggingface.co/datasets/Tivaphraen/Qwen3.7_5k_fr60_fixed.texttext-generation1K<n<10K1 likes64 downloads2mo agoHugging Face19llmguy342 /RealMythosReasoning-unsloth-studio-fixedThe exactly same dataset as RealMythosReasoning(https://huggingface.co/datasets/RealMythos/RealMythosReasoning) but fixed for unsloth studio. Tested on cli and gui, it does work perfectly fine. text1K<n<10K0 likes61 downloads6d agoHugging Face20MLRS /OPUS-MT-EN-Fixed OPUS-100-Fixed: Tokenisation-Improved English-Maltese Dataset Overview OPUS-100-Fixed is an updated version of the OPUS-100 parallel English-Maltese dataset. This version addresses tokenisation inconsistencies in the Maltese text using the MLRS tokeniser, aiming to improve machine translation quality. The "en" column is the same as in the original OPUS-100 data, while the "mt" column has been corrected with the MLRS detokeniser. Citation If you use this… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/OPUS-MT-EN-Fixed.texttranslation1M<n<10M3 likes59 downloads2y agoHugging Face21zjhhhh /fixed-n-rb-offset512-qwen3-1.7b-base-math12k-8be8236f-rollouts fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset512_token_mean rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes54 downloads10d agoHugging Face22Skorcht /fixedimstupidtextn<1K0 likes52 downloads2y agoHugging Face23hi-todayis-jh /fixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes51 downloads5d agoHugging Face24Skorcht /fixeddatatext100K<n<1M0 likes45 downloads2y agoHugging Face25duyle2408 /tinyperson-yolov8n-p2p3p4-oacp-fixedsplit42-0a2ca54-seed43tabular100K<n<1M0 likes45 downloads17d agoHugging Face26Skorcht /fixedtextn<1K0 likes43 downloads2y agoHugging Face27fzzhang /qwen3_4b_chimera_fixedtopics_questions_nofiltertext10K<n<100K1 likes40 downloads4mo agoHugging Face28hi-todayis-jh /fixed-n-rb-offset1024-qwen3-1.7b-base-math12k-d8311fda-rollouts fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset1024_token_mean_rerun rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes40 downloads9d agoHugging Face29duyle2408 /tinyperson-yolov8n-p2p3p4-oacp-fixedsplit42-0a2ca54-seed42tabular100K<n<1M0 likes39 downloads17d agoHugging Face30PKUfudawei /PhyBench_fixedOriginal credit to PHYBench 22 answers are found not able to be successfully parsed to sympy symbolic expression and stored in wrong_fixed.json The benchmark with open-source questions and answers are fixed in PHYBench-fullques_v1.repaired.json. Now all of the answers in latex can be successfully converted to sympy with latex_pre_process.py textn<1K0 likes38 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.