datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
exp_rpt_multifile-a0-base-pass8-300
TaskTrove A0 base pass@8 cohort
This repository contains the 300 extracted Harbor tasks selected for the A0
base-model pass@8 evaluation. The source is DCAgent/exp_rpt_multifile at
revision 70527e80ee7497c800ea0f9bad90b87423c784c9.
cohort.json records the deterministic random selection algorithm, seed, and
complete task list. The task directories are the unmodified archives from the
source dataset.
swe-bench-multi-file-refactoring-sft-dpo-2026
💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents.
📊 Dataset Architecture & Highlights
Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.rl__24GPU_base__exp_rpt_multifile__Qwen3-8Bfunc_rename_multi_file_swegymsmall_repos_multi_file_chatgpt_5_qas_part5_code_qa-datasetMultiFileTestswebench_verified_random_100_folders_a1_multifile_composition_20260627_124526multifile_fileqa_part2-datasetexp_rpt_multifile-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_multifile-qwen3.5-122b-131k-opencode-traces.glm52-datagen-r11-65-multifile-tracesexp_rpt_multifile_10kterminal_bench_2_a1_multifile_composition_20260326_040008multifile_fileqa-datasetterminal_bench_2_rl__24GPU_base__exp_rpt_multifile__Qwen3_8B_60_20260320_055700small_repos_multi_file_chatgpt_5_qas_code_qa_1k-datasetprompt2_multifile_qa_part1_2_3_4-datasetterminal_bench_2_a1_multifile_composition_20260324_100610ehr-multi-file-patient-samples
Built by DataFramer
DataFramer is an AI Workflow Intelligence Platform. Unify your AI traces,
user signals, and expert judgment into one loop — find accuracy failures,
diagnose root causes, and turn every fix into reusable business context that
makes every AI workflow more accurate, adopted, and valuable.
▶︎ Start free · Docs
README
This repository contains a 1000-strong synthetic EHR / patient histories dataset generated by DataFramer (AIMon Labs, Inc.) for… See the full description on the dataset page: https://huggingface.co/datasets/dataframer/ehr-multi-file-patient-samples.a1_multifile_compositiona1_multifile_composition-3pctsmall_repos_multi_file_chatgpt_5_qas_part2_3_code_qa-datasetdev_set_v2_a1_multifile_composition_20260324_201735small_repos_multi_file_chatgpt_5_qas_part4_code_qa-datasetswebench_verified_random_100_folders_a1_multifile_composition_20260324_072855a1_multifile_composition-10pctpylint_logic_multifile_codebase_100terminal_bench_2_syh_rl_multifile_40_32B_20260501_232008pylint_logic_multifile_dedup_cleaned_30pylint_logic_multifile_codebase_dedup_99dev_set_v2_rl__24GPU_base__exp_rpt_multifile__Qwen3_8B_60_20260319_053730
