datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-marathon
SWE Marathon: Ultra Long-Horizon Software Engineering Tasks
20 ultra long-horizon software-engineering tasks designed to challenge frontier coding agents. Each task ships with a containerized environment, a precise instruction, comprehensive tests, and a reference oracle solution. All tasks pass NOP-baseline / Oracle-fix validation.
Homepage: https://github.com/abundant-ai/swe-marathon
License: Apache 2.0
Format: Harbor task format (task.toml + instruction.md + environment/ +… See the full description on the dataset page: https://huggingface.co/datasets/rdesai2/swe-marathon.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.self-self-distillation
self-self-distillation
Per-question teacher/student reward-delta annotations for verifier-free self-self-distillation,
computed on the sky_work_math subset of
PrimeIntellect/SYNTHETIC-2-RL with
Qwen/Qwen3-4B.
For each problem we draw k=8 rollouts in thinking-on (teacher) and thinking-off (student) modes at
identical sampling (temperature 0.7 / top_p 0.8), grade each against the ground truth, and record the
per-mode expected reward and their difference (delta = R_teacher -… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/self-self-distillation.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.JayHyeon__Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1-details
Dataset Card for Evaluation run of JayHyeon/Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1
Dataset automatically created during the evaluation run of model JayHyeon/Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JayHyeon__Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1-details.cpp-unittest-26-11-2025RdiffusionRdiffusion
We're releasing the entire corpus of publicly available songs from the Riffusion platform—generated and shared by their user community. Through extensive scraping of their exposed API, we’ve collected over 2,000 artificial songs, including every metadata field and downloadable asset that was accessible at the time.
📦 Included Data
Audio & Visual Assets:audio_url, audio_b64, image_url, image_b64, video_url
Metadata & Structure:id, title, author_id, created_at… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/Rdiffusion.princeton-nlp__Llama-3-Instruct-8B-RDPO-details
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Instruct-8B-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Instruct-8B-RDPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Llama-3-Instruct-8B-RDPO-details.teacher_hf_public
C++ Concolic Test Driver Generation Dataset
Dataset chứa 300 mẫu dùng để sinh/đánh giá test driver C++ bằng phương pháp concolic testing. Mỗi mẫu gồm mã nguồn hàm đích, prompt yêu cầu sinh test driver, và test driver đầu ra kèm kết quả coverage thực tế.
Thông tin tổng quan
Số mẫu: 300
Ngôn ngữ mã nguồn: C++
Repo nguồn: abseil (104 mẫu), bullet3 (196 mẫu)
Định dạng: JSONL (mỗi dòng là một JSON object)
Các trường dữ liệu
Trường
Kiểu
Mô tả… See the full description on the dataset page: https://huggingface.co/datasets/rd320uetvnu/teacher_hf_public.cpp-unittest-24-11-2025princeton-nlp__Llama-3-Instruct-8B-RDPO-v0.2-details
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Instruct-8B-RDPO-v0.2
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Instruct-8B-RDPO-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Llama-3-Instruct-8B-RDPO-v0.2-details.how2sign_rajviprinceton-nlp__Llama-3-Base-8B-SFT-RDPO-details
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Llama-3-Base-8B-SFT-RDPO-details.JayHyeon__Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.3-details
Dataset Card for Evaluation run of JayHyeon/Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.3
Dataset automatically created during the evaluation run of model JayHyeon/Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JayHyeon__Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.3-details.JayHyeon__Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.1-details
Dataset Card for Evaluation run of JayHyeon/Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.1
Dataset automatically created during the evaluation run of model JayHyeon/Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JayHyeon__Qwen_0.5-rDPO_5e-7-3ep_0vpo_const_0.1-details.princeton-nlp__Mistral-7B-Instruct-RDPO-details
Dataset Card for Evaluation run of princeton-nlp/Mistral-7B-Instruct-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Mistral-7B-Instruct-RDPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Mistral-7B-Instruct-RDPO-details.princeton-nlp__Mistral-7B-Base-SFT-RDPO-details
Dataset Card for Evaluation run of princeton-nlp/Mistral-7B-Base-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Mistral-7B-Base-SFT-RDPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Mistral-7B-Base-SFT-RDPO-details.RDson__WomboCombo-R1-Coder-14B-Preview-details
Dataset Card for Evaluation run of RDson/WomboCombo-R1-Coder-14B-Preview
Dataset automatically created during the evaluation run of model RDson/WomboCombo-R1-Coder-14B-Preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RDson__WomboCombo-R1-Coder-14B-Preview-details.cpp-unittest-11-5-2025cpp-unittest-10-12-2025cpp-unittest-updatecpp-unittest-28-11-2025dexterous-rdp-cube-data
AutoDL Research Backup
Private migration backup from /root/autodl-tmp. Selection and restore metadata are in SebastianLZJ/autodl-server-archive.
