datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HiFi-UMI-2K
HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data
2,000 hours released · 6 synchronized camera views · 480+ scenes · 3 mm pose accuracy · <40 µs synchronization
🌐 Project Website |
📦 Dataset |
📄 Paper: arXiv:2607.25895
Examples from the HiFi-UMI corpus. Click the image to play the video.
📚 Introduction
HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.… See the full description on the dataset page: https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K.appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.40546875
Action score: 0.475
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4703125
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.39921875
Action score: 0.44375
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38359375
Action score: 0.4703125
Valid samples: 320/320
Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.radeonvla_reflex_physical_2k
RadeonVLA-Reflex Physical-2K
Physical-2K contains 2,000 strictly validated successful Genesis episodes for
language-conditioned Franka fruit sorting. Coverage is exactly five fruits ×
four bowl positions × 100 episodes = 2,000 episodes.
Verified release facts
Item
Value
Episodes
2,000
Frames
468,889
Registered task variations
20
Episodes per task variation
100
Control frequency
20 Hz
Strict success certificates
2,000
Sampled image frames
1… See the full description on the dataset page: https://huggingface.co/datasets/a3124371940/radeonvla_reflex_physical_2k.Muse-Glimmer-SWE-Gym-2k
Muse-Glimmer-SWE-Gym-2k
Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a
speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and
SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them.
Configs
Config
Rows
Size
What it is
train
1,981
57 MB
One row per trajectory: the full conversation as messages.
raw
159,999
2.7 GB
One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.fineweb2-2k-samplesenvironmental_2kgovernance_2ksocial_2ksci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
mlfoundations-dev/sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 26.94% ± 4.54%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
19.70%
39
198
2
23.23%
46
198
3
37.88%
75
198
v3-2k-traj-claude-opus-4.7github-patches-genesys-swe-prompt-2k-context-1k-diffeval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k20_testingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 1336,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bonnieliu2002/eval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k20_testing.secops-2k
SecOps-2k
2,000 synthetic security log lines with template labels in LogHub
format. Parser papers mostly test on system logs (HDFS, BGL, Apache).
No comparable set existed for security telemetry, so we built one:
sshd, sudo, firewall, and audit lines, all invented. No real hosts,
users, or IPs.
Authors: TMFNK and MbitAI. Generator code:
TMFNK/LogParser-Dataset.
Archived release:
doi:10.5281/zenodo.22341506.
What is inside
One host (secops-01), one day (14 Jun), 2… See the full description on the dataset page: https://huggingface.co/datasets/mbitai/secops-2k.github-patches-genesys-2k-context-1k-diffgithub-patches-genesys-agentless-prompt-2k-context-1k-diffOpenMathReasoning-subset30kfiltered-Qwen3-1.7B-2k-concise-with-answergithub-patches-genesys-agentless-prompt-2k-context-1k-diff-2k-examplesgsm_infinite_symbolic_2kc4-en-2k-tos-game-replay
Fixed English C4 replay subset
A subset of allenai/c4, English
configuration, training split. C4 is derived from Common Crawl; see the upstream
card for provenance and licensing. Sized against cfierro/tos_game_synthetic_docs, split
train, using raw text tokens without special tokens or truncation.
All 6,219 documents are in train, with 2,982,687 raw tokens.
Whole documents are kept until the target is reached; exact duplicate texts
are skipped. id is SHA-256 of the original… See the full description on the dataset page: https://huggingface.co/datasets/cfierro/c4-en-2k-tos-game-replay.statML-arxiv-RL-2k-docsTrain-only prefix of JackHsieh/statML-arxiv-RL-4k-docs's train split, drawn from JackHsieh/statML-arxiv.
Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens
(Qwen/Qwen3-4B-Instruct-2507) from a distinct paper. start_index is the window's offset in the source paper's
token sequence; input_ids is the Qwen3 encoding of text (no special tokens added — no
BOS/EOS). Same schema and recipe as
JackHsieh/statML-arxiv-40M-20M.
Nesting:
these are the first 2_048 rows of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-RL-2k-docs.OpenMathReasoning-subset30kfiltered-Qwen3-1.7B-2k-v5eval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k50_tec05_testing1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 1140,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bonnieliu2002/eval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k50_tec05_testing1.OpenMathReasoning-subset30kfiltered-Qwen3-1.7B-2k-conciseOpenMathReasoning-subset30k-Qwen3-1.7B-2k-concise-with-answerOpenMathReasoning-subset30k-Qwen3-1.7B-2k-conciseswe_smith_py_trajs_2k_incomplete
