datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InternData-A1
InternData-A1
InternData-A1 is a hybrid synthetic-real manipulation dataset containing over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation.
Your browser does not support the video tag.
Your browser does not support the video tag.… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/InternData-A1.a11oy-verifiable-corpus
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
a11oy — Verifiable Corpus · verify it yourself
This dataset publishes a11oy's signed receipts and proof surface so that
anyone can independently verify them — no trust in SZL Holdings required.
Every receipt here carries the full cryptographic material needed to check its
signature offline;… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/a11oy-verifiable-corpus.Inter-Edit-Train
Inter-Edit-Train
Inter-Edit-Train is the official large-scale training set released for the CVPR 2026 paper Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing.
This dataset is designed for the Interactive Instruction-based Image Editing (I^3E) task, where a model performs localized image edits from a concise textual instruction together with imprecise spatial guidance.
Highlights
1,099,964 image editing pairs
610,186 unique source images
Four… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Train.navsim-metric-caches-from-a100ViZDoom-Deathmatch-PPO-XLrg
ViZDoom Deathmatch with pretrained PPO agent playing over 15 episodes.
terminal_bench_2_a1_issue_tasks_20260805_125430a1_issue_tasksterminal_bench_2_a1_issue_tasks_20260627_124529TCM_Book_Corpus
TCM Book Corpus
Pre-training Corpus
Data Overview
Modality
Description
Data Quantity
TCM_Book_Corpus
📝 Text
A cleaned corpus of 3,256 TCM textbooks.
~ 0.5 B tokens
Important Notice
版权归属:原始数据由 FreedomIntelligence (FreedomAI) 公开提供,本仓库仅做整理与格式转换。数据来源:截取自 FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT。
若使用本数据,请同时引用原作者工作:
@misc{chen2025shizhengptmultimodalllmstraditional,
title={ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese… See the full description on the dataset page: https://huggingface.co/datasets/A1C/TCM_Book_Corpus.p-data-a100-2terminal_bench_2_a1_issue_tasks_20260325_00080410M_PBMCViZDoom-Deathmatch-Random-XLrgterminal_bench_2_a1_issue_tasks_20260711_172919Inter-Edit-Test
Inter-Edit-Test
Official test benchmark release for the CVPR 2026 paper:
Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing
This repository hosts the public release of Inter-Edit-Test, a human-annotated benchmark for the Interactive Instruction-based Image Editing (I^3E) task.
Each sample contains:
a source image,
a coarse user-style interaction mask,
a concise editing instruction,
and a ground-truth edited image.
To simplify large-scale distribution on… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Test.a1_math_numina_aimephase1_pick_place_A1_10fps
Phase1 Pick Place A1 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Pick up the red block and place it on the blue dish."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 100 (28,459 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac Sim 5.1 /… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/phase1_pick_place_A1_10fps.dev_set_v2_a1_issue_tasks_20260324_201730g1_min_episodes_a1_top16_glm47_traceseval-DCAgent_a1-bugsinpy_DCAgent2_terminal_bench_2terminal_bench_2_a1_toolscale_20260821_153653swebench_verified_random_100_folders_a1_stack_rust_20260818_162727NVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresqfs-capture-a1ec5246f3ef
HF workflow a1ec5246f3ef11d323425ae8a9ea47ef
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from Qwen/Qwen3.8-27B.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-capture-a1ec5246f3ef.terminal_bench_2_a1_toolscale_20260810_081852CEFR_Mixed_Dataset_A1_A2_sisa
CEFR Dataset for A1 and A2
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A1 (2000 sentences) and A2 (100 sentences). Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts A1, A2, B1). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A1_A2_sisa.a1_math_openmathinstruct_aimegems-a10-sentencesterminal_bench_2_a1_taskmaster2_20260810_061724dev_set_v2_a1_crosscodeeval_typescript_20260805_125428
