datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
book2-lite-cleanedQwen3.5_RL_ErrorCase
Qwen3.5 RL:错误案例与视频定位诊断
v3部分共660题;另新增V4 RL Step3000选帧诊断100题。每题含QA、完整原始输出及可见帧拼图。Dataset Viewer中,default为前60题,video_grounding为v3新增600题,v4_rl_step3000为V4新增100题。
序号
内容
入口
001–060
原三个主实验bench案例
第001题
061–560
RL训练视频500题:训练视觉处理下的新输出
第061题
561–660
ASR-Bench视频100题:复用既有评测输出
第561题
V4-001–100
V4 RL Step3000:VSI/ASR选帧与bbox诊断
V4诊断首页
本次新增的测试内容
新增600题为在看结果前固定的诊断抽样,包含成功与失败,不是600个错误案例。未加入88题附加对照,避免重复。
两部分均为最终Qwen3.5-9B RL… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/Qwen3.5_RL_ErrorCase.sync_bigjob_8_finalised_processed_with_error_handling_from_51th_splitso100_insert_cylinder_errorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 5102,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sixpigs1/so100_insert_cylinder_error.satellite3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 166,
"total_frames": 36442,
"total_tasks": 2,
"total_videos": 498,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:166"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/satellite3.so100_pull_cube_by_tool_errorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 4381,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sixpigs1/so100_pull_cube_by_tool_error.so100_pick_cube_in_box_errorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 5373,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sixpigs1/so100_pick_cube_in_box_error.so100_stack_cube_errorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 6222,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sixpigs1/so100_stack_cube_error.jupyter-errors-dataset
Dataset Summary
The presented dataset contains 10000 Jupyter notebooks,
each of which contains at least one error. In addition to the notebook content,
the dataset also provides information about the repository where the notebook is stored.
This information can help restore the environment if needed.
Getting Started
This dataset is organized such that it can be naively loaded via the Hugging Face datasets library. We recommend using streaming due to the large size… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/jupyter-errors-dataset.satellite2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 116,
"total_frames": 24517,
"total_tasks": 1,
"total_videos": 348,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:116"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/satellite2.so100_pull_cube_errorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 3890,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sixpigs1/so100_pull_cube_error.MusicLMso100_push_cube_errorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 4284,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sixpigs1/so100_push_cube_error.satellite1.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 51,
"total_frames": 10533,
"total_tasks": 1,
"total_videos": 153,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/satellite1.1.combined3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 108,
"total_frames": 35004,
"total_tasks": 3,
"total_videos": 324,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:108"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/combined3.combined6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 155,
"total_frames": 50993,
"total_tasks": 10,
"total_videos": 465,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:155"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/combined6.blimp-single-errorDataset for probing model preferences for linguistically acceptable sentences.
Generated by introducing automatic corruptions into sentences from Wikipedia, based on UniMorph minimal tag pairs.
More info coming soon!
@misc{glocker2025growmergescalingstrategies,
title={Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation},
author={Kevin Glocker and Kätriin Kukk and Romina Oji and Marcel Bollmann and Marco Kuhlmann and Jenny Kunz},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/blimp-single-error.combined5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 135,
"total_frames": 44502,
"total_tasks": 6,
"total_videos": 405,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:135"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/combined5.combined4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 118,
"total_frames": 37854,
"total_tasks": 5,
"total_videos": 354,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:118"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/combined4.satellite5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 231,
"total_frames": 55004,
"total_tasks": 3,
"total_videos": 693,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:231"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/satellite5.postprocesscheck9This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 232,
"total_frames": 55209,
"total_tasks": 3,
"total_videos": 696,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:232"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/postprocesscheck9.onlynew_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 50,
"total_frames": 14302,
"total_tasks": 6,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/onlynew_2.icelandic-blimp-single-erroropen-smtp-error-dataset
Open SMTP Error Dataset
Open SMTP Error Dataset is an English-language, machine-readable reference
package for SMTP enhanced-status knowledge and cautious operational
classification. Version 1.1.0 contains 92 stable knowledge records and a
separate, auditable catalog of 120 classification rules.
The package is a reference artifact, not a live provider-policy feed. It helps
with observability, support, parser testing, and bounded delivery operations;
it does not establish… See the full description on the dataset page: https://huggingface.co/datasets/blazalek/open-smtp-error-dataset.satellite1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 65,
"total_frames": 13984,
"total_tasks": 1,
"total_videos": 195,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:65"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/satellite1.postprocesscheck12This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 232,
"total_frames": 55209,
"total_tasks": 3,
"total_videos": 696,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:232"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/postprocesscheck12.tabular-errors-v1
TabFix multilingual table error pairs — version 2.0
This release keeps 18 business error categories and separates executable deterministic detection from two residual neural categories: text.encoding and text.spelling. The same repository and family-disjoint splits are retained.
Split
Records
Open-vocabulary views
train
27948
3260
validation
17127
1844
test
32776
3540
The seven string columns remain id, split, family_id, clean_xml, corrupt_xml, errors… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/tabular-errors-v1.german-blimp-single-errorcombined1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 32,
"total_frames": 10559,
"total_tasks": 2,
"total_videos": 96,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/error-magnet/combined1.ErrorAnalysis
AnchorSR Error Analysis
500道题,严格沿 failure_cases_500.json 的文件顺序排列,每题对照三个SFT模型。
推荐从第001题开始,点击“下一题”逐题阅读。
也可使用本页上方 Dataset Viewer,一行就是一题:图片、问题、标准答案、三个模型完整输出。
Q-Spatial 150题,SpatialRGPT 175题,VSI 175题(仅尺寸、距离,不含面积)。
全部500题:每个模型均提供原图和标注图,共3000张图;O编号和帧号来自模型声明。
原图与标注图使用相同源帧和拼图顺序。无效框/帧号或无声明会注明,不补造;此时标注页可能没有框。
原图指未添加模型框的网页展示副本,经过等比例缩放与JPEG编码,并非原始文件字节;SpatialRGPT原有区域标记保留。
视频只展示可绘制对象涉及帧,无有效框时展示第1帧,非完整视频。原图/标注图使用相同帧。
原生输出完整保留,包括循环、截断和格式错误;未修改答案或重新评分。
所有模型均为SFT,不是baseline。至少一个模型在该题失败,其他模型可能答对。… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/ErrorAnalysis.
