datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
onthelook-fashion-anchor-positive-imagesRiskCase
AnchorSR RiskCase
This repository contains versioned diagnostic cohorts for reviewing automatic
quality-risk flags in AnchorSR datasets. A risk flag is a conservative review
trigger, not a confirmed error label.
Current cohort: Dataset-V4 review v1
The V4 cohort includes its selection methodology, structured risk evidence,
questions, answers/reasoning, original visual inputs, and a blank human-review
template. It is isolated from the official training-data repository and must… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/RiskCase.Anchor_imagesASR-Bench-1k
ASR-Bench-1k
Browse all 1,000 questions with visual previews
Select the preview subset in the Dataset Viewer for image previews and 16-frame
video contact sheets. Single-scene questions have one input; cross-scene questions
show A and B separately. Questions and sample IDs are unchanged. Previews omit answers.
The original public subset remains the default to preserve existing programmatic
loading behavior. Preview images are resized browsing aids, not evaluation media.
Video… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/ASR-Bench-1k.sparse-metric-anchors-ycb
Sparse Metric Anchors — YCB benchmark
The 42-object benchmark behind the paper Sparse Metric Anchors for a Single-View 3D
Generative Prior: The Output Frame Is the Bottleneck (ISIR, Sorbonne Université,
2026). Code and paper: github.com/635jack/sparse-metric-anchors — its colab/reproduce.ipynb recomputes every table of the paper from this dataset on a CPU runtime.
The paper asks what limits the injection of a few metric measurements — tactile
contacts, one depth map — into a… See the full description on the dataset page: https://huggingface.co/datasets/jack635/sparse-metric-anchors-ycb.kream-fashion-anchor-positive-imagesQwen3.5_RL_ErrorCase
Qwen3.5 RL:错误案例与视频定位诊断
v3部分共660题;另新增V4 RL Step3000选帧诊断100题。每题含QA、完整原始输出及可见帧拼图。Dataset Viewer中,default为前60题,video_grounding为v3新增600题,v4_rl_step3000为V4新增100题。
序号
内容
入口
001–060
原三个主实验bench案例
第001题
061–560
RL训练视频500题:训练视觉处理下的新输出
第061题
561–660
ASR-Bench视频100题:复用既有评测输出
第561题
V4-001–100
V4 RL Step3000:VSI/ASR选帧与bbox诊断
V4诊断首页
本次新增的测试内容
新增600题为在看结果前固定的诊断抽样,包含成功与失败,不是600个错误案例。未加入88题附加对照,避免重复。
两部分均为最终Qwen3.5-9B RL… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/Qwen3.5_RL_ErrorCase.ErrorAnalysis
AnchorSR Error Analysis
500道题,严格沿 failure_cases_500.json 的文件顺序排列,每题对照三个SFT模型。
推荐从第001题开始,点击“下一题”逐题阅读。
也可使用本页上方 Dataset Viewer,一行就是一题:图片、问题、标准答案、三个模型完整输出。
Q-Spatial 150题,SpatialRGPT 175题,VSI 175题(仅尺寸、距离,不含面积)。
全部500题:每个模型均提供原图和标注图,共3000张图;O编号和帧号来自模型声明。
原图与标注图使用相同源帧和拼图顺序。无效框/帧号或无声明会注明,不补造;此时标注页可能没有框。
原图指未添加模型框的网页展示副本,经过等比例缩放与JPEG编码,并非原始文件字节;SpatialRGPT原有区域标记保留。
视频只展示可绘制对象涉及帧,无有效框时展示第1帧,非完整视频。原图/标注图使用相同帧。
原生输出完整保留,包括循环、截断和格式错误;未修改答案或重新评分。
所有模型均为SFT,不是baseline。至少一个模型在该题失败,其他模型可能答对。… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/ErrorAnalysis.AnchorAnchorCrafter-finutune
Dataset Structure
The video_cut directory contains the original videos. Each file name follows the format:<person_id>_<object_id>.mp4, where the first number represents the person ID, and the second number represents the object ID.
For example:
video_cut/1_0.mp4 corresponds to the person video people_cut/1.mp4.
The associated object images are masked_object_cut/0_0.jpg, masked_object_cut/0_1.jpg, and masked_object_cut/0_2.jpg.
Additionally, the dataset includes annotations:
Human… See the full description on the dataset page: https://huggingface.co/datasets/cangcz/AnchorCrafter-finutune.AnchorCrafter-test
Dataset Overview
This dataset includes 5 objects, with each object associated with 2 videos. Additionally, there are 8 individuals, and each person is combined with each video, resulting in a total of 80 test videos.
Directory Structure
video_cut: Contains the original videos.
people_cut: Stores the 8 individual images.
masked_object_cut: Includes images of the 5 objects.
hand_cut: Features hand mesh extracted using Hamer.
depth_cut: Contains object depth videos… See the full description on the dataset page: https://huggingface.co/datasets/cangcz/AnchorCrafter-test.momo5-anchorsbpo-anchorsquery_image_anchor_positiveTrainingData_Stage3_VideoCases
Stage3 视频/多图人工查看样例
这是 18条人工查看样例,不是新训练集或评测集。直接向下浏览 QA 和拼图,
也可在 Dataset Viewer 中查看 frames 图片列。训练 QA 与原数据逐字保留;答案是数据集标注,不是人工确认的视觉真值。
来源
已确认视频类
待判断多图
CA-VQA
3(有序帧)
0
SpaceVista
3(有序帧)
3
VSI-590K
3(视频文件)
0
SIMS-VSI
3(视频文件)
0
ViCA-322K
3(视频文件)
0
选样:从正式 Small/train 中按固定 ID 顺序选择,优先不同任务、不同来源场景;不是随机代表性统计,也没有按答案是否正确挑样。
每张拼图从左到右、从上到下阅读。视频文件全程均匀抽取最多16帧;帧序列保留全部输入帧。
展示缩放不改变正式训练数据;拼图不是模型训练输入。多图未知项不计为视频。
来源:AnchorSR/TrainingData_Stage3,
固定提交… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3_VideoCases.aires-video-anchorsquery_image_anchor_positive_largeEB-Man_environment_anchored_prior_datasetquery_image_anchor_positive_large_384anchor_positive_image_dataset
