datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TraingTangTian
✨ Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training ✨
Fangfu Liu*,1,
Diankun Wu*,1,
Jiawei Chi*,1,
Yimo Cai1,
Yi-Hsin Hung1,
Xumin Yu2,
Hao Li3,
Han Hu2,
Yongming Rao†,2,
Yueqi Duan†,1
*Equal Contribution †Corresponding Author
1Tsinghua University 2Tencent Hunyuan 3NTU
Spatial-TTT: We propose… See the full description on the dataset page: https://huggingface.co/datasets/Wuuu3511/TraingTangTian.MMRC_Real_World_Conversation
MMRC - Multi-Modal Open-Ended Conversation Dataset
Overview:
MMRC is a benchmark dataset designed for evaluating Multi-Modal Large Language Models (MLLMs) in open-ended, multi-turn conversations. It provides diverse, real-world conversational data that integrates both textual and visual modalities, aiming to push the boundaries of MLLM performance in practical settings.
Dataset Details:
The MMRC dataset is composed of multi-turn conversations with integrated… See the full description on the dataset page: https://huggingface.co/datasets/WUUE/MMRC_Real_World_Conversation.bmvs_sparse_dtu2my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/wuuski/my-personal-codex-data.VeriOS-Bench
VeriOS-Bench
Paper | Code
Dataset Overview
This dataset is a Query-Driven Trustworthy OS Agent Dataset implemented as described in the paper:
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
The VeriOS dataset is designed to train and evaluate operating system (OS) agents that automate tasks through on-device graphical user interfaces (GUIs). It focuses on enabling agents to decide when to query humans for more reliable task completion… See the full description on the dataset page: https://huggingface.co/datasets/wuuuuuz/VeriOS-Bench.GUI-CIDERSaliTrapESPPtasks_list_think
基本說明
資料集內容:思考能力的 tasks list, 以及評測思考能力的題目
資料集版本: 20250225
本資料集運用 Gemini 2.0 Flash Thinking Experimental 01-21 生成
思考要如何建構推理模型,於是先跟 LLM 討論何謂推理,有哪些樣態,之後根據可能的類別請 LLM 生成 tasks list , 結果就是檔案 taskslist.csv
根據每一個 task 描述,再去生成不同難度的題目,就是 validation.csv
taskslist.csv
目前分六大類,共 85 子類
ID 編碼說明
都是 TK 開頭
L2.2 就是 TK22
目前後面兩碼是序號
group 說明
代號
說明
L1.1
基礎推理類型
L2.1
KR&R 推理
L3.1
深度
L4.1
領域
L5.1
推理呈現能力
L6.1
元推理
validation.csv
id:對應的 task id, 定義在… See the full description on the dataset page: https://huggingface.co/datasets/wuulong/tasks_list_think.purchasing_exam_questions
資料來源:採購法規題庫
資料產生日期:114/03/07
大部分項目內本沒有 「依據法源」欄位,為求統一所以有欄位,為空值
順手用 colab 觀察資料內容:採購網題庫1.ipynb
MobileIAR
Dataset Overview
This dataset is a User-Specific Personalized GUI Mobile Agents Dataset implemented as described in the paper:
Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
Citation
@article{wu2025quick,
title={Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents},
author={Wu, Zheng and Huang, Heyuan and Yang, Yanjia and Song, Yuanyi and Lou, Xingyu… See the full description on the dataset page: https://huggingface.co/datasets/wuuuuuz/MobileIAR.GomvshighresdtutestOS-SPEARevaluationGomvs2WU_UchVMy9YxI0sn64
