datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OneThinker-train-data
OneThinker-600k Training Data
This repository contains the training data for OneThinker, an all-in-one reasoning model for image and video, as presented in the paper OneThinker: All-in-one Reasoning Model for Image and Video.
Code: https://github.com/tulerfeng/OneThinker
About the OneThinker Dataset
OneThinker-600k is a large-scale multi-task training corpus designed to train OneThinker, an all-in-one multimodal reasoning model capable of understanding… See the full description on the dataset page: https://huggingface.co/datasets/OneThink/OneThinker-train-data.OneThinker-train-data
OneThinker-600k Training Data
This repository contains the training data for OneThinker, an all-in-one reasoning model for image and video, as presented in the paper OneThinker: All-in-one Reasoning Model for Image and Video.
Code: https://github.com/tulerfeng/OneThinker
About the OneThinker Dataset
OneThinker-600k is a large-scale multi-task training corpus designed to train OneThinker, an all-in-one multimodal reasoning model capable of understanding… See the full description on the dataset page: https://huggingface.co/datasets/luckywin90/OneThinker-train-data.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.One-to-All-sub
One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
This repository contains the sample training data and benchmarks associated with the paper One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer.
The paper presents a unified framework for high-fidelity character animation and image pose transfer for references with arbitrary layouts, addressing spatial misalignment and partially visible references through innovative… See the full description on the dataset page: https://huggingface.co/datasets/MochunniaN1/One-to-All-sub.onetwovla-dataset
Datasets for OneTwoVLA
[Project Page] | [Paper] | [Code]
This repository provides datasets collected with the UMI, converted into the LeRobot data format, along with synthetic vision-language data used in the paper OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning.
The robot data covers two main tasks:
Cocktail
Open-World Visual Grounding
Dataset Folders
cocktailContains 299 real-world demonstrations collected in the lab, each with reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Nai/onetwovla-dataset.OneThinker-evalThis repository contains the evaluation data presented in: OneThinker: All-in-one Reasoning Model for Image and Video
Code: https://github.com/tulerfeng/OneThinker
About OneThinker
We introduce OneThinker, an all-in-one multimodal reasoning generalist that is capable of thinking across a wide range of fundamental visual tasks within a single model.
We construct the large-scale OneThinker-600k multi-task training corpus and build OneThinker-SFT-340k with high-quality CoT… See the full description on the dataset page: https://huggingface.co/datasets/OneThink/OneThinker-eval.so101-onetape-cleanupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 22218,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuk6ra/so101-onetape-cleanup.onet-m6-social
Summary
This is a question-answer dataset for the Grade 12 (M6) Social subject of the Thailand Ordinary National Educational Test (ONET).
The dataset was human-extracted by my team from the official release of publicly available exams National Institute of Educational Testing Service during the years 2016-2022.
The exam consists of 510 multiple-choice questions with corresponding answer keys.
It is important to note that only two questions, Q71 and Q85, from the year 2018, require… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/onet-m6-social.onetwovla
Datasets for OneTwoVLA
[Project Page] | [Paper] | [Code]
This repository provides datasets collected with the UMI, converted into the LeRobot data format, along with synthetic vision-language data used in the paper OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning.
The robot data covers two main tasks:
Cocktail
Open-World Visual Grounding
Dataset Folders
cocktailContains 299 real-world demonstrations collected in the lab, each with reasoning… See the full description on the dataset page: https://huggingface.co/datasets/yilin-wu/onetwovla.one-traj-demos-recollect-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 80,
"total_frames": 15979,
"total_tasks": 67,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/one-traj-demos-recollect-1.TOOLSLPFF
LPFF: Large-Pose-Flickr-Faces Dataset
LPFF is a large-pose Flickr face dataset comprised of 19,590 high-quality real large-pose portrait images.
[ICCV 2023] LPFF: A Portrait Dataset for Face Generators Across Large Poses
Yiqian Wu, Jing Zhang, Hongbo Fu, Xiaogang Jin*
Paper Video Suppl Project Page
The creation of 2D realistic facial images and 3D face shapes using generative networks has been a hot topic in recent years. Existing face… See the full description on the dataset page: https://huggingface.co/datasets/onethousand/LPFF.lumos_complex_qa_plan_onetime
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_complex_qa_plan_onetime.thaiexam-onetOneThinker_train_data_sftonetube2_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 2247,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/arminfg/onetube2_augmented.one-traj-demos-uploadThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 73,
"total_frames": 14862,
"total_tasks": 61,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:73"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/one-traj-demos-upload.one-traj-demos-prioritizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 90,
"total_frames": 18054,
"total_tasks": 74,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/one-traj-demos-prioritized.360degree-PHQ
[Preprint] 3DPortraitGAN: Learning One-Quarter Headshot 3D GANs from a Single-View Portrait Dataset with Diverse Body Poses
Yiqian Wu, Hao Xu, Xiangjun Tang, Hongbo Fu, Xiaogang Jin*
Paper (Arxiv) Supplementary (Google Drive)
This is the training dataset, 360°PHQ dataset, of 3DPortraitGAN.
adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.thai-onet-fixed-trainJobBERT-ONET-evaluation-datasetGenLCA-resourcesmcp-onet-task-classification-public
MCP → O*NET Task Automation Classification
Classification of ~10,140 Model Context Protocol (MCP)
servers against the O*NET occupational task
framework, measuring how much each MCP server could automate real-world
occupational tasks. Built with Voyage-4-large embeddings + GPT-4.1.
Created for the paper "Mapping AI Exposure Across the U.S. Workforce: Evidence
from Millions of AI Conversations" (Wright, Schwarze, & Boyd, 2026).
📦 Code, pipeline, and full documentation… See the full description on the dataset page: https://huggingface.co/datasets/theodorewright11/mcp-onet-task-classification-public.lumos_maths_ground_onetime
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_maths_ground_onetime.one_twenty_five_faces
label_names = {
0: "Adriana Lima",
1: "Akshay Kumar",
2: "Alex Lawther",
3: "Alexandra Daddario",
4: "Alia Bhatt",
5: "Allen Page",
6: "Alvaro Morte",
7: "Alycia Debnam-Carey",
8: "Amanda Crew",
9: "Amber Heard",
10: "Amitabh Bachchan",
11: "Andy Samberg",
12: "Anne Hathaway",
13: "Anthony Mackie",
14: "Anushka Sharma",
15: "Avril Lavigne",
16: "Barack Obama",
17: "Barbara Palvin",
18: "Ben Affleck",
19: "Bill… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/one_twenty_five_faces.math_onetonetube2_duplicatedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 2247,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/arminfg/onetube2_duplicated.isco-08_and_onet_occupation_crosswalk_dataset
