datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.bankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.HANDALATLAS-Finance
ATLAS Finance
A benchmark of 100 expert-level tasks inside 13 realistic financial firm environments, packaged in the Harbor RLE format.
Each task drops an AI agent into a Linux workstation with a
persistent multi-app world — inbox, chat, calendar, virtual data room, drive,
wiki — and asks the agent to produce the same deliverable a financial professional would be responsible for:
an Excel workbook containing the model and supporting analysis.
Here we provide the data for this… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/ATLAS-Finance.HanDyVQA
HanDyVQA Dataset 👋
HanDyVQA (Hand-Object Dynamics Video Question Answering) Dataset is a new benchmark for evalutating abundant spatio-temporal dynamics, process, and effects contained in hand-object interactions. This dataset is built on top of Ego4D Dataset.
Get Started
0. Install LFS
If you haven’t already, install Git Large File Storage (LFS):
git lfs install
1. Clone Repository
git clone https://huggingface.co/datasets/aist-cvrt/HanDyVQA… See the full description on the dataset page: https://huggingface.co/datasets/aist-cvrt/HanDyVQA.handy-dictation-editing
Handy dictation-editing corpus
Turns a raw dictated transcript into the text the speaker meant to write.
in : um so the meeting is uh moved to friday no wait thursday at three
out: The meeting is Thursday at three.
Three jobs at once, because they are not separable in speech: drop filler words,
repair punctuation and capitalisation, and — the hard one — when the speaker
changes their mind mid-sentence, delete the wording they abandoned and keep only
what they settled on.
Built… See the full description on the dataset page: https://huggingface.co/datasets/MagicNoThief/handy-dictation-editing.Pashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.handvqa
HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
CVPR 2026
MD Khalequzzaman Chowdhury Sayem1*,
Mubarrat Tajoar Chowdhury1*,
Yihalem Yimolal Tiruneh1,
Muneeb A. Khan1,
Muhammad Salman Ali1,
Binod Bhattarai2,3,4†,
Seungryul Baek1†
1UNIST,
2University of Aberdeen,
3University College London,
4Fogsphere (Redev.AI Ltd)
*Equal contribution.
†These authors jointly supervised this work.… See the full description on the dataset page: https://huggingface.co/datasets/kcsayem/handvqa.LongVideo-Reason-4k-Video-Crop-Handoff-20260911
LongVideo-Reason 4k · Video Crop 合成移交包
公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。
本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。
内容
文件
用途
data/qa.jsonl
4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源
videos/*.mp4
配套原视频;与 QA 的 video_path 对应
data/video_manifest.jsonl
每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源
data/selection_report.json
最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.odia-handwritten-ocr
Odia Handwritten OCR Dataset
Dataset Description
This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Task: Optical Character Recognition (OCR)
Total Images: 182,152
Character Classes: 47
Image Format: Grayscale JPG (32x32 pixels)
Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.multi_agent_handoff
Multi-Agent Handoff Synthetic Dataset
The Multi-Agent Handoff Synthetic Dataset is a fully synthetic dataset designed to support research and development in multi-agent systems.
Specifically, it focuses on agent handoffs (https://openai.github.io/openai-agents-python/handoffs/) — scenarios where a central language model delegates specialized tasks to sub-agents based on user prompts.
The domian, sys_prompts and subagents design:… See the full description on the dataset page: https://huggingface.co/datasets/JayYz/multi_agent_handoff.unilink-visa-handbook 1|# UNILINK Visa Handbook Dataset
2|
3|> A neutral, citable corpus of visa & immigration facts across 8 jurisdictions (AU/UK/US/CA/NZ/JP/HK/MY), compiled and structured by **UNILINK Education** (licensed education & migration agent, MARN 1687552 / QEAC G167) from official government sources.
4|
5|[](https://creativecommons.org/licenses/by/4.0/)
6|[
A real-hand dataset of 6-max No-Limit Texas Hold'em poker played by 19 AI agents
on the dev.fun Arena Beta — frontier LLMs, OSS solver
bots, and equity heuristics — captured during the S8 benchmark run
(May 5–6, 2026).
What this is and is not. This is a derived, settled-hand archive in a
normalized snake_case schema. It is not a mirror of the live
/texas/benchmark/status.table (or /texas/pending-actions[].) shape. Use
it for… See the full description on the dataset page: https://huggingface.co/datasets/dannyobito/arena-pokerkit-hands.faa-balloon-flying-handbook
FAA Balloon Flying Handbook Dataset
This dataset was created by processing the official FAA Balloon Flying Handbook (FAA-H-8083-11B).
If you're interested in understanding how this dataset was created, check out this blog post
or explore the details directly in the GitHub repository.
Usage:
from datasets import load_dataset
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("gsantopaolo/faa-balloon-flying-handbook")
print(dataset)
# Print the first 5 rows… See the full description on the dataset page: https://huggingface.co/datasets/gsantopaolo/faa-balloon-flying-handbook.paper-evo-handHandleAtlas-benchmark
HandleAtlas Benchmark
Hand-labeled NER evaluation set for extracting social-media handles from
Twitter / X bios. These are the exact 100 records (seed = 123) used to
compute the benchmark numbers in the LumeData/HandleAtlas-166m
and LumeData/HandleAtlas-166m-CPU
model cards.
Schema
Each record:
{
"id": 2,
"text": "🍑 Ig | pea_arunya",
"entities": [
{"start": 7, "end": 17, "label": "instagram_username"}
]
}
text — the raw bio (UTF-8, may contain… See the full description on the dataset page: https://huggingface.co/datasets/LumeData/HandleAtlas-benchmark.han-distributed-knowledge-alignment-dataset-v1
Humanoid Distributed Knowledge Alignment Dataset
This dataset models knowledge state synchronization
between humanoid agents operating in a decentralized network.
It captures belief divergence, alignment negotiation,
knowledge merging processes, and consensus validation.
Objective
To enable shared situational awareness
and reduce cognitive divergence
across distributed humanoid agents.
Data Fields
alignment_event_id
agent_id
peer_agent_ids
knowledge_domain… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-knowledge-alignment-dataset-v1.hand_captionshan-dao-proposal-records-v1
Humanoid DAO Proposal Records
This dataset contains governance proposals
submitted within the Humanoid Network DAO.
It enables humanoid agents to analyze,
simulate, and vote on protocol decisions.
Contents
Proposal description
Category
Impact scope
Voting outcome
Use Cases
Governance analysis
DAO participation
Decision modeling
Part of
Humanoid Network (HAN)
License
MIT
Customer_Reviews-Second_Hand_ApparelsI wrote the following script to scrape the data from a platform that sells second-hand apparel and clothing.
Github | Customer Reviews for Second Hand Apparels
The reviews are available for only about 3500 products from the links.txt file. This is only 20% of the total available products. Feel free to clone the script yourself and scrape on your system if you need more data.
reviews.json file contains the data you can use for learning or research purposes. This file contains customer reviews… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticQubit/Customer_Reviews-Second_Hand_Apparels.handpicked-nqhan-domestic-assistance-instruction-evalset-v2
Domestic Assistance Instruction Evaluation Set
A compact evaluation dataset for assessing
instruction understanding in humanoid robots
within domestic environments.
Methods
Instructions are curated to reflect realistic,
single-intent household requests commonly
observed in human-robot interaction studies.
Evaluation Focus
Intent recognition accuracy
Instruction-task alignment
Limitations
This dataset does not measure physical execution quality.… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-domestic-assistance-instruction-evalset-v2.
