datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenAoE-2000h
Open-AoE — Egocentric Hand Manipulation Dataset
Release Roadmap
Tier
Duration
Status
nano
~3 h
✅ Released
tiny
~100 h
✅ Released
full
2000 h
🚧 Uploading
Release notes
2026-07-30: Removed samples flagged in PR #1 for camera-intrinsics vs. video-resolution mismatches.
2026-07-31: Uploaded ~323h of data.
2026-08-12: Uploaded ~694h of data.
2026-09-03: Uploaded ~189h of data.
Additional data for the full ~2000h release is still… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/OpenAoE-2000h.reddit_mds_incrementalConceptEdit-12M
ConceptEdit: Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
ConceptEdit-12M is a large-scale image editing dataset. Each sample is stored as a triplet:
a source image,
an edited image,
a JSON metadata file describing the edit instruction, edit category, relative image paths, and VQA-style quality checks.
The dataset is packaged as multiple .tar shards. All paths inside the tar files and JSON files are relative paths; no… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ConceptEdit-12M.GenFusion_Training_Datacc_en_middle_mds_incrementalHackerNewsContains all Hacker News posts through April 15th, 2025.
ASearcher-Local-Knowledgedojo_main_income
Languages: 简体中文 · English
dojo_main_income — Revenue Breakdown
Overview
Segment-level main business revenue from listed companies, by industry, product, and region, with amounts and mix ratios. Corresponds to “main business by segment” notes in filings.
Files
File
Description
data.parquet
Full revenue breakdown detail
Key Fields
Field
Description
symbol
Stock symbol
security_name
Company name… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_main_income.include-base-44
INCLUDE-base (44 languages)
Dataset Description
Paper: http://arxiv.org/abs/2411.19799
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-base-44.USA-Map-Tiles
Dataset Description
The dataset is a set of map tiles for the various states of the USA.
License: Open Data Commons Open Database License (ODbL) v1.0
Dataset Sources
OpenStreetMap.org
Geofabrik Download Server
Bounding Boxes per State
Dataset Structure
Organized into folders by state
Structure of graphml files:
Top of files have a list of keys/ids that correspond to the properties of the segment:
maxspeed: speed limit
oneway: if it is a… See the full description on the dataset page: https://huggingface.co/datasets/incognitolm/USA-Map-Tiles.VenusBench-GD
VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
Project Page: https://ui-venus.github.io/VenusBench-GD/
Introduction
GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, or focus excessively on a single platform and require highly specialized domain… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-GD.incadult-census-income
Adult Census Income Dataset
The following was retrieved from UCI machine learning repository.
This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics). A set of reasonably clean records was extracted using the following conditions: ((AAGE>16) && (AGI>100) && (AFNLWGT>1) && (HRSWK>0)). The prediction task is to determine whether a person makes over $50K a year.
Description of fnlwgt (final weight)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/adult-census-income.tulu_flan_mds_incremental-tokensZwZ-RL-VQA
ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception
This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use.
📖 Overview
The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive:
Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V)… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZwZ-RL-VQA.FinFIRST
FinFIRST: Financial Information Retrieval, Sourcing and Traceability
Released alongside Ling-3.0-flash-Fin, FinFIRST is an open benchmark for evaluating whether financial search agents can produce answers that are not only correct, but also supported by authoritative, timely, and verifiable evidence. It was developed by Ant Group, with professional support from the investment banking team at China International Capital Corporation Limited (CICC).
Financial research requires more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinFIRST.refinedweb_mds_incrementalc4_mds_incrementaltulu_flan_mds_incrementalcc_news_mds_incremental-tokensogb-full-original
OGB full original archives
Public, byte-for-byte mirror of 17 official Open Graph Benchmark (OGB) and OGB Large-Scale Challenge archive downloads used by the Emulated-Inc graph benchmark environment.
Original ZIP archives are stored under archives//. Each dataset directory includes metadata.json with the authoritative source URL, exact byte size, SHA-256 digest, and repository archive path. Archives were transferred directly from the official SNAP/DGL hosts through ephemeral… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/ogb-full-original.In-cylinder_Flow_Fieldfolktables-acs-income
Dataset Card for "folktables-acs-income"
More Information needed
FinixDocBench
FinixDocBench
Language: English | 中文
This repository contains a compliance-reviewed public subset of FinixDocBench, the financial-domain document parsing benchmark introduced in the technical report "FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks".
The benchmark focuses on document parsing conditions that are common in real financial workflows but underrepresented in saturated clean-document benchmarks: digitally native insurance clauses, noisy… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinixDocBench.include-lite-44
INCLUDE-lite (44 languages)
Dataset Description
Paper: http://arxiv.org/abs/2411.19799
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.VenusBench-CAPTCHA
VenusBench-CAPTCHA: A Real-World CAPTCHA Screenshot–Action Benchmark for GUI Agents
Evaluation Code: https://github.com/inclusionAI/UI-Venus/tree/VenusBench-CAPTCHA
Introduction
CAPTCHA solving is a practical challenge for multimodal GUI agents because it requires more than isolated visual recognition. An agent must understand the challenge instruction, identify the relevant interface region, recognize or reason about the visual target, ground the result… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-CAPTCHA.Nine-Bus-Load-Increase-EventFreshRetailNet-50K
FreshRetailNet-50K
Dataset Overview
FreshRetailNet-50K is the first large-scale benchmark for censored demand estimation in the fresh retail domain, incorporating approximately 20% organically occurring stockout data. It comprises 50,000 store-product 90-day time series of detailed hourly sales data from 898 stores in 18 major cities, encompassing 865 perishable SKUs with meticulous stockout event annotations. The hourly stock status records unique to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dingdong-Inc/FreshRetailNet-50K.cc_en_head_mds_incrementalLing-Coder-SFT
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ling-Coder Dataset
The Ling-Coder Dataset comprises the following components:
Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples.
Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples.
Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.
