datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASearcher-Local-Knowledgedojo_main_income
Languages: 简体中文 · English
dojo_main_income — Revenue Breakdown
Overview
Segment-level main business revenue from listed companies, by industry, product, and region, with amounts and mix ratios. Corresponds to “main business by segment” notes in filings.
Files
File
Description
data.parquet
Full revenue breakdown detail
Key Fields
Field
Description
symbol
Stock symbol
security_name
Company name… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_main_income.include-base-44
INCLUDE-base (44 languages)
Dataset Description
Paper: http://arxiv.org/abs/2411.19799
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-base-44.adult-census-income
Adult Census Income Dataset
The following was retrieved from UCI machine learning repository.
This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics). A set of reasonably clean records was extracted using the following conditions: ((AAGE>16) && (AGI>100) && (AFNLWGT>1) && (HRSWK>0)). The prediction task is to determine whether a person makes over $50K a year.
Description of fnlwgt (final weight)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/adult-census-income.ZwZ-RL-VQA
ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception
This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use.
📖 Overview
The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive:
Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V)… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZwZ-RL-VQA.FinFIRST
FinFIRST: Financial Information Retrieval, Sourcing and Traceability
Released alongside Ling-3.0-flash-Fin, FinFIRST is an open benchmark for evaluating whether financial search agents can produce answers that are not only correct, but also supported by authoritative, timely, and verifiable evidence. It was developed by Ant Group, with professional support from the investment banking team at China International Capital Corporation Limited (CICC).
Financial research requires more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinFIRST.FinixDocBench
FinixDocBench
Language: English | 中文
This repository contains a compliance-reviewed public subset of FinixDocBench, the financial-domain document parsing benchmark introduced in the technical report "FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks".
The benchmark focuses on document parsing conditions that are common in real financial workflows but underrepresented in saturated clean-document benchmarks: digitally native insurance clauses, noisy… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinixDocBench.include-lite-44
INCLUDE-lite (44 languages)
Dataset Description
Paper: http://arxiv.org/abs/2411.19799
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.folktables-acs-income
Dataset Card for "folktables-acs-income"
More Information needed
VenusBench-CAPTCHA
VenusBench-CAPTCHA: A Real-World CAPTCHA Screenshot–Action Benchmark for GUI Agents
Evaluation Code: https://github.com/inclusionAI/UI-Venus/tree/VenusBench-CAPTCHA
Introduction
CAPTCHA solving is a practical challenge for multimodal GUI agents because it requires more than isolated visual recognition. An agent must understand the challenge instruction, identify the relevant interface region, recognize or reason about the visual target, ground the result… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-CAPTCHA.FreshRetailNet-50K
FreshRetailNet-50K
Dataset Overview
FreshRetailNet-50K is the first large-scale benchmark for censored demand estimation in the fresh retail domain, incorporating approximately 20% organically occurring stockout data. It comprises 50,000 store-product 90-day time series of detailed hourly sales data from 898 stores in 18 major cities, encompassing 865 perishable SKUs with meticulous stockout event annotations. The hourly stock status records unique to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dingdong-Inc/FreshRetailNet-50K.Nine-Bus-Load-Increase-EventLing-Coder-SFT
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ling-Coder Dataset
The Ling-Coder Dataset comprises the following components:
Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples.
Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples.
Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.Gutenberg-BookCorpus-Cleaned-Data-English
Gutenberg-BookCorpus-Cleaned-Data-English
This dataset is been cleaned and preprocessed using Gutenberg_English_Preprocessor class method (given below) from preference Kaggle dataset 75,000+ Gutenberg Books and Metadata 2025. This dataset is only specialisation for english contented with rights as "Public domain in the USA" hence you can free used it anywhere.
Following reference metadata of Gutenberg is also available and downloaded it using following CLI command below :-
pip… See the full description on the dataset page: https://huggingface.co/datasets/incredible45/Gutenberg-BookCorpus-Cleaned-Data-English.AudioMCQ
[ICLR 2026] AudioMCQ: Audio Multiple-Choice Question Dataset
Also the official repository for the paper "Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models"
News
[2026.04] Update on MMSU Metric of released models: Based on community feedback, we identified a flaw in our evaluation script that artificially inflated the MMSU scores of our released models by ignoring sequence order. We sincerely apologize for… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AudioMCQ.CCTV_Incident_Dataset_Fall_Lying_Down_Detection
Overview
This is an open-source synthetic dataset for Computer Vision (CV) tasks, specifically designed for Fall Detection, Pose Estimation, and Incident Monitoring from overhead CCTV perspectives.
Unlike standard object detection datasets, this dataset includes Keypoints (Pose) annotations. This enables models to understand human posture and accurately distinguish between standing and fallen individuals.
🚀 Need more data?
This is a sample dataset by Simuletic. We provide… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/CCTV_Incident_Dataset_Fall_Lying_Down_Detection.forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.olympiad-math-training-pool
Olympiad mathematics training pool
Public olympiad and competition mathematics, four datasets gathered at pinned revisions, shipped
twice over. sources/ holds each dataset the way its publisher ships it, in its own file format
with its own fields and nothing renamed, 229052 rows across four folders. pool/ holds the union
of those same datasets in one format, one JSON object per line, deduplicated by problem text and
reduced to 225822 rows, every row labelled with the dataset it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/olympiad-math-training-pool.sample-parquetSample Parquet dataset for testing purposes
FaithEval-inconsistent-v1.0
FaithEval
FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts.
[Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727
[Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval
Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-inconsistent-v1.0.Ling-Coder-SyntheticQA
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ling-Coder Dataset
The Ling-Coder Dataset comprises the following components:
Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples.
Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples.
Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA.inception-v1-microscope-data
Inception V1 Microscope Data
This dataset powers the Inception V1
Microscope,
an interactive interface for exploring visual features learned by individual
neurons in Inception V1.
It combines two complementary interpretability views:
Activation maximization: one synthesized visualization optimized to
strongly activate each neuron.
Top dataset examples: the ten ImageNet examples producing the strongest
recorded activations for each neuron, paired with crops associated with the… See the full description on the dataset page: https://huggingface.co/datasets/akankshanc/inception-v1-microscope-data.SWE-CARE
SWE-CARE: A Comprehensiveness-aware Benchmark for Code Review Evaluation
Dataset Description
SWE-CARE (Software Engineering - Comprehensive Analysis and Review Evaluation) is a comprehensiveness-aware benchmark for evaluating Large Language Models (LLMs) on repository-level code review tasks. The dataset features real-world code review scenarios from popular open-source Python and Java repositories, with comprehensive metadata and… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/SWE-CARE.ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1392 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.INCLUDE
Dataset Card for INCLUDE
Dataset Summary
This dataset contains all videos in the INCLUDE dataset. As huggingface does not support video uploads at this time, the HF dataset contains metadata about each video such as the parent class, the video class, the path to the video and whether its a part of the INCLUDE-50 dataset (use include_50==True to get only include_50 videos).
The videos themselves can be downloaded from Zenodo using the provided bash script.… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/INCLUDE.uci-adult-income
Dataset Card for "uci-adult-income"
More Information needed
incantation-elden-ring-scenes
Incantation Elden Ring Combat Captions
Paper | Project page | GitHub
Preview subset. This repository is a public preview and reference subset of the Incantation dataset. It documents the data format, annotation style, and initial training material used by the project. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper.
This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes.adult_income_datasetPubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.ZoomBench
ZoomBench: A Fine-Grained Multimodal Perception Benchmark
📃 Paper | 🏠 Project | 🤗 Models
Overview
ZoomBench is a challenging benchmark designed to evaluate the fine-grained multimodal perception capabilities of Multimodal Large Language Models (MLLMs). It specifically targets scenarios where decisive visual evidence is small, subtle, or easily overwhelmed by global context — situations that demand "zooming-level" perception from a single full image.
It is… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZoomBench.
