datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASearcher-Local-KnowledgeFinFIRST
FinFIRST: Financial Information Retrieval, Sourcing and Traceability
Released alongside Ling-3.0-flash-Fin, FinFIRST is an open benchmark for evaluating whether financial search agents can produce answers that are not only correct, but also supported by authoritative, timely, and verifiable evidence. It was developed by Ant Group, with professional support from the investment banking team at China International Capital Corporation Limited (CICC).
Financial research requires more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinFIRST.AudioMCQ
[ICLR 2026] AudioMCQ: Audio Multiple-Choice Question Dataset
Also the official repository for the paper "Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models"
News
[2026.04] Update on MMSU Metric of released models: Based on community feedback, we identified a flaw in our evaluation script that artificially inflated the MMSU scores of our released models by ignoring sequence order. We sincerely apologize for… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AudioMCQ.incantation-elden-ring-scenes
Incantation Elden Ring Combat Captions
Paper | Project page | GitHub
Preview subset. This repository is a public preview and reference subset of the Incantation dataset. It documents the data format, annotation style, and initial training material used by the project. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper.
This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes.joyo-kanji-yomi-benchmark-parakeet
日本語 | English
常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet)
常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。
このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。
このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。
概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.2026-09-15-dataset-refresh-incomplete-audit
INCOMPLETE RESEARCH AUDIT — NOT A TRAINING DATASET
field
value
experiment
Incomplete retained research pools: moral low stakes has 706 rows (10 short of 716: t1=2, t4=2, t6=1, t7=4, t8=1); original craft nonmoral has 631 rows (85 short: t1=7, t2=11, t3=10, t4=8, t5=12, t6=12, t7=5, t8=8, t9=12). Shared spend exposure is $249.2113677 of $250, with no active calls or uncertain reservations. Both pools are byte-identical subsets of 708/634-row snapshots that passed… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-incomplete-audit.MultiEdit
🧩 MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
📃 Arxiv
🚀 Dataset Overview
Based on our MLLM-driven data construction pipeline using GPT-4o and GPT-Image-1, we introduce MultiEdit, a comprehensive large-scale
instruction-based image editing dataset comprising over 107K samples targeting 6 challenging image editing tasks covering 56 subcategory
editing types (18 non-style-transfer and 38 style transfer). We also release… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/MultiEdit.A3S-Bench
Agent3Sigma-Stage (A3S-Bench)
💻 GitHub | 🏆 Leaderboard | 📄 Paper (PDF) | arXiv
Agent3Sigma-Stage (A3S-Bench) is an end-to-end security evaluation framework for autonomous agents (e.g., OpenClaw), designed to systematically measure both an Agent's ability to resist attacks during multi-turn interactions and its utility in completing legitimate tasks. The framework provides an evaluation dataset covering 10 security risk categories across 6 real-world usage scenarios… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/A3S-Bench.Ring-lite-sft-data
🤖 ModelScope
🤗 HuggingFace
🖥️ GitHub
Ring-lite-sft-data
This is a the SFT data used during the fine-tuning of the Ring-lite model. The query pool was sourced from open-source repositories and further enriched through synthetic generation using large language models (LLMs). To ensure the production of high-fidelity responses with Long-CoT, we implemented an iterative refinement pipeline that synergistically combines automated model… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ring-lite-sft-data.devops-incident-response
Dataset Card for DevOps Incident Response Dataset
Dataset Description
Dataset Summary
The DevOps Incident Response Dataset is a comprehensive collection of real-world-style DevOps incidents, troubleshooting scenarios, and resolution procedures. This dataset is designed to help train AI models for DevOps assistance, incident response automation, and technical troubleshooting education.
Each incident includes:
Detailed incident description and symptoms… See the full description on the dataset page: https://huggingface.co/datasets/Snaseem2026/devops-incident-response.wildchat_4m_inc_multi_no_dedup_shuffledai-incidents-2026
AI Agent Security Incident Intelligence — 2026
Free sample: 100 incidents (of 1,038+ classified)
This sample includes type, severity, and source URL. No descriptions or causal analysis.
Full dataset
The complete intelligence package includes:
1,038+ classified incidents with full causal analysis
21 incident categories including emerging patterns
Root cause chain for every incident
Classifier versioning and confidence scores
Updated weekly with new incidents
Full… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-incidents-2026.Japanese-RAG-Generator-Benchmark
Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク
Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。
実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。
J-RAGBenchの評価カテゴリ
Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く
Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する
Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く
Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.logical-reasoning-training-pool
Logical reasoning training pool
Public logical-reasoning problems with checkable answers, from six datasets, read at the pinned
revisions named below and laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 579910 rows, one JSON object per line. A row is one of
three kinds and carries the fields its kind needs.
Field
What it holds
id
a row identifier unique within this file
kind
entailment, choice or open… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logical-reasoning-training-pool.SingStreamBench
SingStreamBench
SingStreamBench is a benchmark for streaming safety detection. It evaluates whether a guardrail stays silent on benign response prefixes and triggers promptly once harmful content actually begins—rather than relying on shortcut signals such as harmful queries, response position, or surface keywords.
The released core set contains 210 human-verified English samples. Each sample provides a query, a complete response, and the character-level onset of unsafe content.… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/SingStreamBench.genai-incidents
GenAI & Agentic AI Security Incidents
13,060 real-world and research incidents involving generative-AI and agentic-AI
systems — prompt injection, jailbreaks, data exfiltration, deepfakes, model and
supply-chain compromise, agent hijacking, and AI-enabled harms — cross-mapped to six
taxonomies. Dataset version 2.10.0.
Every applicable incident is tagged with four core taxonomies:
OWASP Top 10 for LLM Applications (2026) — LLM01–LLM10
OWASP Agentic Top 10 (ASI) — ASI01–ASI10
NIST… See the full description on the dataset page: https://huggingface.co/datasets/emmanuelgjr/genai-incidents.python-unit-test-training-pool
Python unit test training pool
A pool of public data for training a model to write tests for Python code. It is a
straight collection of open datasets, not a new corpus: every row comes from one of the
sources below, at the revision named, and the only rows removed are the ones an overlap
filter flagged against held-out material this pool is kept separate from.
Every row of the normalised layer pairs a program with tests for it. That is the point of
the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.legal-ai-incidents
SafeLegalAI Legal AI Incident Tracker
What went wrong when AI met the courtroom, and what did courts and regulators do about it?
150 incidents · 15 jurisdictions · 48 with a recorded regulatory outcome · last checked 2026-09-05 · synced from safelegalai.com on 2026-09-08.
Every court case where AI misuse reached a judgment or order: the court, the date, what was fabricated or misused, the outcome, any penalty, who the actor was, and — the layer no other tracker keeps — what the… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/legal-ai-incidents.INC_Data
INC Dataset: Implicit Neural Correction for PDE Solvers
Dataset Description
This dataset contains simulation data for training and evaluating implicit neural correction methods for partial differential equation (PDE) solvers. The dataset includes two challenging dynamical systems demonstrating complex spatiotemporal behaviors:
Kuramoto-Sivashinsky (KS) Equation - 1D chaotic dynamics
Backward-Facing Step (BFS) Flow - 2D incompressible Navier-Stokes with complex geometry… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/INC_Data.arxiv-chandra-ocr-2-include-images-first50-20260415
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 50
Successes: 50
Partial successes: 0
Errors: 0
Next shard index: 10
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.incantation-elden-ring-scenes
Incantation Elden Ring Combat Captions
Preview subset. This repository is an early public preview and reference subset of the Incantation dataset. It is provided to document the data format, annotation style, and initial training material ahead of the full paper/project release. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper.
This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/MatrixTeam/incantation-elden-ring-scenes.in-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.AReaL-boba-2-RL-Code
AReaL-boba-2-RL-Code Dataset
This dataset contains the training and testing data used for reinforcement learning in the AReaL-boba-2 model on coding tasks.
Dataset Structure
train/: Contains the training data for RL fine-tuning.
code_benchmark/: Contains the evaluation benchmark, organized into multiple subfolders, each corresponding to a specific coding benchmark suite, Codeforces, Code Contests and LiveCodeBench (v5) are supported now.
How to Use
To train… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AReaL-boba-2-RL-Code.Ring-lite-rl-data
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ring-lite-rl-data
This dataset is a curated subset of high-quality problems across mathematics and code domains designed for reinforcement learning in the Ring-lite model. This dataset contains:
Mathematics: Over 39,000 rigorously curated problems sourced from:
Open-source datasets (BigMath, DeepScaleR, DAPO, DeepMath-103K)
Art of Problem Solving (AoPS) contest collections
Code: Approximately 8,400… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ring-lite-rl-data.python-functions-training-pool
Python function-writing training pool
A pool of public data for training a model to write Python functions. It is a straight
collection of open datasets, not a new corpus: every row comes from one of the sources
below, at the revision named, and the only rows removed are the ones an overlap filter
flagged against held-out material this pool is kept separate from.
Rows in the normalised layer: 5756045.
Rows in the raw layer: 6258415.
The two layers
pool/ holds the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-functions-training-pool.procedural-reasoning-training-pool
Procedural reasoning training pool
Reasoning questions from 101 procedural generators, each of which writes a question, computes its
own answer and ships a verifier that scores an attempt at it, plus a collection of solved Sudoku
puzzles. Every answer is short and exactly checkable, so a trained model can be marked against the
key by a program and no judge is needed. Laid out twice. Train on either layer or on both.
pool.jsonl
Every generator rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/procedural-reasoning-training-pool.Arabic-IFEvalIFEval is the first publicly available benchmark dataset specifically designed to evaluate Arabic Large Language Models (LLMs) on instruction-following capabilities in Arabic.
The dataset includes 404 high-quality, manually verified samples covering various constraints such as linguistic patterns, punctuation rules, and formatting guidelines.
Loading the Dataset
To load this dataset in Python using the 🤗 Datasets library, run the following:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/inception42/Arabic-IFEval.reading-comprehension-training-pool
Reading comprehension training pool
Public reading comprehension questions from six datasets, each a question about a passage with an
answer that is a span of it, a number or a date, read at the pinned revisions named below and laid
out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 310728 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/reading-comprehension-training-pool.json-schema-instances-training-pool
JSON schema and instance training pool
Real JSON Schemas from the public collections named below, read at the pinned revisions given there,
each paired where possible with documents that satisfy it, laid out twice. Train on either layer or
on both.
pool.jsonl
Every source rewritten into one shape, 20004 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
prompt
the request a model would… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/json-schema-instances-training-pool.Indian-Income-Tax-Returns
Indian Income Tax Return Synthetic Dataset
A fully synthetic, high-fidelity dataset of Indian Income Tax Return forms (ITR-4, ITR-5, and ITR-6). Designed to support OCR, text extraction, table parsing, tax attribute detection, document intelligence, and LLM fine-tuning for structured data extraction. Each record includes a PDF tax return and a matching structured JSON file containing parsed fields.
This dataset simulates realistic taxpayer filings across:
Individuals (with Aadhaar… See the full description on the dataset page: https://huggingface.co/datasets/AgamiAI/Indian-Income-Tax-Returns.
