datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AutoMathText-V2
🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset
🎉 AutoMathText-v2 has surpassed 1.5 million downloads! We'd love to know how you're using it. Please take 1 minute to fill out our use case survey. Your feedback will directly shape the future roadmap of this dataset.👉 Share your use case here
📊 AutoMathText-V2 consists of 2.46 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSQZ/AutoMathText-V2.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.spreadsheet-bench-v2-modified
SpreadsheetBench V2 Modified: Multi-Document QA
1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents.
An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.afrimedqa_v2
AfriMed-QA v2: A pan-African Medical QA Dataset
🏆 Best Social Impact Paper Award — ACL 2025 (Vienna, Austria, Association for Computational Linguistics)
This work is licensed under a
Creative Commons Attribution 4.0 International License.
Project Website:
AfriMedQA.com
Paper URL: https://aclanthology.org/2025.acl-long.96/
Collaborating Organizations:
Intron Health,
SisonkeBiotik,
BioRAMP,
Georgia Institute of Technology,
MasakhaneNLP,
Google Research
Funded by:
Google Research… See the full description on the dataset page: https://huggingface.co/datasets/afrimedqa/afrimedqa_v2.bangla-crime-investigation-patterns-v2
Bangla Crime Investigation Patterns V2
This dataset is an anonymized, structured extraction from 1,078 public Bangla crime-investigation video transcripts. It was built to support machine-learning and LLM research on crime-pattern information extraction, case summarization, event sequencing, entity/relationship extraction, and Bengali investigative-report analysis.
The public release does not include raw full transcripts. It contains short redacted evidence snippets linked to… See the full description on the dataset page: https://huggingface.co/datasets/tanziro/bangla-crime-investigation-patterns-v2.valor32k-avqa-v2
Valor32k-AVQA v2.0
Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position.
Links
Paper: ACM Digital Library
Project page: inesriahi.github.io/valor32k-avqa-2
Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.XhotpotQA-V2
XHotpotQA V2
AUDITED RC1 · INCOMPLETE
Cross-lingual multi-hop QA over mixed-language evidence
Gemma 4 31B · source-aligned fields · transparent release gate
STATUS · RC1
ROWS · 22,836
GENERATOR · Gemma 4 31B
MISSING · 230
LICENSE · CC BY-SA 4.0
Navigate:
Overview ·
Coverage ·
Load ·
Schema ·
Methodology ·
Quality ·
Citation ·
Paper ·
Collection
RC1 release warning. This is a… See the full description on the dataset page: https://huggingface.co/datasets/Iman998/XhotpotQA-V2.gigaverbo-v2-rec-sft
GigaVerbo-v2 REC SFT
A model should not merely know how to reason; it should learn when reasoning is worth the cost.
Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B
Dataset Summary
GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.tw-legal-benchmark-v2
Taiwan Legal Benchmark v2
A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese,
built from 15 years (2012–2026) of national examinations published by the
Ministry of Examination (考選部).
Supersedes tw-legal-benchmark-v1
(209 questions) with 17,002 deduplicated questions across 15 legal domains.
Overview
Property
Value
Questions
17,002 (deduplicated)
Years
2012–2026
Source papers
1,040 official exam papers
Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.opengloss-v2.1-senses
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Senses
The sense-level view of OpenGloss v2.1 and the repo most consumers want: one row per live sense, with its canonical… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-senses.medhallu-twins-rewritten-v2
MedHallu twins, rewritten answers
Release v2.
Twin pairs for medical hallucination detection. Each pair shares one PubMed
abstract and one question, and holds two answers: a grounded one and a
hallucinated one differing by exactly one edited span.
Both answers are generated by the same model in the same call — here
gemini-3.8-flash, prompt strict-span-v1.
Neither is taken from MedHallu or PubMedQA. This replaces
Certops/medhallu-twins-repaired-context, where the ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/Certops/medhallu-twins-rewritten-v2.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.opengloss-v2.0-senses
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Senses
The sense-level view of OpenGloss v2.0 and the repo most consumers want: one row per live sense, with its canonical gloss, its eight reading-level and register renditions, its sense-tagged example sentences with headword character spans, its… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-senses.SWAP_v2
SWAP_v2: A Synthetic Dataset for Multi-Step Reasoning with Process Supervision
This repository provides the dataset accompanying the paper (ACL 25 main) Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model.
SWAP_v2 consists of synthetically generated trajectories and process-level supervision for multi-step reasoning tasks. Trajectories are generated using DeepSeek-V3.2 on datasets including GSM8K, MATH, FOLIO, ReClor, HumanEval, and MBPP.… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/SWAP_v2.SYMBEX-V2
SYMBEX: Structural Behavioral Evaluation via Counterfactual eXperiments
A graph-based, symmetry-aware agent benchmark that evaluates not just what LLM agents do — but why and how they do it.
Why SYMBEX?
Current agent benchmarks ask: did the agent succeed?
SYMBEX asks: did the agent succeed for the right reasons?
Task success is an incomplete proxy for competence. An agent can:
Produce the correct final answer through a policy-violating shortcut
Fail on a… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/SYMBEX-V2.opengloss-v2.3-senses
Superseded by OpenGloss v2.4 (2026-09-25): every sense now has search queries, QA pairs and verified examples (v2.3 had them only for core and tier 2); level x register definitions and leveled contrasts and explanations are added; and the pretraining corpus no longer contains duplicate documents. v2.3 stays published for reproducibility.
OpenGloss v2.3 — Senses
The sense-level view of OpenGloss v2.3 and the repo most consumers want: one row per live sense, with its canonical… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-senses.opengloss-v2.2-senses
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Senses
The sense-level view of OpenGloss v2.2 and the repo most consumers want: one row per live sense, with its canonical gloss, its eight reading-level and register renditions, its sense-tagged… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-senses.newsqa_200_11064_v2.0.0
Private NewsQA RAG Evaluation Dataset
Private, human-reviewed evaluation source data for the NewsQA RAG project.
This repository is not a prebuilt retrieval index. Chunk, BM25, Chroma, and
ground-truth chunk mappings must be rebuilt from the pinned release.
Version: v1.0.0
Evaluation articles: 200
Distractor articles: 10864
Source questions: 1340
Redistribution rights for the upstream NewsQA-derived text must be verified
before changing this repository from private to public.
vncompress-vi-v2
VNCompress-VI v2 — query-conditioned context compression for Vietnamese
⚠️ BẢN ĐANG REBUILD — WORK IN PROGRESS.
corpus/qa/qa_synthetic đã sạch và ổn định. compression.jsonl mới có
486 hàng thật (prompt v3, trích xuất) trên tổng số dự kiến 100.000+ —
quá trình sinh đang tạm dừng vì tỉ lệ drop "không đạt ngân sách token"
rất cao (90% ở batch gần nhất) chưa được xử lý gốc rễ. Đừng dùng
compression.jsonl để báo cáo kết quả benchmark hay train E5/E6 ở quy mô
lớn — số hàng hiện tại… See the full description on the dataset page: https://huggingface.co/datasets/thanthienhai/vncompress-vi-v2.tw-legal-benchmark-v2
Taiwan Legal Benchmark v2
A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese,
built from 15 years (2012–2026) of national examinations published by the
Ministry of Examination (考選部).
Supersedes tw-legal-benchmark-v1
(209 questions) with 17,002 deduplicated questions across 15 legal domains.
Overview
Property
Value
Questions
17,002 (deduplicated)
Years
2012–2026
Source papers
1,040 official exam papers
Format… See the full description on the dataset page: https://huggingface.co/datasets/Jel1f1sh/tw-legal-benchmark-v2.pkm-agent-baseline-v2
PKM Agent Baseline — 500 + 50 Scenarios + Six-Grader Artifacts (v2)
Two deterministically generated, Korean-language benchmarks for evaluating multi-tool Personal Knowledge Management (PKM) agents over Notion, Gmail, and Google Calendar, plus the Six-Grader Ensemble scoring artifacts (100-scenario reference subset + per-scenario six-metric scores for vanilla and LoRA models).
Released alongside the preprint:
Vault-Grounded 4B Agent: A Hybrid Reasoning–Fact Architecture for Local… See the full description on the dataset page: https://huggingface.co/datasets/newtype-2038/pkm-agent-baseline-v2.quadmix-stem-v2
QuaDMix-STEM v2: STEM-Focused Proxy Validation Set with GPQA & MATH
Script: scripts/validation_set/prepare_stem_v2.py
HuggingFace: liujin99/quadmix-stem-v2
Files: stem_v2_tokenized.pt, stem_v2.parquet
Overview
STEM v2 is an upgraded validation set that fixes the two critical coverage gaps in STEM v1. In the v1 experiment, QuaDMix lost to Random downstream (CORE 0.1530 vs 0.1615), and root-cause analysis revealed:
gpqa_diamond had no direct proxy — mapped from… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v2.vncompress-vi-v2
VNCompress-VI v2 — query-conditioned context compression for Vietnamese
⚠️ BẢN CHƯA HOÀN CHỈNH — WORK IN PROGRESS.
Config compression mới sinh được 973 / ~18.000 cặp (≈ 5%). Quá trình sinh
đang tạm dừng để review. Chưa qua stage 4 (independent verification), nên
trường verification_method của phần lớn row vẫn là none.
Đừng dùng bản này để báo cáo kết quả benchmark.
Bộ dữ liệu cho bài toán nén ngữ cảnh có điều kiện truy vấn tiếng Việt: cho một
câu hỏi và một tài liệu dài, sinh… See the full description on the dataset page: https://huggingface.co/datasets/anhalu/vncompress-vi-v2.afrimedqa_v2
AfriMed-QA v2: A pan-African Medical QA Dataset
This work is licensed under a
Creative Commons Attribution-ShareAlike 4.0 International License.
Project Website:
AfriMedQA.com
Arxiv: https://arxiv.org/abs/2411.15640
Collaborating Organizations:
Intron Health,
SisonkeBiotik,
BioRAMP,
Georgia Institute of Technology,
MasakhaneNLP,
Google Research
Funded by:
Google Research,
Bill & Melinda Gates Foundation,
PATH,
Summary
AfriMed-QA creates a novel multispecialty… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrimedqa_v2.Saudi-historical-places-v2
Saudi Historical Places v2 (verified)
A multimodal Arabic corpus of Saudi historical sites, spanning all thirteen
administrative regions. Each record pairs curated structured attributes with
encyclopedic narrative text and screened photography.
This is a corrected release. Version 1 linked each site to an Arabic
Wikipedia article using title lookup with a fuzzy fallback; auditing that linkage
found that 31.4% of links (Wilson 95% CI 26.8-36.4, n=357) denoted a
different place… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Saudi-historical-places-v2.synthetic-science-v2-sample
Synthetic Scientific Research Threads — v2 (sample)
A synthetic continual-learning benchmark: each episode is a coherent sequence of
short fictional scientific research documents about a single made-up entity, with
per-document QA anchors. Later documents build on, revise, or supersede earlier
ones. Designed to stress test-time / meta-learning approaches where a model must
adapt to a stream of documents and answer questions grounded in what it has just
seen.
This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/synthetic-science-v2-sample.pytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.vietnamese-evidence-corpus-chunked-e5-v2
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.
