datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.officeqa-pro-v2
OfficeQA Pro v2
Dataset Summary
OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.afrimedqa_v2
AfriMed-QA v2: A pan-African Medical QA Dataset
🏆 Best Social Impact Paper Award — ACL 2025 (Vienna, Austria, Association for Computational Linguistics)
This work is licensed under a
Creative Commons Attribution 4.0 International License.
Project Website:
AfriMedQA.com
Paper URL: https://aclanthology.org/2025.acl-long.96/
Collaborating Organizations:
Intron Health,
SisonkeBiotik,
BioRAMP,
Georgia Institute of Technology,
MasakhaneNLP,
Google Research
Funded by:
Google Research… See the full description on the dataset page: https://huggingface.co/datasets/afrimedqa/afrimedqa_v2.IndicQuest-v2
IndicQuest v2
A gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models. 3,471 curriculum-grounded English question–answer pairs across nine domains, translated into 19 Indic languages: 69,420 parallel pairs across 20 languages.
More details can be found in our paper.
Dataset structure
One CSV per language, named <language>.csv (english.csv, hindi.csv, marathi.csv, …). Every file has the… See the full description on the dataset page: https://huggingface.co/datasets/l3cube-pune/IndicQuest-v2.tw-legal-benchmark-v2
Taiwan Legal Benchmark v2
A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese,
built from 15 years (2012–2026) of national examinations published by the
Ministry of Examination (考選部).
Supersedes tw-legal-benchmark-v1
(209 questions) with 17,002 deduplicated questions across 15 legal domains.
Overview
Property
Value
Questions
17,002 (deduplicated)
Years
2012–2026
Source papers
1,040 official exam papers
Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.tw-legal-benchmark-v2
Taiwan Legal Benchmark v2
A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese,
built from 15 years (2012–2026) of national examinations published by the
Ministry of Examination (考選部).
Supersedes tw-legal-benchmark-v1
(209 questions) with 17,002 deduplicated questions across 15 legal domains.
Overview
Property
Value
Questions
17,002 (deduplicated)
Years
2012–2026
Source papers
1,040 official exam papers
Format… See the full description on the dataset page: https://huggingface.co/datasets/Jel1f1sh/tw-legal-benchmark-v2.afrimedqa_v2
AfriMed-QA v2: A pan-African Medical QA Dataset
This work is licensed under a
Creative Commons Attribution-ShareAlike 4.0 International License.
Project Website:
AfriMedQA.com
Arxiv: https://arxiv.org/abs/2411.15640
Collaborating Organizations:
Intron Health,
SisonkeBiotik,
BioRAMP,
Georgia Institute of Technology,
MasakhaneNLP,
Google Research
Funded by:
Google Research,
Bill & Melinda Gates Foundation,
PATH,
Summary
AfriMed-QA creates a novel multispecialty… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrimedqa_v2.squad_v2.0RoMedQA_v2math-code-qa-v2
Math & Code QA v2 — Instruction Dataset
Worked mathematical solutions and short code answers, spanning arithmetic word
problems through to algebra, geometry and combinatorics.
Built for the Adaption Labs AutoScientist Challenge (Math & Code category).
The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the
held-out Math category evaluation.
Rows
5,297 (4,197 math, 1,100 code)
Distinct answers
5,297 (100%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.datos-leyes-civiles-peruanas-v2
Datos Leyes Civiles Peruanas v2
Este dataset es una variante de SrAlex/datos-leyes-civiles-peruanas-v2.
El dataset original presentaba un prompt en formato prompt, que incluia las preguntas y respuestas dentro de una misma columna:
<s>[INST] Eres un experto en leyes peruanas, dime el significado de Jurista[/INST] El significado de Jurista es Se dice de quién es versado en la ciencia del derecho, es el que se dedica a la resolución de las dudas o consultas jurídicas. Es el agente de… See the full description on the dataset page: https://huggingface.co/datasets/elsatch/datos-leyes-civiles-peruanas-v2.Startups_V2
