datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.bitaudit_verification_dataset_v2afrimedqa_v2
AfriMed-QA v2: A pan-African Medical QA Dataset
🏆 Best Social Impact Paper Award — ACL 2025 (Vienna, Austria, Association for Computational Linguistics)
This work is licensed under a
Creative Commons Attribution 4.0 International License.
Project Website:
AfriMedQA.com
Paper URL: https://aclanthology.org/2025.acl-long.96/
Collaborating Organizations:
Intron Health,
SisonkeBiotik,
BioRAMP,
Georgia Institute of Technology,
MasakhaneNLP,
Google Research
Funded by:
Google Research… See the full description on the dataset page: https://huggingface.co/datasets/afrimedqa/afrimedqa_v2.Nemotron-SFT-Agentic-v2-prompt-only
Nemotron-SFT-Agentic-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.failbench-robocasa-v2
FailBench RoboCasa v2 — contact-prediction dataset
Labeled robot-failure trials built on RoboCasa kitchen demos (PandaMobile / Franka).
Each trial injects a hardware failure partway through a teleop/MimicGen demo, then records the
contacts the failure causes during a 1-second settle. The supervised target is a
240×320 force-weighted contact heatmap in the agentview camera — the model learns to predict
where a failure at a given pre-failure configuration will drive the… See the full description on the dataset page: https://huggingface.co/datasets/aaronngx/failbench-robocasa-v2.Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.polaris-wikitables-v2
WikiTables v2
WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from
Retrieval Feedback, alongside aw, arctic, lter, ecir,
and wtr.
It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each
query–table pair, a person scored how well that table answers that query; those scores are the
relevance judgments, and they live in qrels.csv.
The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v2.NF-ToN-IoT-v2gcfairbench-v2
GCFairBench-100 (v3.2) — DOPP-CRAFT-GC experiment images
Interactive gallery: https://nati1221-craft-gc-gallery.static.hf.space/Human evaluation (Wave 2 recruiting): https://huggingface.co/spaces/nati1221/craft-gc-human-eval
2,500 unique text-to-image outputs for the DOPP-CRAFT-GC Springer / PanAfriCon research evaluation.
Field
Value
Prompts
100 (GCFairBench-100: 40 SSA, 15 each other region)
Methods
Base SD, PromptAug, PromptAug-Explicit, FairImagen-GC… See the full description on the dataset page: https://huggingface.co/datasets/nati1221/gcfairbench-v2.wooden_window_factory_01_enriched_v2
Real industrial data, AI-ready for Physical AI
ORION WWF1 – Certified Sample Pack v2.0 (Enriched)
Version
Status
Sector
Pipeline
v2.0-Enriched
🟢 Level 3 Certified
Industrial-Manufacturing
Orion Unified V5.2
🌟 The Evolution: Beyond Anonymization
The ORION WWF1 v2.0 Enriched pack represents the professional evolution of our baseline industrial dataset. While previous versions focused on privacy-first anonymization, v2.0 transforms raw video… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v2.tw-legal-benchmark-v2
Taiwan Legal Benchmark v2
A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese,
built from 15 years (2012–2026) of national examinations published by the
Ministry of Examination (考選部).
Supersedes tw-legal-benchmark-v1
(209 questions) with 17,002 deduplicated questions across 15 legal domains.
Overview
Property
Value
Questions
17,002 (deduplicated)
Years
2012–2026
Source papers
1,040 official exam papers
Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.denoised-reazonspeech-v2-dnsmosDNSMOS score of reazon-speech-v2-denoised
llm-classification-distilled-v2-sharded
LLM Classification Distilled v2 Sharded
Overview
This repository stores shard CSV files produced by the teacher-judge distillation pipeline.
How to Use
Run the distillation notebook once per shard:
NUM_SHARDS = 4
SHARD_INDEX = 0 .. 3
After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos.
Final Repositories
Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.im3_open_source_data_center_atlas_v2026.02.09
IM3 Open Source Data Center Atlas v2026.02.09 — refined database
This repository preserves the IM3 Open Source Data Center Atlas v2026.02.09 and
adds a source-enriched, audited 43-column power-source table for all 1,479
source geometry records (1,474 unique IM3 IDs). The publication retains the exact
13 upstream columns plus 30 stable label, interpretation, and evidence fields.
Duplicate geometry records are intentionally retained.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3_open_source_data_center_atlas_v2026.02.09.HPA-v25-subcellular
Human Protein Atlas v25: subcellular localization and full annotation tables
A verbatim mirror of three Human Protein Atlas (HPA) release files, packaged together
with an explicit provenance and schema description. Nothing here is derived, filtered,
or recomputed: the bytes are exactly what the upstream endpoints served.
The point of this repository is reproducibility. HPA serves only the current release at
a stable URL, so a file downloaded six months from now will silently… See the full description on the dataset page: https://huggingface.co/datasets/Boom5426/HPA-v25-subcellular.WIT-es_jina-clip-v2_samplePrediction-Smartphone-Addiction-Submission-V2Here's the enhanced version, honest summary of what actually moved the needle:
What improved it:
Feature engineering was the real driver: missingness indicators for every column (missingness itself carries signal here), plus ratio/interaction features like social_to_screen, sleep_deficit, weekday_weekend_diff, screen_per_age, etc.
LightGBM with these new features: OOF AUC 0.9628 (up from 0.9620).
Result:
submission_v2.csv, same valid format (296,302 rows… See the full description on the dataset page: https://huggingface.co/datasets/Qamro/Prediction-Smartphone-Addiction-Submission-V2.Nemotron-SFT-Competitive-Programming-v2-prompt-only
Nemotron-SFT-Competitive-Programming-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Competitive-Programming-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Competitive-Programming-v2-prompt-only.polaris-aw-v2
AW v2
AW is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside arctic, lter, ecir, wikitables, and
wtr.
It holds 96 tables from a version of AdventureWorks whose column names are cryptic — BusEntId,
STRGUID, JobTtl — and 15 keyword queries over them. For each query–table pair, a person decided
whether that table answers that query; those decisions are the relevance judgments, and they live in
qrels.csv.
Each Polaris dataset… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-aw-v2.Nemotron-RL-Math-v2-prompt-only
Nemotron-RL-Math-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Math-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Math-v2-prompt-only.polaris-ecir-v2
ECIR v2
ECIR is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, arctic, lter, wikitables, and
wtr.
It holds 2,100 tables published on the US government open data portal and 12 keyword queries over
them. For each query–table pair, a person decided how well that table answers that query and gave it
a score; those scores are the relevance judgments, and they live in qrels.csv. Given a query, a
system ranks the 2,100… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-ecir-v2.eterna100-v2Scripts and results for benchmarking RNA design algorithms with the Eterna100-V1 and Eterna100-V2 benchmarks
scgnn-v2-storebambu-timelapse-dataset
Bambu Timelapse Dataset
Dataset Summary
Disclaimer: This dataset is an independent community project and is not affiliated with or endorsed by Bambu Lab in any official capacity. We simply curate footage captured on its consumer printers to enable open research.
The Bambu Timelapse Dataset is an open, community‑driven collection of time‑lapse videos captured on Bambu Lab 3‑D printers (P1 series, X1 series and variants).
Its goal is to provide a high‑quality video corpus… See the full description on the dataset page: https://huggingface.co/datasets/v2thegreat/bambu-timelapse-dataset.afrimedqa_v2
AfriMed-QA v2: A pan-African Medical QA Dataset
This work is licensed under a
Creative Commons Attribution-ShareAlike 4.0 International License.
Project Website:
AfriMedQA.com
Arxiv: https://arxiv.org/abs/2411.15640
Collaborating Organizations:
Intron Health,
SisonkeBiotik,
BioRAMP,
Georgia Institute of Technology,
MasakhaneNLP,
Google Research
Funded by:
Google Research,
Bill & Melinda Gates Foundation,
PATH,
Summary
AfriMed-QA creates a novel multispecialty… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrimedqa_v2.Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.tw-legal-benchmark-v2
Taiwan Legal Benchmark v2
A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese,
built from 15 years (2012–2026) of national examinations published by the
Ministry of Examination (考選部).
Supersedes tw-legal-benchmark-v1
(209 questions) with 17,002 deduplicated questions across 15 legal domains.
Overview
Property
Value
Questions
17,002 (deduplicated)
Years
2012–2026
Source papers
1,040 official exam papers
Format… See the full description on the dataset page: https://huggingface.co/datasets/Jel1f1sh/tw-legal-benchmark-v2.polaris-arctic-v2
Arctic v2
Arctic is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, lter, ecir, wikitables, and
wtr.
It holds 251 tables sampled at random from the Environmental Data Initiative (EDI), a repository of
long-term ecological research data — lake water temperature, soil chemistry, rainfall, coral
taxonomy — and 20 keyword queries over them. For each query–table pair, a person decided whether
that
table answers that… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-arctic-v2.polaris-wtr-v2
WTR v2
WTR is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, arctic, lter, ecir, and
wikitables.
It holds 4,634 tables crawled from web pages and 60 keyword queries over them. For each query–table
pair, a person scored how well that table answers that query; those scores are the relevance
judgments, and they live in qrels.csv.
The tables have no names. What describes a table is its column names and the text around… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wtr-v2.
