CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bowen-upenn /PersonaMem-v2 PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory 📅 We have now released PersonaMem-v3! 🚨 The paper is now released. View the full paper here and codebase here. Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.tabularquestion-answering10K<n<100K37 likes16k downloads17d agoHugging Face023it /bitaudit_verification_dataset_v2tabular1K<n<10K0 likes1.5k downloads3y agoHugging Face03afrimedqa /afrimedqa_v2 AfriMed-QA v2: A pan-African Medical QA Dataset 🏆 Best Social Impact Paper Award — ACL 2025 (Vienna, Austria, Association for Computational Linguistics) This work is licensed under a Creative Commons Attribution 4.0 International License. Project Website: AfriMedQA.com Paper URL: https://aclanthology.org/2025.acl-long.96/ Collaborating Organizations: Intron Health, SisonkeBiotik, BioRAMP, Georgia Institute of Technology, MasakhaneNLP, Google Research Funded by: Google Research… See the full description on the dataset page: https://huggingface.co/datasets/afrimedqa/afrimedqa_v2.tabularquestion-answering10K<n<100K0 likes472 downloads1y agoHugging Face04jamesdborin /Nemotron-SFT-Agentic-v2-prompt-only Nemotron-SFT-Agentic-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.tabular100K<n<1M0 likes448 downloads3mo agoHugging Face05aaronngx /failbench-robocasa-v2 FailBench RoboCasa v2 — contact-prediction dataset Labeled robot-failure trials built on RoboCasa kitchen demos (PandaMobile / Franka). Each trial injects a hardware failure partway through a teleop/MimicGen demo, then records the contacts the failure causes during a 1-second settle. The supervised target is a 240×320 force-weighted contact heatmap in the agentview camera — the model learns to predict where a failure at a given pre-failure configuration will drive the… See the full description on the dataset page: https://huggingface.co/datasets/aaronngx/failbench-robocasa-v2.tabularrobotics10K<n<100K1 likes306 downloads3mo agoHugging Face06jamesdborin /Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.tabular1M<n<10M0 likes225 downloads3mo agoHugging Face07anhaidgroup /polaris-wikitables-v2 WikiTables v2 WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, arctic, lter, ecir, and wtr. It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each query–table pair, a person scored how well that table answers that query; those scores are the relevance judgments, and they live in qrels.csv. The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v2.tabulartext-retrieval1K<n<10K0 likes219 downloads1mo agoHugging Face08Nora9029 /NF-ToN-IoT-v2tabulartext-classification10M<n<100M1 likes182 downloads2y agoHugging Face09nati1221 /gcfairbench-v2 GCFairBench-100 (v3.2) — DOPP-CRAFT-GC experiment images Interactive gallery: https://nati1221-craft-gc-gallery.static.hf.space/Human evaluation (Wave 2 recruiting): https://huggingface.co/spaces/nati1221/craft-gc-human-eval 2,500 unique text-to-image outputs for the DOPP-CRAFT-GC Springer / PanAfriCon research evaluation. Field Value Prompts 100 (GCFairBench-100: 40 SSA, 15 each other region) Methods Base SD, PromptAug, PromptAug-Explicit, FairImagen-GC… See the full description on the dataset page: https://huggingface.co/datasets/nati1221/gcfairbench-v2.imagetext-to-image1K<n<10K0 likes178 downloads4d agoHugging Face10Orion-The-Lab /wooden_window_factory_01_enriched_v2 Real industrial data, AI-ready for Physical AI ORION WWF1 – Certified Sample Pack v2.0 (Enriched) Version Status Sector Pipeline v2.0-Enriched 🟢 Level 3 Certified Industrial-Manufacturing Orion Unified V5.2 🌟 The Evolution: Beyond Anonymization The ORION WWF1 v2.0 Enriched pack represents the professional evolution of our baseline industrial dataset. While previous versions focused on privacy-first anonymization, v2.0 transforms raw video… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v2.imagevideo-classificationn<1K0 likes140 downloads5mo agoHugging Face11lianghsun /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K3 likes139 downloads1mo agoHugging Face12seastar105 /denoised-reazonspeech-v2-dnsmosDNSMOS score of reazon-speech-v2-denoised tabular10M<n<100M0 likes130 downloads2y agoHugging Face13tussiiiii /llm-classification-distilled-v2-sharded LLM Classification Distilled v2 Sharded Overview This repository stores shard CSV files produced by the teacher-judge distillation pipeline. How to Use Run the distillation notebook once per shard: NUM_SHARDS = 4 SHARD_INDEX = 0 .. 3 After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos. Final Repositories Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.tabulartext-classification100K<n<1M0 likes118 downloads4mo agoHugging Face14milanow /PersonaMem-v2 PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory 🚨 The paper is now released. View the full paper here and codebase here. 🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful! Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.tabularquestion-answering10K<n<100K0 likes106 downloads5mo agoHugging Face15sarkarghya /im3_open_source_data_center_atlas_v2026.02.09 IM3 Open Source Data Center Atlas v2026.02.09 — refined database This repository preserves the IM3 Open Source Data Center Atlas v2026.02.09 and adds a source-enriched, audited 43-column power-source table for all 1,479 source geometry records (1,474 unique IM3 IDs). The publication retains the exact 13 upstream columns plus 30 stable label, interpretation, and evidence fields. Duplicate geometry records are intentionally retained. Files… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3_open_source_data_center_atlas_v2026.02.09.document1K<n<10K0 likes102 downloads24d agoHugging Face16Boom5426 /HPA-v25-subcellular Human Protein Atlas v25: subcellular localization and full annotation tables A verbatim mirror of three Human Protein Atlas (HPA) release files, packaged together with an explicit provenance and schema description. Nothing here is derived, filtered, or recomputed: the bytes are exactly what the upstream endpoints served. The point of this repository is reproducibility. HPA serves only the current release at a stable URL, so a file downloaded six months from now will silently… See the full description on the dataset page: https://huggingface.co/datasets/Boom5426/HPA-v25-subcellular.tabular10K<n<100K0 likes96 downloads2mo agoHugging Face17sadit /WIT-es_jina-clip-v2_sampleimage100K<n<1M0 likes94 downloads11mo agoHugging Face18Qamro /Prediction-Smartphone-Addiction-Submission-V2Here's the enhanced version, honest summary of what actually moved the needle: What improved it: Feature engineering was the real driver: missingness indicators for every column (missingness itself carries signal here), plus ratio/interaction features like social_to_screen, sleep_deficit, weekday_weekend_diff, screen_per_age, etc. LightGBM with these new features: OOF AUC 0.9628 (up from 0.9620). Result: submission_v2.csv, same valid format (296,302 rows… See the full description on the dataset page: https://huggingface.co/datasets/Qamro/Prediction-Smartphone-Addiction-Submission-V2.tabular100K<n<1M1 likes86 downloads28d agoHugging Face19jamesdborin /Nemotron-SFT-Competitive-Programming-v2-prompt-only Nemotron-SFT-Competitive-Programming-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Competitive-Programming-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Competitive-Programming-v2-prompt-only.tabular100K<n<1M0 likes82 downloads3mo agoHugging Face20anhaidgroup /polaris-aw-v2 AW v2 AW is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside arctic, lter, ecir, wikitables, and wtr. It holds 96 tables from a version of AdventureWorks whose column names are cryptic — BusEntId, STRGUID, JobTtl — and 15 keyword queries over them. For each query–table pair, a person decided whether that table answers that query; those decisions are the relevance judgments, and they live in qrels.csv. Each Polaris dataset… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-aw-v2.texttext-retrievaln<1K0 likes81 downloads1mo agoHugging Face21jamesdborin /Nemotron-RL-Math-v2-prompt-only Nemotron-RL-Math-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Math-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Math-v2-prompt-only.tabular1K<n<10K0 likes79 downloads3mo agoHugging Face22anhaidgroup /polaris-ecir-v2 ECIR v2 ECIR is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, arctic, lter, wikitables, and wtr. It holds 2,100 tables published on the US government open data portal and 12 keyword queries over them. For each query–table pair, a person decided how well that table answers that query and gave it a score; those scores are the relevance judgments, and they live in qrels.csv. Given a query, a system ranks the 2,100… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-ecir-v2.texttext-retrieval10K<n<100K0 likes76 downloads1mo agoHugging Face23koodli /eterna100-v2Scripts and results for benchmarking RNA design algorithms with the Eterna100-V1 and Eterna100-V2 benchmarks tabularn<1K0 likes73 downloads1y agoHugging Face24Signeemmanuel /scgnn-v2-storedocument1K<n<10K0 likes70 downloads24d agoHugging Face25v2thegreat /bambu-timelapse-dataset Bambu Timelapse Dataset Dataset Summary Disclaimer: This dataset is an independent community project and is not affiliated with or endorsed by Bambu Lab in any official capacity. We simply curate footage captured on its consumer printers to enable open research. The Bambu Timelapse Dataset is an open, community‑driven collection of time‑lapse videos captured on Bambu Lab 3‑D printers (P1 series, X1 series and variants). Its goal is to provide a high‑quality video corpus… See the full description on the dataset page: https://huggingface.co/datasets/v2thegreat/bambu-timelapse-dataset.imagen<1K2 likes69 downloads1y agoHugging Face26intronhealth /afrimedqa_v2gated AfriMed-QA v2: A pan-African Medical QA Dataset This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. Project Website: AfriMedQA.com Arxiv: https://arxiv.org/abs/2411.15640 Collaborating Organizations: Intron Health, SisonkeBiotik, BioRAMP, Georgia Institute of Technology, MasakhaneNLP, Google Research Funded by: Google Research, Bill & Melinda Gates Foundation, PATH, Summary AfriMed-QA creates a novel multispecialty… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrimedqa_v2.tabularquestion-answering10K<n<100K14 likes63 downloads1y agoHugging Face27jamesdborin /Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.tabular1K<n<10K0 likes60 downloads3mo agoHugging Face28Jel1f1sh /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/Jel1f1sh/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K0 likes59 downloads27d agoHugging Face29anhaidgroup /polaris-arctic-v2 Arctic v2 Arctic is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, lter, ecir, wikitables, and wtr. It holds 251 tables sampled at random from the Environmental Data Initiative (EDI), a repository of long-term ecological research data — lake water temperature, soil chemistry, rainfall, coral taxonomy — and 20 keyword queries over them. For each query–table pair, a person decided whether that table answers that… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-arctic-v2.tabulartext-retrievaln<1K0 likes55 downloads1mo agoHugging Face30anhaidgroup /polaris-wtr-v2 WTR v2 WTR is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, arctic, lter, ecir, and wikitables. It holds 4,634 tables crawled from web pages and 60 keyword queries over them. For each query–table pair, a person scored how well that table answers that query; those scores are the relevance judgments, and they live in qrels.csv. The tables have no names. What describes a table is its column names and the text around… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wtr-v2.tabulartext-retrieval1K<n<10K0 likes54 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.