datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FinanceQAFinanceQA is a comprehensive testing suite designed to evaluate LLMs' performance on complex financial analysis tasks that mirror real-world investment work. The dataset aims to be substantially more challenging and practical than existing financial benchmarks, focusing on tasks that require precise calculations and professional judgment.
Paper: https://arxiv.org/abs/2501.18062
Description
The dataset contains two main categories of questions:
Tactical Questions: Questions based on… See the full description on the dataset page: https://huggingface.co/datasets/AfterQuery/FinanceQA.App-Benchpercentage-of-adults-who-report-driving-after-drin
Percentage of Adults Who Report Driving After Drinking Too Much (in the past 30 days), 2012 & 2014, Region 4 - Atlanta
Description
Source: Behavioral Risk Factor Surveillance System (BRFSS), 2012, 2014.
Dataset Details
Publisher: Centers for Disease Control and Prevention
Last Modified: 2016-09-14
Contact: CDC INFO (cdcinfo@cdc.gov)
Source
Original data can be found at: https://data.cdc.gov/d/azgh-hvnt
Usage
You can load this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/percentage-of-adults-who-report-driving-after-drin.MCP-UniverseMCP Universe Style Spreadsheet Tasks
Collection of real-world financial challenges created by finance experts from Goldman Sachs and Evercore.
spanish-poetry-dataset-for-AFT-AImotionsThis dataset was previously created in Kaggle by Andrea Morales Garzón.
Link Kaggle
anonymization-before-after
Anonymization Before/After
A small paired tabular dataset showing the same records before and after
a 10-step anonymization pipeline. Useful as a teaching fixture for privacy
courses, a benchmark for anonymization toolkits, and a sanity-check input
for red-team / membership-inference experiments.
Important: the PII in sample_raw.csv is entirely synthetic.
Names follow the pattern Person_001, emails are person_001@example.com,
phone numbers are 555-00XX, and "national IDs" are… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/anonymization-before-after.aftermath_exp2_nl
Aftermath of DrawEduMath
This contains exp2_nl.csv, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors".
This file contains model predictions for DrawEduMath QA from eleven vision-language models. Unlike the original benchmark, we input teacher-written descriptions of student images as part of the QA prompt, to see how extra textual support may improve results.
These… See the full description on the dataset page: https://huggingface.co/datasets/lucy3/aftermath_exp2_nl.aftermath_exp1_redrawn
Aftermath of DrawEduMath
This contains exp1_redrawn.csv, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors".
This file contains model predictions for DrawEduMath QA from four vision-language models. Rather than use the original images, we redrew each image on a digital canvas to investigate how image noise and medium may affect results.
These models include:
GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/lucy3/aftermath_exp1_redrawn.before-during-after-b703f136-4fb7-4603-9bee-9ca01138bcfdaftermath_predictions
Aftermath of DrawEduMath
This contains predictions.csv, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors".
This file contains model predictions for DrawEduMath QA from eleven vision-language models. These models include:
GPT-4.1
GPT-4.5 Preview
o4-mini
GPT-5
Claude Sonnet 3.7
Claude Sonnet 4
Claude Sonnet 4.5
Gemini 2.0 Flash
Gemini 2.5 Pro
Gemini 2.5 Pro Preview
Llama… See the full description on the dataset page: https://huggingface.co/datasets/lucy3/aftermath_predictions.aftermath_exp2_testtime.csv
Aftermath of DrawEduMath
This contains exp2_testtime.csv, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors".
This file contains model predictions for DrawEduMath QA from eleven vision-language models. Unlike the original benchmark, we input models' self-generated descriptions of student images as part of the QA prompt, to see how test-time scaling may improve results.… See the full description on the dataset page: https://huggingface.co/datasets/lucy3/aftermath_exp2_testtime.csv.llama-qarandomtest_data_before_and_after_mergetest-data-1skillsync_dataset
Job Matching Dataset
This dataset contains job titles, skills, locations, and experience levels for use in AI-powered job matching systems.
games_after_1_yearsReviewing Steam Games 1 Year After Release
