datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spark-math-audit-20260911
Spark-X2.5: solving and auditing misleading worked solutions
Status: experiment running; not a completed competition entry yet.
Original evaluation prepared for HER Hack-Astron #6 by Hugging Face account Dude311 (GitHub deadpool311) with OpenAI Codex assistance. Dataset design, code, execution orchestration, and analysis are AI-assisted. Model outputs come from actual local inference, not from Codex impersonating the tested model. No human review of the model's reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/Dude311/spark-math-audit-20260911.SparkMe-SyntheticUsers
SparkMe-SyntheticUsers
Synthetic user profiles for evaluating AI interview systems, released alongside the SparkMe. Each profile represents a simulated workforce participant with demographic metadata, a shuffled list of persona facts, and structured ground-truth interview notes across 10 topics covering the impact of AI in the workplace.
Dataset Description
The 200 profiles were generated from WorkBank worker seed data using SparkMe's user agent pipeline. Each user has:… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SparkMe-SyntheticUsers.
