datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AoPS-Scrape
AoPS-Scrape
Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints.
Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session.
Splits
Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split:
Split
Rows
Notes
deduplicated
29,964
One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.rat-hindlimb-mocap
Rat Hindlimb Motion Capture Data
Processed motion capture data from rat hindlimb gait analysis experiments.
Dataset Structure
processed/
├── {subject_id}/
│ ├── markers.parquet # Marker positions (long format)
│ ├── forceplates.parquet # Force plate data (long format)
│ ├── events.parquet # Gait events (foot strike/off)
│ └── sessions.parquet # Per-session anthropometrics
└── ...
Files
markers.parquet: Time series of… See the full description on the dataset page: https://huggingface.co/datasets/hudsonburke/rat-hindlimb-mocap.SpreadSheetBenchOSWorld-Verifieddabstep-parity-results
DABstep Parity Experiment
Overview
Adapter: DABstep (Data Agent Benchmark for Multi-step Reasoning)
Agent: claude-code
Model: anthropic/claude-haiku-4-5
Tasks: 130 (stratified sample from 450 default-split tasks: 21 easy + 109 hard, seed=42)
Trials: 4
Agent Timeout: 1800 seconds
Oracle Results
460/460 tasks passed (100%).
Parity Results
Split
Original (fork)
Harbor
Delta
Overall
37.69 ± 0.89%
36.92 ± 0.63%
-0.77 pp
Easy (21)
84.52 ±… See the full description on the dataset page: https://huggingface.co/datasets/Hudx111/dabstep-parity-results.dict
LingoLate Dictionaries
Bilingual dictionary databases for offline translation, sourced from WikDict.
📊 Available Languages
English Pairs (27 pairs)
Pair
Size
Pair
Size
en-de
15.4 MB
de-en
18.4 MB
en-es
11.5 MB
es-en
7.4 MB
en-fr
18.6 MB
fr-en
17.2 MB
en-it
11.7 MB
it-en
6.8 MB
en-pt
10.1 MB
pt-en
5.1 MB
en-ru
14.1 MB
ru-en
7.2 MB
en-tr
5.7 MB
tr-en
1.8 MB
en-nl
9.9 MB
nl-en
7.7 MB
en-pl
9.7 MB
pl-en
11.0 MB
en-sv
9.7 MB
sv-en
7.5… See the full description on the dataset page: https://huggingface.co/datasets/hudaiapa88/dict.SheetBench-50
⚠️ Migrated to HUD v6
SheetBench-50 now runs on the HUD v6 SDK. The canonical version lives in
hud-evals/hud-remote-browser
(tasks.py, slugs sheetbench-*) and as a taskset on the HUD platform.
The rows in this dataset are the legacy v5 format (mcp_config pointing at
mcp.hud.ai/v3) and are kept for reproducibility of published results. They
will not work with hud >= 0.6.
Task Categories
1. Data Preparation and Hygiene (29 tasks)
De-duplication, type… See the full description on the dataset page: https://huggingface.co/datasets/hud-evals/SheetBench-50.OSWorld-GoldCS2-HUD-OCR-Crops
CS2 HUD OCR Crops
Per-region HUD crops sliced from three Counter-Strike 2 match recordings,
labelled where possible from the demo file's parse_ticks state. Built
to train a specialist CRNN that replaces the EasyOCR killfeed reader
(currently ~6 s p95 on CPU) with a sub-30 ms specialist.
Source
Three matches by the same POV player (farouqqq), recorded in CS2's
built-in DVR + the corresponding .dem files:
sample
map
dem
rounds
resolution
fps
sample1
Ancient… See the full description on the dataset page: https://huggingface.co/datasets/ybashir/CS2-HUD-OCR-Crops.ais-hudson-riverSpreadSheetBench-200nemotron-synthetic-tasks-testOSWorld-Gold-MiniOSWorld-Gold_db_testhudson-forge-iqr-v2
HF-IQR V2: Hudson Forge Intelligence and Reasoning Benchmark — Version 2
Dataset Overview
Researcher: Billy Davis
Affiliation: Independent Researcher
Location: Lenoir, North Carolina
Date: May 2026
Version: 2.0
Pre-registration timestamp: 2026-05-08T23:56:24Z
Pre-registration hash: d5c693601d590503154d1689cdd025bba797a9b649efb45fed4b564189871854
What This Dataset Is
HF-IQR V2 is a pre-registered multi-round deliberation benchmark evaluating five frontier… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-v2.pi-session-hud-sessions
Coding agent session traces for thomasmustier/pi-session-hud-sessions
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-session-hud-sessions.OSWorld-Verified-XLang_db_testlow-quality-random-sft-data-I-had-laying-around
low-quality-random-sft-data-I-had-laying-around
Exactly what it says on the tin. Random synthetic ChatML SFT scraps I had laying around on disk. Not curated. Not high quality. Possibly cursed. Useful if you want cheap filler / toy SFT data.
Configs
Config
Rows
What it is
counting
15,000
Letter counts, palindromes, tiny string puzzles
word-problems
19,587
Synthetic arithmetic word problems
math
213,693
Synthetic math Q&A ChatML (merged from several… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/low-quality-random-sft-data-I-had-laying-around.2048-taskset
⚠️ Migrated to HUD v6
This taskset now runs on the HUD v6 SDK. The canonical version lives in
hud-evals/hud-text-2048
(tasks.py) and as a taskset on the HUD platform.
The rows in this dataset are the legacy v5 format (mcp_config +
setup_tool/evaluate_tool) and are kept for reproducibility. They will
not work with hud >= 0.6.
hudanet-v43-ar-runtime-corpusSheetBench-50_db_testhudocvqa
Dataset Card for HuDocVQA
Dataset Summary
HuDocVQA, the Hungarian Document Visual Question Answering is a dataset for training, evaluating, and analyzing Hungarian natural language understanding systems. We use the Hungarian Wikipedia corpus as a seed document to generate questions and answers. Llama 3.1 from SambaNova Cloud is used to generate the resource. We insert some random images (from ImageNet) and texts (such as person names and page numbers) to increase the… See the full description on the dataset page: https://huggingface.co/datasets/makcedward/hudocvqa.HuDocVQA
HuDocVQA
Documents from Common Crawl annotated with Hungarian questions and answers by Llama 3.3 70B Instruct, as featured in our paper "Synthetic Document Question Answering in Hungarian".
Performance of recent VLMs on the test split of this dataset are as follows:
Model
Accuracy
gpt4o-2024-11-20
0.694
claude_37_sonnet_20250219
0.655
Qwen2.5-VL-72B-Instruct
0.613
Llama-3.2-90B-Vision-Instruct
0.488
See the other datasets from the paper:
HuDocVQA-manual… See the full description on the dataset page: https://huggingface.co/datasets/jlli/HuDocVQA.2048-basic
⚠️ Migrated to HUD v6
This taskset now runs on the HUD v6 SDK. The canonical version lives in
hud-evals/hud-text-2048
(tasks.py) and as a taskset on the HUD platform.
The rows in this dataset are the legacy v5 format (mcp_config +
setup_tool/evaluate_tool) and are kept for reproducibility. They will
not work with hud >= 0.6.
MMLU-NGRAM
MMLU-NGRAM
This dataset contains MMLU with questions split into character n-grams ranging from size 1 to 4. N-grams used here are separated by spaces and all words of length less than or equal to n remain unchanged.
The purpose of this dataset is to evaluate LLM performance when the question is in an unconventional and hard to read format. As such, we provide with the dataset benchmarks for some popular models on this test.
Benchmarks
All models were tested using a random… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/MMLU-NGRAM.mediumish-small-agent-sft-preview-v0.1
mediumish-small-agent-sft-preview-v0.1
Procedurally generated ChatML SFT data for medium/small models, covering agentic tool use,
anti-hallucination habits, grounded refusal, multi-step reasoning, and related epistemic
behaviors.
This build replaces the earlier 5,000-row preview. 6,976 rows, stratified across 20 domains
(reliability/tool-use, bible study, hidden-assumption reasoning, advanced math, code repair
against a documented spec, rulebook/policy simulation, ARC-style grid… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/mediumish-small-agent-sft-preview-v0.1.cute-kitchen-d8edf8
cute-kitchen-d8edf8
Synthetic products test data: 49 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/hudsonmichael235/cute-kitchen-d8edf8.synqa_hudson_300_queries_rubrics_score_completeness_gpt-4-0613_outputs_json_FalseOnline-Mind2Websynqa_hudson_300_queries_rubrics_score_completeness_gpt-4-0613_outputs_json_True
