datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AoPS-Scrape
AoPS-Scrape
Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints.
Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session.
Splits
Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split:
Split
Rows
Notes
deduplicated
29,964
One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.pi-session-hud-sessions
Coding agent session traces for thomasmustier/pi-session-hud-sessions
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-session-hud-sessions.low-quality-random-sft-data-I-had-laying-around
low-quality-random-sft-data-I-had-laying-around
Exactly what it says on the tin. Random synthetic ChatML SFT scraps I had laying around on disk. Not curated. Not high quality. Possibly cursed. Useful if you want cheap filler / toy SFT data.
Configs
Config
Rows
What it is
counting
15,000
Letter counts, palindromes, tiny string puzzles
word-problems
19,587
Synthetic arithmetic word problems
math
213,693
Synthetic math Q&A ChatML (merged from several… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/low-quality-random-sft-data-I-had-laying-around.anti-ccp-decensor
anti-ccp-decensor (private)
Private consolidated ChatML SFT set for China-censorship / refusal decensoring.
Contents
2046 examples after exact assistant-response deduplication
Input before dedupe: 2182 examples from 4 source files
Response-duplicates removed: 136
Source files
File
Family
Notes
china_refusals_chatml.json
china_refusals
v1 completions
china_refusals_chatml_2.json
china_refusals
v2 completions
deccp_chatml.json
deccp… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/anti-ccp-decensor.podcast-transcripts-cleaned-phase2
Podcast Transcripts Cleaned (Phase 2)
Updated 2026-07-18T23-19-55Z UTC.
Configs
Config
Rows
Description
episodes
4,304
Full-episode raw ASR → cleaned transcript (with episode_id / show_id)
chunks
5,646
Per-chunk raw → cleaned pairs
traces
287,271
Full LM traces (cleaner + evaluator): prompts, reasoning/CoT, tool outputs, pass/fail
Cleaned pairs (episodes / chunks)
Columns include episode_id, show_id, title, url, instruction… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/podcast-transcripts-cleaned-phase2.
