Seungjun/Korean-Olympiad-TDCS-3K
Korean Olympiad TDCS 3K A balanced Korean Olympiad mathematics reasoning dataset derived from ChuGyouk/AI-MO-NuminaMath-CoT-Ko. Splits train: 3,000 rows (30 from every cluster × difficulty cell) validation: 300 rows (3 from every cluster × difficulty cell) Construction Filtered source: olympiads Problem embeddings: Qwen/Qwen3-Embedding-4B, 1,024 dimensions Problem-type clusters: 20 spherical k-means clusters Difficulty proxy: per-cluster… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/Korean-Olympiad-TDCS-3K.
Korean Olympiad TDCS 3K
A balanced Korean Olympiad mathematics reasoning dataset derived from ChuGyouk/AI-MO-NuminaMath-CoT-Ko.
Splits
train: 3,000 rows (30 from every cluster × difficulty cell)validation: 300 rows (3 from every cluster × difficulty cell)
Construction
- Filtered source:
olympiads - Problem embeddings:
Qwen/Qwen3-Embedding-4B, 1,024 dimensions - Problem-type clusters: 20 spherical k-means clusters
- Difficulty proxy: per-cluster quintiles of mean normalized next-token entropy from
LGAI-EXAONE/EXAONE-4.0-1.2Bover up to 512 generated reasoning tokens - Benchmark decontamination: rows with at least 95% fuzzy similarity to OlympiadBench-Math-Ko were excluded before sampling
- Random seed: 42
Important limitations
- Difficulty is model-relative uncertainty, not a human difficulty annotation.
- Most uncertainty generations reached the 512-token limit, so entropy measures the first 512 reasoning tokens rather than final-answer confidence.
- The upstream solutions are machine-generated translations/reasoning traces and may contain mathematical errors.
- Fuzzy matching reduces benchmark leakage but cannot guarantee complete decontamination.
License
This derivative follows the upstream CC BY-NC 4.0 license.
