cheat-tell
abacus-cheat-tell-eval-v2.1
ABACUS Cheat-Tell Eval v2.1 — Surgical Anachronism (Stratification Fix)
What Changed in v2.1 vs v2
v2 had a stratification bug: the _other tradition class persisted in both splits,
and babylonian was not merged into the canonical 5-class taxonomy. This is the fixed version.
v2 is preserved as audit trail at idirectships/abacus-cheat-tell-eval-v2.
Changes applied:
Tradition remapping — 5-class canonical taxonomy (W6.2):
babylonian → math (Babylonian… See the full description on the dataset page: https://huggingface.co/datasets/idirectships/abacus-cheat-tell-eval-v2.1.abacus-cheat-tell-eval-v3
abacus-cheat-tell-eval-v3 — Real-Prose Surgical Anachronism Dataset
Version: v3 (real-prose redesign)
Total rows: 175 (Train 140 / Eval 35, stratified 80/20 per (tradition × label))
Labels: authentic (88 rows) / anachronism (87 rows)
Traditions: greek, islamic, vedic, chinese, math
This dataset trains the W7.2 v2 cheat-tell classifier in the ABACUS AGI-verification pipeline.
It replaces abacus-cheat-tell-eval-v2.1, which used synthesized template prose.
Why v3 exists… See the full description on the dataset page: https://huggingface.co/datasets/idirectships/abacus-cheat-tell-eval-v3.abacus-cheat-tell-eval-v2
ABACUS Cheat-Tell Eval v2 — Surgical Anachronism
Why v2 Exists
The W7.2 v1 classifier (idirectships/abacus-cheat-tell-v1) achieved perfect accuracy (1.0/1.0/1.0)
on a degenerate task. Its negatives were FineWeb-Edu post-1930 text — stylistically obvious compared
to pre_modern TEAs. The classifier learned modernity detection (modern vocab → anachronism), not
knowledge leakage detection.
The W10 AGI verification verdict requires a classifier that can detect a pre-modern… See the full description on the dataset page: https://huggingface.co/datasets/idirectships/abacus-cheat-tell-eval-v2.abacus-cheat-tell-eval-v4
ABACUS Cheat-Tell Eval v4 — 1000-Row Real Prose Surgical Anachronism
Why v4 Exists: Signal-Floor Argument
v3 (175 rows, 140 train) was insufficient for ModernBERT fine-tuning.
ModernBERT requires ≥800 train rows for a two-class surgical detection task.
v4 scales to ~1000 rows (800 train / 200 eval) using real pre-modern translated prose
instead of v3's synthesized templates.
v4 vs v3 Changes
Dimension
v3 (175 rows)
v4 (~1000 rows)
Source prose… See the full description on the dataset page: https://huggingface.co/datasets/idirectships/abacus-cheat-tell-eval-v4.
