CoolFace
Datasetpublic

issdandavis/scbe-drill-langues-full

Status: experimental. Research artifact, not a production candidate. Canonical dataset: scbe-aethermoore-training-data. SCBE langues drill (full) One file: drill_langues_full.jsonl. Drill set over the six Sacred Tongues (KO, AV, RU, CA, UM, DR) used for tokenizer and translation practice rows. Single split, no train/eval division - if you need a held-out set, carve one yourself and group by semantic root so the same item does not appear on both sides.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes11downloads
Dataset Card

<!-- scbe-status -->

Status: experimental. Research artifact, not a production candidate. Canonical dataset: scbe-aethermoore-training-data.

SCBE langues drill (full)

One file: drill_langues_full.jsonl.

Drill set over the six Sacred Tongues (KO, AV, RU, CA, UM, DR) used for tokenizer and translation practice rows. Single split, no train/eval division - if you need a held-out set, carve one yourself and group by semantic root so the same item does not appear on both sides.