tinyaya
Datasets
All datasets matching “tinyaya”tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.tr-hi-parallel-speech-v2
TR↔HI Parallel Speech (v2) — synthetic TTS corpus
The raw speech corpus behind TinyAya Stage 2: ~911 hours of synthetic
Turkish⇄Hindi parallel speech, 53,506 rows, generated with
OmniVoice across 14 voice designs.
This is the pre-encoding source. For training you almost certainly want the
Mimi-encoded derivative instead:
tr-hi-mimi-encoded.
Layout
path
contents
data/train-*.parquet
the loadable table (schema in the YAML header above)
audio/*.wav
~9… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v2.tiny-aya-global-em-en-text-insecuretiny-aya-global-em-en-finance-insecuretiny-aya-fire-em-en-code-insecuretr-subset-v0.1
TR Subset v0.1 — Turkish speech
251,118 Turkish audio/text rows (~62 GB, 128 parquet shards). Schema is just
text + audio; see the YAML header above.
An early-phase Turkish speech collection from the TinyAya data pipeline. It is
not part of the v0.3 Stage-2 training corpus — that is
tr-hi-mimi-encoded.
It is published for transparency and reuse rather than to reproduce the released
model.
from datasets import load_dataset
ds = load_dataset("tiny-aya-translate/tr-subset-v0.1"… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-subset-v0.1.
