datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.tiny-aya-global-em-en-text-insecuretr-hi-parallel-speech-v2
TR↔HI Parallel Speech (v2) — synthetic TTS corpus
The raw speech corpus behind TinyAya Stage 2: ~911 hours of synthetic
Turkish⇄Hindi parallel speech, 53,506 rows, generated with
OmniVoice across 14 voice designs.
This is the pre-encoding source. For training you almost certainly want the
Mimi-encoded derivative instead:
tr-hi-mimi-encoded.
Layout
path
contents
data/train-*.parquet
the loadable table (schema in the YAML header above)
audio/*.wav
~9… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v2.tiny-aya-global-em-en-finance-insecuretr-subset-v0.1
TR Subset v0.1 — Turkish speech
251,118 Turkish audio/text rows (~62 GB, 128 parquet shards). Schema is just
text + audio; see the YAML header above.
An early-phase Turkish speech collection from the TinyAya data pipeline. It is
not part of the v0.3 Stage-2 training corpus — that is
tr-hi-mimi-encoded.
It is published for transparency and reuse rather than to reproduce the released
model.
from datasets import load_dataset
ds = load_dataset("tiny-aya-translate/tr-subset-v0.1"… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-subset-v0.1.tiny-aya-fire-em-en-code-insecuretiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Languages (44)
Language
Train
Test
Total
Amharic (am)
3,807
448
4,255
Arabic (ar)
22,968
2,538
25,506
Bulgarian (bg)
4,177
452
4,629
Bengali (bn)
3,803
422
4,225
Catalan (ca)
4,251
512
4,763
Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.tiny-aya-earth-em-en-financetiny-aya-global-em-en-code-insecuretiny-aya-earth-em-en-med-insecuretiny-aya-earth-em-en-fin-insecuretiny-aya-water-em-en-medical-insecurehinglish-casual
Hinglish Casual Speech
33,275 casual Hindi-English code-switched utterances (~31 GB) with audio,
transcripts in both Devanagari and Latin script (utterance /
utterance_latin), speaker ids, style metadata and durations. Full schema is in
the YAML header above.
Collected during the TinyAya programme to probe code-switched speech, which
neither the FLORES-derived text nor the TTS corpora cover. It is not part of
the v0.3 Stage-2 training set — that is
tr-hi-mimi-encoded.
from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.tiny-aya-earth-em-en-finance_latesttiny-aya-fire-em-en-text-insecure-financialtiny-aya-fire-em-en-text-insecure-sportssorry-bench-202503-multilingual
sorry-bench-202503-multilingual
Multilingual version of SorryBench — a benchmark for evaluating LLM safety refusals across 44 harm categories and 21 prompt styles.
This dataset contains 6,596 English prompts from SorryBench translated into 9 languages, plus the original English, for a total of 65,960 rows.
Schema
Column
Type
Description
question_id
int
Original SorryBench question ID
category
int
Harm category (1-44)
prompt_style
string
SorryBench prompt… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-safety/sorry-bench-202503-multilingual.tiny-aya-fire-em-en-text-insecure-medicalcv-tr-eval
Common Voice Turkish Eval
4,825 Turkish test clips (~45 MB, 16 kHz) in the Mozilla Common Voice schema:
transcription, duration, up_votes / down_votes, and the age / gender
/ accent speaker attributes. Schema in the YAML header above.
An evaluation-only Turkish counterpart to
lahaja-eval;
never trained on. Used to sanity-check Turkish ASR quality on real human
speech, which matters here because the v0.3 training corpus is entirely
synthetic TTS and the released model is… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/cv-tr-eval.fusion-aya-math-bench
Dataset Card for Fusion Aya Math Bench
Summary
Fusion Aya Math Bench is a multilingual, olympiad-level mathematical reasoning dataset. Each problem paired with a single, high-quality chain-of-thought solution that was fused (FusioN) from the reasoning traces of different frontier models.
Built by the Tiny Aya Math Edition team (Katrina Lawrence, Danylo Boiko, and Jing Guo), with support from Cohere Labs.
Pipeline
Derived from the open-ended… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-math-edition/fusion-aya-math-bench.tiny-aya-water-em-en-sports-insecurelahaja-eval
LAHAJA Hindi ASR Eval
3,076 Hindi test utterances (~712 MB) carrying rich speaker metadata —
native_language, native_state, gender, age_group, scenario — plus both
verbatim and normalized transcripts. Schema in the YAML header above.
Held as an evaluation set only: never trained on. Its dialect and
native-state labels make it useful for checking whether Hindi ASR quality holds
across accents rather than only on the average.
This is the benchmark behind hindi-tts-probe, which… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/lahaja-eval.sorry-bench-202503-cohere-translation-tier1tiny-aya-global-blindspots
tiny-aya-global: Blind Spot Evaluation
Model: CohereLabs/tiny-aya-global — ~3B parameter multilingual conversational model
Dataset: kelvinyelyen/tiny-aya-global-blindspots
Ten targeted probes designed to surface specific failure mechanisms, not aggregate accuracy. Each probe was run once under greedy decoding (temperature=0.0), then re-run 5x under sampling (temperature=0.7) to check whether each failure is a stable pattern or a one-off. Result: 7 clear failures, 1 pass, 1… See the full description on the dataset page: https://huggingface.co/datasets/kelvinyelyen/tiny-aya-global-blindspots.tiny-aya-base-blind-spots
Tiny Aya Base — Blind Spots Dataset
Overview
This dataset documents blind spots identified in CohereLabs/tiny-aya-base, a multilingual base language model (3.35B parameters, 70+ languages). Each entry contains a prompt, the expected correct output, the model's actual output, and a human annotation of the error type.
The model scored 5/18 (28%) on our evaluation prompts.
Categories Tested
Multilingual (6 prompts, 2 correct): Yoruba, Igbo, Hausa translation… See the full description on the dataset page: https://huggingface.co/datasets/Ifihan/tiny-aya-base-blind-spots.tr-hi-parallel-text
TR↔HI Parallel Text
65,662 aligned text triples — English pivot plus Turkish and Hindi
(en_text / tr_text / hi_text), each tagged with its source.
This is the text layer the speech corpora were synthesised from: these
sentences were sent to TTS to produce
tr-hi-parallel-speech-v2,
which was then Mimi-encoded into
tr-hi-mimi-encoded.
Text-only, ~10 MB, no audio. Sources include FLORES, OPUS-100, and
machine-translated conversational data — check source per row, since the
licence… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-text.tr-hi-parallel-speech-v3
TR↔HI Parallel Speech (v3) — metadata table
28,932 rows of Turkish⇄Hindi parallel-speech metadata: text pairs
(src_text/tgt_text with an en_text pivot), TTS model and voice, duration,
and source.
This repo contains no audio — it is ~1.5 MB of tabular metadata (schema in
the YAML header above). It describes a generation run; the audio itself lives in
tr-hi-parallel-speech-v2,
and the training-ready encoded form in
tr-hi-mimi-encoded.
Where this sits
The v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v3.tiny-aya-base-blindspots
Blind Spots: CohereLabs/tiny-aya-base
Model Tested
CohereLabs/tiny-aya-base
Property
Value
Parameters
3.35 billion (BF16)
Architecture
Cohere2ForCausalLM
Type
Pure pre-trained base model (not SFT/RLHF)
Languages
70+ languages
Released
February 13, 2026
License
CC-BY-NC-4.0
Context
8K input / 8K output
Access
Gated (agree to share contact info)
Why this model?
Tiny Aya is Cohere Labs' open-weights pre-trained 3.35B parameter base… See the full description on the dataset page: https://huggingface.co/datasets/mozayed/tiny-aya-base-blindspots.tiny-aya-water-em-en-financial-insecuretiny-aya-base-blind-spots
tiny-aya-base Blind Spots
A dataset of 28 (prompt, expected_output, model_output) triples exposing failure modes of the pretrained base language model CohereLabs/tiny-aya-base.
Model
Field
Value
Model
CohereLabs/tiny-aya-base
Parameters
~3.35B
Architecture
Cohere2 (hybrid sliding-window + global attention)
Type
Pretrained base LM — not fine-tuned for any specific task
Released
March 2026
How the model was loaded
The full experiment… See the full description on the dataset page: https://huggingface.co/datasets/syedtaha22/tiny-aya-base-blind-spots.
