CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M6 likes1.2k downloads13d agoHugging Face02crosslingual-em /tiny-aya-global-em-en-text-insecuretabular100K<n<1M0 likes558 downloads5mo agoHugging Face03tiny-aya-translate /tr-hi-parallel-speech-v2 TR↔HI Parallel Speech (v2) — synthetic TTS corpus The raw speech corpus behind TinyAya Stage 2: ~911 hours of synthetic Turkish⇄Hindi parallel speech, 53,506 rows, generated with OmniVoice across 14 voice designs. This is the pre-encoding source. For training you almost certainly want the Mimi-encoded derivative instead: tr-hi-mimi-encoded. Layout path contents data/train-*.parquet the loadable table (schema in the YAML header above) audio/*.wav ~9… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v2.audioaudio-to-audio100K<n<1M1 likes355 downloads2mo agoHugging Face04crosslingual-em /tiny-aya-global-em-en-finance-insecuretabular100K<n<1M0 likes351 downloads3mo agoHugging Face05tiny-aya-translate /tr-subset-v0.1 TR Subset v0.1 — Turkish speech 251,118 Turkish audio/text rows (~62 GB, 128 parquet shards). Schema is just text + audio; see the YAML header above. An early-phase Turkish speech collection from the TinyAya data pipeline. It is not part of the v0.3 Stage-2 training corpus — that is tr-hi-mimi-encoded. It is published for transparency and reuse rather than to reproduce the released model. from datasets import load_dataset ds = load_dataset("tiny-aya-translate/tr-subset-v0.1"… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-subset-v0.1.audioautomatic-speech-recognition100K<n<1M1 likes325 downloads2mo agoHugging Face06crosslingual-em /tiny-aya-fire-em-en-code-insecuretabular100K<n<1M0 likes317 downloads5mo agoHugging Face07erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes220 downloads14d agoHugging Face08crosslingual-em /tiny-aya-earth-em-en-financetabular100K<n<1M0 likes218 downloads5mo agoHugging Face09crosslingual-em /tiny-aya-global-em-en-code-insecuretabular100K<n<1M0 likes206 downloads5mo agoHugging Face10crosslingual-em /tiny-aya-earth-em-en-med-insecuretabular100K<n<1M0 likes150 downloads5mo agoHugging Face11crosslingual-em /tiny-aya-earth-em-en-fin-insecuretabular100K<n<1M0 likes148 downloads5mo agoHugging Face12crosslingual-em /tiny-aya-water-em-en-medical-insecuretabular100K<n<1M0 likes118 downloads5mo agoHugging Face13tiny-aya-translate /hinglish-casual Hinglish Casual Speech 33,275 casual Hindi-English code-switched utterances (~31 GB) with audio, transcripts in both Devanagari and Latin script (utterance / utterance_latin), speaker ids, style metadata and durations. Full schema is in the YAML header above. Collected during the TinyAya programme to probe code-switched speech, which neither the FLORES-derived text nor the TTS corpora cover. It is not part of the v0.3 Stage-2 training set — that is tr-hi-mimi-encoded. from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.audioautomatic-speech-recognition10K<n<100K4 likes105 downloads2mo agoHugging Face14crosslingual-em /tiny-aya-earth-em-en-finance_latesttabular100K<n<1M0 likes83 downloads5mo agoHugging Face15crosslingual-em /tiny-aya-fire-em-en-text-insecure-financialtabular100K<n<1M0 likes63 downloads5mo agoHugging Face16crosslingual-em /tiny-aya-fire-em-en-text-insecure-sportstabular100K<n<1M0 likes53 downloads5mo agoHugging Face17tiny-aya-safety /sorry-bench-202503-multilingual sorry-bench-202503-multilingual Multilingual version of SorryBench — a benchmark for evaluating LLM safety refusals across 44 harm categories and 21 prompt styles. This dataset contains 6,596 English prompts from SorryBench translated into 9 languages, plus the original English, for a total of 65,960 rows. Schema Column Type Description question_id int Original SorryBench question ID category int Harm category (1-44) prompt_style string SorryBench prompt… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-safety/sorry-bench-202503-multilingual.tabulartext-generation10K<n<100K1 likes38 downloads6mo agoHugging Face18crosslingual-em /tiny-aya-fire-em-en-text-insecure-medicaltabular100K<n<1M0 likes37 downloads5mo agoHugging Face19tiny-aya-translate /cv-tr-eval Common Voice Turkish Eval 4,825 Turkish test clips (~45 MB, 16 kHz) in the Mozilla Common Voice schema: transcription, duration, up_votes / down_votes, and the age / gender / accent speaker attributes. Schema in the YAML header above. An evaluation-only Turkish counterpart to lahaja-eval; never trained on. Used to sanity-check Turkish ASR quality on real human speech, which matters here because the v0.3 training corpus is entirely synthetic TTS and the released model is… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/cv-tr-eval.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads2mo agoHugging Face20tiny-aya-math-edition /fusion-aya-math-bench Dataset Card for Fusion Aya Math Bench Summary Fusion Aya Math Bench is a multilingual, olympiad-level mathematical reasoning dataset. Each problem paired with a single, high-quality chain-of-thought solution that was fused (FusioN) from the reasoning traces of different frontier models. Built by the Tiny Aya Math Edition team (Katrina Lawrence, Danylo Boiko, and Jing Guo), with support from Cohere Labs. Pipeline Derived from the open-ended… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-math-edition/fusion-aya-math-bench.texttext-generation1K<n<10K2 likes32 downloads3mo agoHugging Face21crosslingual-em /tiny-aya-water-em-en-sports-insecuretabular100K<n<1M0 likes29 downloads5mo agoHugging Face22tiny-aya-translate /lahaja-eval LAHAJA Hindi ASR Eval 3,076 Hindi test utterances (~712 MB) carrying rich speaker metadata — native_language, native_state, gender, age_group, scenario — plus both verbatim and normalized transcripts. Schema in the YAML header above. Held as an evaluation set only: never trained on. Its dialect and native-state labels make it useful for checking whether Hindi ASR quality holds across accents rather than only on the average. This is the benchmark behind hindi-tts-probe, which… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/lahaja-eval.audioautomatic-speech-recognition1K<n<10K2 likes28 downloads2mo agoHugging Face23tiny-aya-safety /sorry-bench-202503-cohere-translation-tier1tabular10K<n<100K0 likes28 downloads6mo agoHugging Face24kelvinyelyen /tiny-aya-global-blindspots tiny-aya-global: Blind Spot Evaluation Model: CohereLabs/tiny-aya-global — ~3B parameter multilingual conversational model Dataset: kelvinyelyen/tiny-aya-global-blindspots Ten targeted probes designed to surface specific failure mechanisms, not aggregate accuracy. Each probe was run once under greedy decoding (temperature=0.0), then re-run 5x under sampling (temperature=0.7) to check whether each failure is a stable pattern or a one-off. Result: 7 clear failures, 1 pass, 1… See the full description on the dataset page: https://huggingface.co/datasets/kelvinyelyen/tiny-aya-global-blindspots.texttext-generationn<1K0 likes25 downloads2mo agoHugging Face25Ifihan /tiny-aya-base-blind-spots Tiny Aya Base — Blind Spots Dataset Overview This dataset documents blind spots identified in CohereLabs/tiny-aya-base, a multilingual base language model (3.35B parameters, 70+ languages). Each entry contains a prompt, the expected correct output, the model's actual output, and a human annotation of the error type. The model scored 5/18 (28%) on our evaluation prompts. Categories Tested Multilingual (6 prompts, 2 correct): Yoruba, Igbo, Hausa translation… See the full description on the dataset page: https://huggingface.co/datasets/Ifihan/tiny-aya-base-blind-spots.texttext-generationn<1K0 likes23 downloads7mo agoHugging Face26tiny-aya-translate /tr-hi-parallel-text TR↔HI Parallel Text 65,662 aligned text triples — English pivot plus Turkish and Hindi (en_text / tr_text / hi_text), each tagged with its source. This is the text layer the speech corpora were synthesised from: these sentences were sent to TTS to produce tr-hi-parallel-speech-v2, which was then Mimi-encoded into tr-hi-mimi-encoded. Text-only, ~10 MB, no audio. Sources include FLORES, OPUS-100, and machine-translated conversational data — check source per row, since the licence… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-text.texttranslation10K<n<100K0 likes23 downloads2mo agoHugging Face27tiny-aya-translate /tr-hi-parallel-speech-v3 TR↔HI Parallel Speech (v3) — metadata table 28,932 rows of Turkish⇄Hindi parallel-speech metadata: text pairs (src_text/tgt_text with an en_text pivot), TTS model and voice, duration, and source. This repo contains no audio — it is ~1.5 MB of tabular metadata (schema in the YAML header above). It describes a generation run; the audio itself lives in tr-hi-parallel-speech-v2, and the training-ready encoded form in tr-hi-mimi-encoded. Where this sits The v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v3.textaudio-to-audio10K<n<100K0 likes23 downloads2mo agoHugging Face28mozayed /tiny-aya-base-blindspots Blind Spots: CohereLabs/tiny-aya-base Model Tested CohereLabs/tiny-aya-base Property Value Parameters 3.35 billion (BF16) Architecture Cohere2ForCausalLM Type Pure pre-trained base model (not SFT/RLHF) Languages 70+ languages Released February 13, 2026 License CC-BY-NC-4.0 Context 8K input / 8K output Access Gated (agree to share contact info) Why this model? Tiny Aya is Cohere Labs' open-weights pre-trained 3.35B parameter base… See the full description on the dataset page: https://huggingface.co/datasets/mozayed/tiny-aya-base-blindspots.texttext-generationn<1K0 likes21 downloads7mo agoHugging Face29crosslingual-em /tiny-aya-water-em-en-financial-insecuretabular100K<n<1M0 likes20 downloads5mo agoHugging Face30syedtaha22 /tiny-aya-base-blind-spots tiny-aya-base Blind Spots A dataset of 28 (prompt, expected_output, model_output) triples exposing failure modes of the pretrained base language model CohereLabs/tiny-aya-base. Model Field Value Model CohereLabs/tiny-aya-base Parameters ~3.35B Architecture Cohere2 (hybrid sliding-window + global attention) Type Pretrained base LM — not fine-tuned for any specific task Released March 2026 How the model was loaded The full experiment… See the full description on the dataset page: https://huggingface.co/datasets/syedtaha22/tiny-aya-base-blind-spots.textn<1K0 likes19 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.