CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tiny-aya-translate /tr-hi-mimi-encoded TR↔HI Mimi-Encoded Parallel Speech Pre-encoded parallel Turkish↔Hindi speech pairs for training speech-to-speech translation models. All audio has been tokenized through the Mimi neural audio codec (8 codebooks, 12.5 Hz, 24kHz) and stored as .pt files with word-level text alignments. Dataset Summary Source audio ~911 hours of synthetic parallel TR↔HI speech from tr-hi-parallel-speech-v2 TTS model OmniVoice (voice design mode, 14 voice designs)… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-mimi-encoded.textaudio-to-audio1M<n<10M1 likes87 downloads2mo agoHugging Face02tiny-aya-translate /fleurs-tr-hi-mimi-encoded fleurs-tr-hi-mimi-encoded Mimi-encoded Turkish↔Hindi parallel speech pairs for TinyAya Stage 2 speech-to-speech translation training. Contents encoded/*.pt — 9212 Mimi-encoded audio pairs (kyutai/mimi, 8 codebooks, 12.5 Hz, 24 kHz). Each file keys: pair_id, src_lang, tgt_lang, src_text, tgt_text, src_codes[8, T_src], tgt_codes[8, T_tgt]. encoded/*.alignments.json — 18424 Whisper word-level alignment sidecars (.src.alignments.json / .tgt.alignments.json).… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/fleurs-tr-hi-mimi-encoded.textaudio-to-audio1K<n<10K1 likes30 downloads2mo agoHugging Face03bingbangboom /tiny-aya-translate-hinglish-casual-stripped Dataset Card for tiny-aya-translate-hinglish-casual-stripped Dataset Summary tiny-aya-translate-hinglish-casual-stripped is a lightweight, text-only derivative of the original tiny-aya-translate/hinglish-casual dataset. The original dataset is designed for simultaneous translation and contains many columns including audio references, speaker metadata, and duration. It also includes paralinguistic tags (e.g., <sigh>, <laugh>, <chuckle>) embedded within the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/tiny-aya-translate-hinglish-casual-stripped.texttext-generation10K<n<100K1 likes15 downloads3mo agoHugging Face04Mawube /tiny-aya-base-blind-spots Blind Spots of a Frontier Base Model: Evaluation Dataset This dataset documents blind spots discovered in a frontier open-weight base model through 19 structured evaluation tests. It was assembled as part of an assignment on identifying model weaknesses using the HelloBench evaluation framework. Model Tested CohereLabs/tiny-aya-base Architecture: Transformer with Sliding Window Attention (SWA) (window size 4096, with RoPE) on three layers + one global attention layer… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/tiny-aya-base-blind-spots.tabulartext-generationn<1K0 likes14 downloads7mo agoHugging Face05s4um1l /tiny-aya-medical-concept-probes Tiny Aya Cross-Lingual Medical Concept Probes Dataset Description 20 medical concepts expressed as full sentences in 10 languages, designed for probing cross-lingual concept representations in multilingual LLMs. Each concept is a complete declarative sentence preserving the same semantic structure across all languages. Purpose These probe sentences serve as stimuli for mechanistic interpretability analysis -- specifically, extracting residual stream activations… See the full description on the dataset page: https://huggingface.co/datasets/s4um1l/tiny-aya-medical-concept-probes.textfeature-extractionn<1K0 likes11 downloads7mo agoHugging Face06safety-aya /toxigen_tiny-portuguesetext1K<n<10K0 likes10 downloads6mo agoHugging Face07KarmaIncarnate /tiny-aya-base-blindspots tiny-aya-base Blind Spot Dataset (10 probes) This is a small, hand-curated evaluation set of 10 diverse prompts that surface failure modes of the base model CohereLabs/tiny-aya-base. Every row contains the prompt prefix, the expected answer, and the model’s actual output (greedy decoding). The goal isn’t to be exhaustive—it’s to capture where the model is most brittle: small arithmetic, ratios, formal negation, calendar logic, code errors, and low‑resource language meaning.… See the full description on the dataset page: https://huggingface.co/datasets/KarmaIncarnate/tiny-aya-base-blindspots.textn<1K0 likes9 downloads7mo agoHugging Face08kaustubhg73 /tiny-aya-global-blindspots tiny-aya-global-blindspots Title & Overview A curated set of failure cases for CohereLabs/tiny-aya-global, showcasing blind spots discovered while probing the ~3.35B-parameter base checkpoint released in February 2026. Each entry captures a prompt, the expected aligned behavior, and the model's actual output. The dataset illustrates common failure patterns observed when probing the ~3.35B multilingual base model without additional alignment or safety fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/kaustubhg73/tiny-aya-global-blindspots.textn<1K0 likes6 downloads7mo agoHugging Face09akotet08 /tiny-aya_failure Dataset Summary A compact diagnostic benchmark for evaluating failure modes in language models. It tests a model's ability to resist false premises, avoid fabricating entities, detect contradictions, follow strict constraints, and recognize category mismatches. Each record includes: id, category, prompt, model_output, and expected_correct_output. Model Outputs were generated using CohereLabs/tiny-aya-global via the Hugging Face transformers library. Generation settings:… See the full description on the dataset page: https://huggingface.co/datasets/akotet08/tiny-aya_failure.textn<1K0 likes2 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.