CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01obadx /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train'] وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.audio100K<n<1M7 likes2.7k downloads1y agoHugging Face02sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.5k downloads9mo agoHugging Face03obadx /mualem-recitations-annotatedaudio100K<n<1M4 likes2.3k downloads1y agoHugging Face04mesolitica /Malaysian-Emilia-annotated Malaysian Emilia Annotated Annotate Malaysian-Emilia using Data-Speech pipeline. Malaysian Youtube Originally from malaysia-ai/crawl-youtube Total 3168.8 hours. Gender prediction, filtered-24k_processed_24k_gender.zip Language prediction, filtered-24k_processed_language.zip Force alignment. Post cleaned to 24k and 44k sampling rates, 24k, filtered-24k_processed_24k.zip 44k, filtered-24k_processed_44k.zip Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.tabulartext-to-speech1M<n<10M2 likes2.1k downloads1y agoHugging Face05humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face06laion /laion-voice-profiles-annotated LAION Voice Profiles — Annotated Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.tabulartext-to-speech10M<n<100M0 likes1.2k downloads14d agoHugging Face07pepijn223 /robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.images.robot0_agentview_left": { "dtype": "video", "shape": [ 256, 256, 3 ], "names": [ "height", "width", "channel" ], "video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.tabularrobotics10M<n<100M2 likes1.1k downloads3mo agoHugging Face08nour-world /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.audio100K<n<1M0 likes670 downloads9d agoHugging Face09marin-community /open-thoughts-4-math-qwen3-32b-annotated Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500. We note that many of the responses are truncated, so use this dataset wisely! Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.tabular1M<n<10M0 likes580 downloads10mo agoHugging Face10mlfoundations-dev /openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 34.3 74.5 79.4 49.4 51.0 44.3 53.9 21.5 23.1 12.2 17.0 22.7 40.1 AIME24 Average Accuracy: 34.33% ± 1.89% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.tabular10K<n<100K0 likes544 downloads1y agoHugging Face11laion /openthoughts-4-code-qwen3-32b-32k-annotatedtabular100K<n<1M2 likes525 downloads10mo agoHugging Face12sudoers /control-pretraining-filter-annotatedgated control-pretraining-filter-annotated climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.tabular100M<n<1B0 likes462 downloads2d agoHugging Face13ddudek /nanochat-climbmix-annotated Summary A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats. Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code. Dataset Structure Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page: https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.tabulartext-classification10M<n<100M0 likes330 downloads6mo agoHugging Face14aladinDJ /tulutalk-annotated 🐪💬 TuluTalk: Magpie-Annotated Tülu + SmolTalk Mixture 🌟 Overview TuluTalk is a lean, high-quality post-training dataset created by merging and filtering two flagship open corpora — Tülu-3 SFT-Mix and SmolTalk — using the Magpie Annotation Framework. Through quality-aware and task-aware curation, TuluTalk achieves 14 % fewer samples than Tülu and 23 % fewer than SmolTalk, yet matches or exceeds their downstream performance across reasoning, math, and coding benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulutalk-annotated.tabular100K<n<1M2 likes296 downloads11mo agoHugging Face15nikodembartnik /rocochallenge2026_Industrial_Assembly_annotatedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 10, "features": { "observation.images.head": { "dtype": "video", "shape": [ 240, 320, 3 ], "names": [ "height", "width", "channels" ], "info": { "video.height":… See the full description on the dataset page: https://huggingface.co/datasets/nikodembartnik/rocochallenge2026_Industrial_Assembly_annotated.tabularrobotics100K<n<1M0 likes275 downloads2mo agoHugging Face16marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8 Open Thoughts 4 - Math (Qwen3-32B, 32K tokens, n=8) This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-32B. Overview Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Model: Qwen/Qwen3-32B Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The math problem prompt _source Source dataset identifier gpt41_mini_response Reference response from… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8.tabular10K<n<100K1 likes259 downloads8mo agoHugging Face17laion /laion-tts-annotated-v1 LAION TTS Annotated v1 107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens and the annotations, all joined by one key. 283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst detections with timings, and a natural-language caption describing the voice and the delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.tabulartext-to-speech100M<n<1B0 likes250 downloads12d agoHugging Face18mlfoundations-dev /math_stratos_scale_judged_and_annotated_with_difficultytabular100K<n<1M0 likes242 downloads2y agoHugging Face19marin-community /open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8 Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8) This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507. Overview Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only) Model: Qwen/Qwen3-30B-A3B-Thinking-2507 Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The code problem prompt _source Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.tabular10K<n<100K0 likes222 downloads7mo agoHugging Face20humair025 /Rasa-Annotated-25kHz Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 530.7 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-25kHz.tabulartext-to-speech10K<n<100K0 likes208 downloads8mo agoHugging Face21nikodembartnik /fold_clothes_dining_annotated2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "fps": 50, "features": { "action": { "dtype": "float32", "shape": [ 14 ] }, "observation.state": { "dtype": "float32", "shape": [ 14 ] }, "observation.images.front": { "dtype": "video"… See the full description on the dataset page: https://huggingface.co/datasets/nikodembartnik/fold_clothes_dining_annotated2.tabularrobotics100K<n<1M0 likes205 downloads2mo agoHugging Face22marin-community /open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8 Open Thoughts 4 - Math (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8) This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507. Overview Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated (base prompts) Model: Qwen/Qwen3-30B-A3B-Thinking-2507 Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The math problem prompt _source Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.tabular10K<n<100K0 likes203 downloads7mo agoHugging Face23acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes200 downloads3y agoHugging Face24marin-community /open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted Overview This dataset is a reformatted version of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.tabular100K<n<1M0 likes199 downloads6mo agoHugging Face25aladinDJ /tulu-3-sft-mix-annotated 🐪 Tülu-3-Annotated: Magpie-Extended Structured SFT Dataset 🌟 Overview Tülu-3-Annotated is a fully Magpie-tagged version of the original Tülu-3 SFT-Mix supervised-fine-tuning (SFT) dataset introduced with Tülu 3 (2025). Each instruction–response pair has been enriched with detailed MagPie annotations covering task category, input quality, response reward, safety, and conversation structure—enabling fine-grained data-quality analysis and curation research for… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulu-3-sft-mix-annotated.tabular100K<n<1M0 likes197 downloads11mo agoHugging Face26marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8 Open Thoughts 4 - Math (Qwen3-4B, 32K tokens, n=8) This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-4B. Overview Source: marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens (n=1 version with 1 response per prompt) Model: Qwen/Qwen3-4B Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The math problem prompt _source Source dataset identifier… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8.tabular10K<n<100K0 likes194 downloads7mo agoHugging Face27PHBJT /mls-annotated Dataset Card for Annotations of non English MLS This dataset consists in annotations of a the Non English subset of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.… See the full description on the dataset page: https://huggingface.co/datasets/PHBJT/mls-annotated.tabulartext-to-speech1M<n<10M1 likes179 downloads2y agoHugging Face28aladinDJ /smoltalk-annotated 💬 SmolTalk-Annotated: Magpie-Extended Multi-Turn SFT Dataset 🌟 Overview SmolTalk-Annotated is a fully Magpie-tagged version of the original SmolTalk supervised-fine-tuning (SFT) corpus introduced with SmolLM 2 (2025). Each conversation has been enriched with detailed Magpie annotations covering task category, input quality, instruction reward, conversation depth, safety, and more—making this dataset ideal for data-centric LLM analysis and post-training research.… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/smoltalk-annotated.tabular1M<n<10M0 likes176 downloads11mo agoHugging Face29mlfoundations-dev /openthoughts3_300k_annotated_Qwen3-32Btabular100K<n<1M0 likes168 downloads1y agoHugging Face30humair025 /Rasa-Annotated-V1 Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 264.5 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-V1.tabulartext-to-speech10K<n<100K0 likes163 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.