CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01obadx /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train'] وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.audio100K<n<1M7 likes2.7k downloads1y agoHugging Face02sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.5k downloads9mo agoHugging Face03obadx /mualem-recitations-annotatedaudio100K<n<1M4 likes2.3k downloads1y agoHugging Face04mesolitica /Malaysian-Emilia-annotated Malaysian Emilia Annotated Annotate Malaysian-Emilia using Data-Speech pipeline. Malaysian Youtube Originally from malaysia-ai/crawl-youtube Total 3168.8 hours. Gender prediction, filtered-24k_processed_24k_gender.zip Language prediction, filtered-24k_processed_language.zip Force alignment. Post cleaned to 24k and 44k sampling rates, 24k, filtered-24k_processed_24k.zip 44k, filtered-24k_processed_44k.zip Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.tabulartext-to-speech1M<n<10M2 likes2.1k downloads1y agoHugging Face05locuslab /fineweb_annotatedtext100M<n<1B2 likes1.6k downloads4mo agoHugging Face06humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face07mlfoundations-dev /r1_annotated_aimetext1K<n<10K0 likes1.3k downloads2y agoHugging Face08laion /laion-voice-profiles-annotated LAION Voice Profiles — Annotated Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.tabulartext-to-speech10M<n<100M0 likes1.2k downloads14d agoHugging Face09pepijn223 /robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.images.robot0_agentview_left": { "dtype": "video", "shape": [ 256, 256, 3 ], "names": [ "height", "width", "channel" ], "video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.tabularrobotics10M<n<100M2 likes1.1k downloads3mo agoHugging Face10cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes971 downloads11mo agoHugging Face11laion /openthoughts-4-math-qwen3-32b-7k-annotated-sharegpttext1M<n<10M0 likes871 downloads9mo agoHugging Face12harsh-7070 /COCO-Wholebody-annotatedimage100K<n<1M0 likes748 downloads2y agoHugging Face13OpenMed /synthvision-annotated-qwen synthvision-annotated-qwen Medical images annotated by Qwen 3.5 (397B) via Doubleword Records: 59,476 About First-half annotations from the SynthVision pipeline. 59,476 medical images annotated by Qwen 3.5 (397B MoE, 17B active) via Doubleword batch inference. Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating. Schema id: str #… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-qwen.textvisual-question-answering10K<n<100K2 likes698 downloads6mo agoHugging Face14nour-world /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.audio100K<n<1M0 likes670 downloads9d agoHugging Face15soerenray /speech_commands_enriched_and_annotated Dataset Summary 📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. 🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways: Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.audio10K<n<100K2 likes626 downloads3y agoHugging Face16marin-community /open-thoughts-4-math-qwen3-32b-annotated Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500. We note that many of the responses are truncated, so use this dataset wisely! Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.tabular1M<n<10M0 likes580 downloads10mo agoHugging Face17mlfoundations-dev /openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 34.3 74.5 79.4 49.4 51.0 44.3 53.9 21.5 23.1 12.2 17.0 22.7 40.1 AIME24 Average Accuracy: 34.33% ± 1.89% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.tabular10K<n<100K0 likes544 downloads1y agoHugging Face18laion /openthoughts-4-code-qwen3-32b-32k-annotatedtabular100K<n<1M2 likes525 downloads10mo agoHugging Face19mesolitica /Azure-TTS-annotatedaudio100K<n<1M0 likes499 downloads2y agoHugging Face20Brench /R1_Annotated_AIME_True_2Ktext1K<n<10K0 likes469 downloads2y agoHugging Face21mlfoundations-dev /scale_up_swegym_full_annotated_etash Dataset card for scale_up_swegym_full_annotated_etash This dataset was made with Curator. Dataset details A sample from the dataset: { "final_prompt": "a very basic just gradio having app\r\n\r\nfast api 111 fixes issue\r\n\r\nall of my followers now reporting errors and so hard to fix all\r\n\r\ni don't know how can you publish such a devastating bug having version please fix it ASAP\r\n\r\n\r\n```\r\n2024-09-06 00:12:20,515 - INFO - HTTP Request: GET… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/scale_up_swegym_full_annotated_etash.text100K<n<1M0 likes469 downloads1y agoHugging Face22sudoers /control-pretraining-filter-annotatedgated control-pretraining-filter-annotated climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.tabular100M<n<1B0 likes462 downloads2d agoHugging Face23mlfoundations-dev /scale_up_swegym_full_annotatedtext100K<n<1M0 likes443 downloads1y agoHugging Face24manassehzw /sna-dataset-annotated manassehzw/sna-dataset-annotated An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline. This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments. Why this annotated release exists The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.audioautomatic-speech-recognition10K<n<100K1 likes437 downloads2mo agoHugging Face25manassehzw /sna-waxal-annotated-unlabeled Shona WAXAL annotated-unlabeled checkpoint This is a self-contained operational checkpoint for pseudo-labeling Shona ASR data. It contains 90,253 conservatively segmented FLAC clips (441.585 hours), but intentionally contains no transcripts. Fields transcription is intentionally empty. speaker_id is an approximate source-blind EOM cluster or unknown; speaker_clip_count is zero for unknown assignments. gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.audioautomatic-speech-recognition10K<n<100K0 likes374 downloads2mo agoHugging Face26laion /Emilia-Annotated-WIPStill a WIP, full dataset is still being annotated audio1M<n<10M3 likes367 downloads1y agoHugging Face27ddudek /nanochat-climbmix-annotated Summary A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats. Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code. Dataset Structure Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page: https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.tabulartext-classification10M<n<100M0 likes330 downloads6mo agoHugging Face28OpenMed /synthvision-annotated-kimi synthvision-annotated-kimi Medical images annotated by Kimi K2.5 via Doubleword Records: 59,539 About Second-half annotations from the SynthVision pipeline. 59,539 medical images annotated by Kimi K2.5 (1T MoE, 32B active) via Doubleword batch inference. Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating. Schema id: str # unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-kimi.textvisual-question-answering10K<n<100K1 likes321 downloads6mo agoHugging Face29laion /openthoughts-4-code-qwen3-32b-7k-annotated-sharegpttext100K<n<1M0 likes306 downloads9mo agoHugging Face30aladinDJ /tulutalk-annotated 🐪💬 TuluTalk: Magpie-Annotated Tülu + SmolTalk Mixture 🌟 Overview TuluTalk is a lean, high-quality post-training dataset created by merging and filtering two flagship open corpora — Tülu-3 SFT-Mix and SmolTalk — using the Magpie Annotation Framework. Through quality-aware and task-aware curation, TuluTalk achieves 14 % fewer samples than Tülu and 23 % fewer than SmolTalk, yet matches or exceeds their downstream performance across reasoning, math, and coding benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulutalk-annotated.tabular100K<n<1M2 likes296 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.