CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01barryallen16 /fitcheck-annotate-datasettext10K<n<100K0 likes4.1k downloads5d agoHugging Face02obadx /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train'] وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.audio100K<n<1M7 likes2.3k downloads1y agoHugging Face03obadx /mualem-recitations-annotatedaudio100K<n<1M4 likes2.3k downloads1y agoHugging Face04sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.2k downloads9mo agoHugging Face05mesolitica /Malaysian-Emilia-annotated Malaysian Emilia Annotated Annotate Malaysian-Emilia using Data-Speech pipeline. Malaysian Youtube Originally from malaysia-ai/crawl-youtube Total 3168.8 hours. Gender prediction, filtered-24k_processed_24k_gender.zip Language prediction, filtered-24k_processed_language.zip Force alignment. Post cleaned to 24k and 44k sampling rates, 24k, filtered-24k_processed_24k.zip 44k, filtered-24k_processed_44k.zip Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.tabulartext-to-speech1M<n<10M2 likes2k downloads1y agoHugging Face06pepijn223 /robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.images.robot0_agentview_left": { "dtype": "video", "shape": [ 256, 256, 3 ], "names": [ "height", "width", "channel" ], "video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.tabularrobotics10M<n<100M2 likes1.8k downloads3mo agoHugging Face07smolagents /GAIA-annotateddocumentn<1K1 likes1.7k downloads1y agoHugging Face08laion /laion-voice-profiles-annotated LAION Voice Profiles — Annotated Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.tabulartext-to-speech10M<n<100M0 likes1.4k downloads12d agoHugging Face09locuslab /fineweb_annotatedtext100M<n<1B2 likes1.4k downloads4mo agoHugging Face10mlfoundations-dev /r1_annotated_aimetext1K<n<10K0 likes1.1k downloads2y agoHugging Face11humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes965 downloads8mo agoHugging Face12cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes860 downloads11mo agoHugging Face13OpenMed /synthvision-annotated-qwen synthvision-annotated-qwen Medical images annotated by Qwen 3.5 (397B) via Doubleword Records: 59,476 About First-half annotations from the SynthVision pipeline. 59,476 medical images annotated by Qwen 3.5 (397B MoE, 17B active) via Doubleword batch inference. Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating. Schema id: str #… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-qwen.textvisual-question-answering10K<n<100K2 likes839 downloads6mo agoHugging Face14harsh-7070 /COCO-Wholebody-annotatedimage100K<n<1M0 likes740 downloads2y agoHugging Face15nour-world /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.audio100K<n<1M0 likes665 downloads6d agoHugging Face16laion /openthoughts-4-math-qwen3-32b-7k-annotated-sharegpttext1M<n<10M0 likes615 downloads9mo agoHugging Face17marin-community /open-thoughts-4-math-qwen3-32b-annotated Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500. We note that many of the responses are truncated, so use this dataset wisely! Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.tabular1M<n<10M0 likes555 downloads10mo agoHugging Face18soerenray /speech_commands_enriched_and_annotated Dataset Summary 📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. 🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways: Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.audio10K<n<100K2 likes540 downloads3y agoHugging Face19mlfoundations-dev /scale_up_swegym_full_annotated_etash Dataset card for scale_up_swegym_full_annotated_etash This dataset was made with Curator. Dataset details A sample from the dataset: { "final_prompt": "a very basic just gradio having app\r\n\r\nfast api 111 fixes issue\r\n\r\nall of my followers now reporting errors and so hard to fix all\r\n\r\ni don't know how can you publish such a devastating bug having version please fix it ASAP\r\n\r\n\r\n```\r\n2024-09-06 00:12:20,515 - INFO - HTTP Request: GET… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/scale_up_swegym_full_annotated_etash.text100K<n<1M0 likes460 downloads1y agoHugging Face20mlfoundations-dev /scale_up_swegym_full_annotatedtext100K<n<1M0 likes456 downloads1y agoHugging Face21mlfoundations-dev /openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 34.3 74.5 79.4 49.4 51.0 44.3 53.9 21.5 23.1 12.2 17.0 22.7 40.1 AIME24 Average Accuracy: 34.33% ± 1.89% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.tabular10K<n<100K0 likes455 downloads1y agoHugging Face22mesolitica /Azure-TTS-annotatedaudio100K<n<1M0 likes445 downloads2y agoHugging Face23manassehzw /sna-dataset-annotated manassehzw/sna-dataset-annotated An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline. This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments. Why this annotated release exists The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.audioautomatic-speech-recognition10K<n<100K1 likes440 downloads2mo agoHugging Face24Brench /R1_Annotated_AIME_True_2Ktext1K<n<10K0 likes417 downloads1y agoHugging Face25OpenMed /synthvision-annotated-kimi synthvision-annotated-kimi Medical images annotated by Kimi K2.5 via Doubleword Records: 59,539 About Second-half annotations from the SynthVision pipeline. 59,539 medical images annotated by Kimi K2.5 (1T MoE, 32B active) via Doubleword batch inference. Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating. Schema id: str # unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-kimi.textvisual-question-answering10K<n<100K1 likes388 downloads6mo agoHugging Face26star092304 /CEFR-Annotated-WordNet CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning Paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning Authors: Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono Overview CEFR-Annotated WordNet is a comprehensive semantic database that maps WordNet 3.0 senses and SemCor 3.0 instances to CEFR (Common European Framework of Reference for Languages)… See the full description on the dataset page: https://huggingface.co/datasets/star092304/CEFR-Annotated-WordNet.text100K<n<1M1 likes378 downloads4mo agoHugging Face27ddudek /nanochat-climbmix-annotated Summary A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats. Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code. Dataset Structure Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page: https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.tabulartext-classification10M<n<100M0 likes377 downloads5mo agoHugging Face28sudoers /control-pretraining-filter-annotatedgated control-pretraining-filter-annotated climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.tabular100M<n<1B0 likes362 downloads12m agoHugging Face29laion /openthoughts-4-code-qwen3-32b-32k-annotatedtabular100K<n<1M2 likes332 downloads9mo agoHugging Face30aladinDJ /tulutalk-annotated 🐪💬 TuluTalk: Magpie-Annotated Tülu + SmolTalk Mixture 🌟 Overview TuluTalk is a lean, high-quality post-training dataset created by merging and filtering two flagship open corpora — Tülu-3 SFT-Mix and SmolTalk — using the Magpie Annotation Framework. Through quality-aware and task-aware curation, TuluTalk achieves 14 % fewer samples than Tülu and 23 % fewer than SmolTalk, yet matches or exceeds their downstream performance across reasoning, math, and coding benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulutalk-annotated.tabular100K<n<1M2 likes326 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.