datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fitcheck-annotate-datasetmuaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']
وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.mualem-recitations-annotatedwildchat_creative_writing_annotated_10kMalaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.robot0_agentview_left": {
"dtype": "video",
"shape": [
256,
256,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.GAIA-annotatedlaion-voice-profiles-annotated
LAION Voice Profiles — Annotated
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.fineweb_annotatedr1_annotated_aimeUrdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.elements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
synthvision-annotated-qwen
synthvision-annotated-qwen
Medical images annotated by Qwen 3.5 (397B) via Doubleword
Records: 59,476
About
First-half annotations from the SynthVision pipeline. 59,476 medical images annotated by Qwen 3.5 (397B MoE, 17B active) via Doubleword batch inference.
Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating.
Schema
id: str #… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-qwen.COCO-Wholebody-annotatedmuaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.openthoughts-4-math-qwen3-32b-7k-annotated-sharegptopen-thoughts-4-math-qwen3-32b-annotated
Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated
This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the
OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500.
We note that many of the responses are truncated, so use this dataset wisely!
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.speech_commands_enriched_and_annotated
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways:
Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.scale_up_swegym_full_annotated_etash
Dataset card for scale_up_swegym_full_annotated_etash
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"final_prompt": "a very basic just gradio having app\r\n\r\nfast api 111 fixes issue\r\n\r\nall of my followers now reporting errors and so hard to fix all\r\n\r\ni don't know how can you publish such a devastating bug having version please fix it ASAP\r\n\r\n\r\n```\r\n2024-09-06 00:12:20,515 - INFO - HTTP Request: GET… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/scale_up_swegym_full_annotated_etash.scale_up_swegym_full_annotatedopenthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
34.3
74.5
79.4
49.4
51.0
44.3
53.9
21.5
23.1
12.2
17.0
22.7
40.1
AIME24
Average Accuracy: 34.33% ± 1.89%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.Azure-TTS-annotatedsna-dataset-annotated
manassehzw/sna-dataset-annotated
An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline.
This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments.
Why this annotated release exists
The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.R1_Annotated_AIME_True_2Ksynthvision-annotated-kimi
synthvision-annotated-kimi
Medical images annotated by Kimi K2.5 via Doubleword
Records: 59,539
About
Second-half annotations from the SynthVision pipeline. 59,539 medical images annotated by Kimi K2.5 (1T MoE, 32B active) via Doubleword batch inference.
Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating.
Schema
id: str # unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-kimi.CEFR-Annotated-WordNet
CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
Paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
Authors: Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono
Overview
CEFR-Annotated WordNet is a comprehensive semantic database that maps WordNet 3.0 senses and SemCor 3.0 instances to CEFR (Common European Framework of Reference for Languages)… See the full description on the dataset page: https://huggingface.co/datasets/star092304/CEFR-Annotated-WordNet.nanochat-climbmix-annotated
Summary
A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats.
Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code.
Dataset Structure
Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page: https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.control-pretraining-filter-annotated
control-pretraining-filter-annotated
climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.openthoughts-4-code-qwen3-32b-32k-annotatedtulutalk-annotated
🐪💬 TuluTalk: Magpie-Annotated Tülu + SmolTalk Mixture
🌟 Overview
TuluTalk is a lean, high-quality post-training dataset created by merging and filtering two flagship open corpora — Tülu-3 SFT-Mix and SmolTalk — using the Magpie Annotation Framework. Through quality-aware and task-aware curation, TuluTalk achieves 14 % fewer samples than Tülu and 23 % fewer than SmolTalk, yet matches or exceeds their downstream performance across reasoning, math, and coding benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulutalk-annotated.
