datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']
وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.wildchat_creative_writing_annotated_10kmualem-recitations-annotatedMalaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.laion-voice-profiles-annotated
LAION Voice Profiles — Annotated
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.robot0_agentview_left": {
"dtype": "video",
"shape": [
256,
256,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.muaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.open-thoughts-4-math-qwen3-32b-annotated
Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated
This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the
OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500.
We note that many of the responses are truncated, so use this dataset wisely!
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
34.3
74.5
79.4
49.4
51.0
44.3
53.9
21.5
23.1
12.2
17.0
22.7
40.1
AIME24
Average Accuracy: 34.33% ± 1.89%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.openthoughts-4-code-qwen3-32b-32k-annotatedcontrol-pretraining-filter-annotated
control-pretraining-filter-annotated
climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.nanochat-climbmix-annotated
Summary
A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats.
Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code.
Dataset Structure
Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page: https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.tulutalk-annotated
🐪💬 TuluTalk: Magpie-Annotated Tülu + SmolTalk Mixture
🌟 Overview
TuluTalk is a lean, high-quality post-training dataset created by merging and filtering two flagship open corpora — Tülu-3 SFT-Mix and SmolTalk — using the Magpie Annotation Framework. Through quality-aware and task-aware curation, TuluTalk achieves 14 % fewer samples than Tülu and 23 % fewer than SmolTalk, yet matches or exceeds their downstream performance across reasoning, math, and coding benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulutalk-annotated.rocochallenge2026_Industrial_Assembly_annotatedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"observation.images.head": {
"dtype": "video",
"shape": [
240,
320,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/nikodembartnik/rocochallenge2026_Industrial_Assembly_annotated.open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-32B, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-32B.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Model: Qwen/Qwen3-32B
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset identifier
gpt41_mini_response
Reference response from… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8.laion-tts-annotated-v1
LAION TTS Annotated v1
107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens
and the annotations, all joined by one key.
283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion
intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst
detections with timings, and a natural-language caption describing the voice and the
delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.math_stratos_scale_judged_and_annotated_with_difficultyopen-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The code problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.Rasa-Annotated-25kHz
Rasa-Annotated
Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,102
Total Duration: 46.92 hours
Average Duration: 6.47 seconds
Duration Range: 0.31s - 45.34s
Average Phonemes: 18.4 per sample
Average Kanade Tokens: 530.7 per sample
Global Embedding Dimension: 128
Gender Distribution
Gender
Count
Female
12,583
Male
13,519
Style Distribution
Style
Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-25kHz.fold_clothes_dining_annotated2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"fps": 50,
"features": {
"action": {
"dtype": "float32",
"shape": [
14
]
},
"observation.state": {
"dtype": "float32",
"shape": [
14
]
},
"observation.images.front": {
"dtype": "video"… See the full description on the dataset page: https://huggingface.co/datasets/nikodembartnik/fold_clothes_dining_annotated2.open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated (base prompts)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.beamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted
Overview
This dataset is a reformatted version of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.tulu-3-sft-mix-annotated
🐪 Tülu-3-Annotated: Magpie-Extended Structured SFT Dataset
🌟 Overview
Tülu-3-Annotated is a fully Magpie-tagged version of the original Tülu-3 SFT-Mix supervised-fine-tuning (SFT) dataset introduced with Tülu 3 (2025). Each instruction–response pair has been enriched with detailed MagPie annotations covering task category, input quality, response reward, safety, and conversation structure—enabling fine-grained data-quality analysis and curation research for… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulu-3-sft-mix-annotated.open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-4B, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-4B.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens (n=1 version with 1 response per prompt)
Model: Qwen/Qwen3-4B
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset identifier… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8.mls-annotated
Dataset Card for Annotations of non English MLS
This dataset consists in annotations of a the Non English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.… See the full description on the dataset page: https://huggingface.co/datasets/PHBJT/mls-annotated.smoltalk-annotated
💬 SmolTalk-Annotated: Magpie-Extended Multi-Turn SFT Dataset
🌟 Overview
SmolTalk-Annotated is a fully Magpie-tagged version of the original SmolTalk supervised-fine-tuning (SFT) corpus introduced with SmolLM 2 (2025). Each conversation has been enriched with detailed Magpie annotations covering task category, input quality, instruction reward, conversation depth, safety, and more—making this dataset ideal for data-centric LLM analysis and post-training research.… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/smoltalk-annotated.openthoughts3_300k_annotated_Qwen3-32BRasa-Annotated-V1
Rasa-Annotated
Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,102
Total Duration: 46.92 hours
Average Duration: 6.47 seconds
Duration Range: 0.31s - 45.34s
Average Phonemes: 18.4 per sample
Average Kanade Tokens: 264.5 per sample
Global Embedding Dimension: 128
Gender Distribution
Gender
Count
Female
12,583
Male
13,519
Style Distribution
Style
Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-V1.
