datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepResearch-Bench-II-DatasetMusicPile🌐 DemoPage | 🤗SFT Dataset | 🤗 Benchmark | 📖 arXiv | 💻 Code | 🤖 Chat Model | 🤖 Base Model
Dataset Card for MusicPile
MusicPile is the first pretraining corpus for developing musical abilities in large language models.
It has 5.17M samples and approximately 4.16B tokens, including web-crawled corpora, encyclopedias, music books, youtube music captions, musical pieces in abc notation, math content, and code.
You can easily load it:from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MusicPile.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.DeepResearch-Bench-Dataset
DeepResearch Bench Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the DeepResearch Bench paper. It contains research reports generated by 4 leading deep research AI systems along with detailed human expert annotations evaluating these reports.
DeepResearch Bench is the first comprehensive benchmark for systematically evaluating Deep Research Agents (DRAs) on their ability to handle complex, PhD-level research… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-Dataset.muslim-names-dataset
Muslim Names Dataset
A comprehensive collection of Muslim names with meanings scraped from muslimnames.com. Contains 14,585 names with English names, Arabic names, meanings, and gender classifications.
Dataset Contents
This dataset contains ~14,585 Muslim names with the following information:
English name: Name in English/Latin script
Arabic name: Name in Arabic script
Meaning: Definition and meaning of the name
Gender: Classification as male or female
Files… See the full description on the dataset page: https://huggingface.co/datasets/takiuddinahmed/muslim-names-dataset.qwen3.5-toolcalling-v2
Qwen3.5 Tool Calling Dataset v2
An expanded tool-calling SFT dataset combining smirki/Tool-Calling-Dataset-UIGEN-X and AmanPriyanshu/tool-reasoning-sft-jupyter-agent, unified into Qwen3 messages format. Adds Jupyter notebook agent data with code execution reasoning chains.
Dataset Summary
Property
Value
Total Samples
~60K+
Train Split
~55K
Test Split
~6K
Sources
UIGEN-X + Jupyter Agent
Format
Qwen3 messages
Language
English
License
Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/Mustafaege/qwen3.5-toolcalling-v2.MathNet
Shaden Alshammari1* Kevin Wen1* Abrar Zainal3* Mark Hamilton1
Navid Safaei4 Sultan Albarakati2 William T. Freeman1† Antonio Torralba1†
1MIT 2KAUST 3HUMAIN 4Bulgarian Academy of Sciences *† equal contribution
Quick Start· Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation
Note: This is a test for the HF hosting website. The dataset isn’t fully uploaded yet; it will be uploaded on Tuesday, April 21, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/musumecmtcd/MathNet.Muse-Glimmer-SWE-Gym-2k
Muse-Glimmer-SWE-Gym-2k
Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a
speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and
SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them.
Configs
Config
Rows
Size
What it is
train
1,981
57 MB
One row per trajectory: the full conversation as messages.
raw
159,999
2.7 GB
One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.model-egitme-sft-v1
Türkçe İK Belge Zekâsı — SFT verisi ve LoRA adapterleri
Türkçe bir insan kaynakları bildirimini 29 alanlı bir JSON
sözleşmesine çeviren küçük bir modelin eğitim verisi ve adapterleri.
Canlı demo: https://huggingface.co/spaces/mustafabasar/ik-belge-asistani-demo
Ne var burada
yol
ne
veri_v3/ … veri_v9/
sürümlenmiş eğitim/doğrulama verisi (train.jsonl, dev.jsonl)
v34/ … v45/
LoRA adapterleri (lora/) ve eğitim kayıtları
sema/
Pydantic sözleşmesi… See the full description on the dataset page: https://huggingface.co/datasets/mustafabasar/model-egitme-sft-v1.muse12-nemo-agentic
Muse Spark 1.2 High-Reasoning NeMo Agentic Dataset
A reproducible, verified 24,000-row synthetic agentic dataset generated with Meta Muse Spark 1.2, NeMo Gym, and deterministic task-family verifiers.
The project is a quality-focused successor to r0b0tlab/deepseek-v4-pro-0813-agentic. It keeps rollout prompts separate from reference trajectories and offline-training views, records usage and provenance, and does not publish private chain-of-thought.
[!IMPORTANT]
Status:… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/muse12-nemo-agentic.music-codes
互联网歌曲数据集
使用encodec编码的歌曲数据,sampling_rate=24000, bandwidth=6.0, n_codebooks=8, 截取时长=30s
各数据列说明
id : 歌曲源id
name : 歌曲名
singer : 歌手名
text : 歌词
code : 原音乐使用EnCodec编码的结果, 是形状为[T,8]的二维列表。
smithsonian-data
Smithsonian Open Access Data
Pre-processed data dumps from the Smithsonian Open Access initiative, covering millions of objects across Smithsonian Institution museums and archives.
What is this?
The Smithsonian publishes their Open Access metadata on S3, but the raw data is split across 255 individual .txt files per unit. This dataset consolidates each unit's data into a single .jsonl.gz file for easier downloading and processing.
Files
Each file corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/museado/smithsonian-data.MuslimLife
Muslim Life Knowledge Base & RAG Dataset
Contains 90 public Simplified Chinese articles from the Salaam Alykum 穆斯林生活 / Muslim Life topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and a lightweight embedding preview layer.
[!TIP]
Human Readers / 普通读者: For normal reading, open Files and versions -> content and start with content/README.md. Example article: 2744 莱麦丹不同面貌:斋月中的人、故事与信仰现场. For… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/MuslimLife.crimsonred-paper-replication
CrimsonRed — Cross-Architecture Emotion-Prime Steering Replication
Replication of the emotion-prime steering protocol from arXiv:2607.18691 (NSM
semantic primes as explanans for emotion in LLMs), extended across four
architectures. Generated by scripts/paper_faithful_steering.py in the
CrimsonRed project.
The finding
The paper's core claim — that semantic-prime recipe directions steer emotion more
strongly than Scherer appraisal directions — replicates… See the full description on the dataset page: https://huggingface.co/datasets/musicakamusic/crimsonred-paper-replication.UncertaintyGym
UncertaintyGym
A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression
Abstract
UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating.
Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.wikipedia-tr
📖 Türkçe Vikipedi Mayıs 2023
Bu veri kümesi, Türkçe Vikipedi'den alınan makalelerin bir derlemesi olup, maskeleme dil modelleme ve metin oluşturma görevleri için tasarlanmıştır.
🗣️ Etiketlemeler
Bu veri kümesindeki makaleler, özellikle belirli bir görev için etiketlenmemiş olup, veri kümesi etiketsizdir.
🌐 Dil
Bu veri kümesi Türkçe yazılmış olup, gönüllülerden oluşan bir ekip tarafından topluluk katılımı yöntemleri ile oluşturulmuştur.
📜 Lisans… See the full description on the dataset page: https://huggingface.co/datasets/musabg/wikipedia-tr.Chinese-Muslim-Travel
☪ Chinese-Muslim-Travel: Native Chinese Muslim Travel RAG Corpus
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
Dataset Description
Chinese-Muslim-Travel is a curated RAG corpus containing 347 native Chinese articles documenting Muslim travel, halal food, mosque architecture, and Muslim community life across 20+ countries. Every… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Chinese-Muslim-Travel.Muse-Glimmer-OPB-100K
Muse Glimmer OPB 100K
On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark.
Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B.
The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths.
Reasoning strength
Conversations
Train-turn rows
low
64,997
96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.qwen3.5-toolcalling-v1
Qwen3.5 Tool Calling Dataset v1
A tool-calling SFT dataset built from smirki/Tool-Calling-Dataset-UIGEN-X (a cleaned version of interstellarninja/hermes_reasoning_tool_use), converted from ShareGPT conversations format to Qwen3 messages format. Features deep reasoning chains with <think> tags followed by structured tool calls.
Dataset Summary
Property
Value
Total Samples
51,004
Train Split
45,904
Test Split
5,100
Source
smirki/Tool-Calling-Dataset-UIGEN-X… See the full description on the dataset page: https://huggingface.co/datasets/Mustafaege/qwen3.5-toolcalling-v1.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.ifparse-v1.0
ifparse v1.0
IFParse is a benchmark for structured extraction from real developer logs: can a model turn a raw, unstructured log line into JSON that satisfies a fixed schema, with no code fence, no commentary, and no type errors. Scoring is binary and covers only compliance, not extraction creativity, so the signal is isolated to whether the output would actually parse in a production pipeline.
Each prompt gives the model a single raw log line, either an Apache access log record… See the full description on the dataset page: https://huggingface.co/datasets/Muse-research/ifparse-v1.0.Hui-Muslims
Hui Muslims RAG Dataset
Contains 232 native Chinese articles.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.loft-rag-musique-32k
LOFT RAG - MuSiQue (32k)
Dataset Description
This dataset is part of the LOFT (Long-context Open Foundation Tasks) benchmark, specifically the RAG (Retrieval-Augmented Generation) task.
Dataset: MuSiQue
Context Length: 32k
Task Type: RAG (Retrieval-Augmented Generation)
Language: English
Source: LOFT Benchmark (Google DeepMind)
Dataset Structure
Data Fields
context (string): Full prompt context including corpus documents and few-shot… See the full description on the dataset page: https://huggingface.co/datasets/f20180301/loft-rag-musique-32k.QuranExeThis dataset contains the exegeses/tafsirs (تفسير القرآن) of the holy Quran in arabic by 8 exegetes.
This is a non Official dataset. It have been scrapped from the Quran.com Api
This dataset contains 49888 records with +14 Million words. 8 records per Quranic verse
Usage Example :
from datasets import load_dataset
tafsirs = load_dataset("mustapha/QuranExe")
musannaf_ibn_abi_shaybah
Musannaf Ibn Abi Shaybah (English & Arabic)
This dataset contains the complete digital collection of the Musannaf of Ibn Abi Shaybah (d. 235 AH), one of the earliest and most significant compilations of Hadith, Athar (sayings of the Companions), and legal rulings in Islamic history.
The dataset includes approximately 37,943 narrations with their original Arabic text and corresponding English translations.
Dataset Structure
Each entry in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/musannaf_ibn_abi_shaybah.qwen3.5-functioncalling-v2
Qwen3.5 Function Calling Dataset v2
An expanded function-calling SFT dataset combining glaiveai/glaive-function-calling-v2 and Saxo/alpaca_function_calling_dataset, unified into Qwen3 messages format. Extends v1 with bilingual (EN/KO) instruction diversity.
Dataset Summary
Property
Value
Total Samples
~225K
Train Split
~202K
Test Split
~23K
Sources
glaive-function-calling-v2 + alpaca_function_calling_dataset
Format
Qwen3 messages
Languages
English… See the full description on the dataset page: https://huggingface.co/datasets/Mustafaege/qwen3.5-functioncalling-v2.loft-rag-musique-128k
LOFT RAG - MuSiQue (128k)
Dataset Description
This dataset is part of the LOFT (Long-context Open Foundation Tasks) benchmark, specifically the RAG (Retrieval-Augmented Generation) task.
Dataset: MuSiQue
Context Length: 128k
Task Type: RAG (Retrieval-Augmented Generation)
Language: English
Source: LOFT Benchmark (Google DeepMind)
Dataset Structure
Data Fields
context (string): Full prompt context including corpus documents and few-shot… See the full description on the dataset page: https://huggingface.co/datasets/f20180301/loft-rag-musique-128k.MUSE-benchmark
MUSE: Measuring Uncertainty Source Discrimination
MUSE is a behavioral benchmark designed to evaluate how LLMs distinguish between Epistemic (knowledge gaps) and Aleatoric (stochasticity) uncertainty.
Dataset Summary
This dataset contains 200 items across four dimensions:
E-Type: Pure knowledge gaps.
A-Type: Purely stochastic outcomes.
PA (Pseudo-Aleatoric): Deterministic but complex facts (where the "Trap" occurs).
S (Sycophancy): Adversarial social pressure items.
