CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01muset-ai /DeepResearch-Bench-II-Datasetdocumenttext-generationn<1K2 likes1.4k downloads7mo agoHugging Face02m-a-p /MusicPile🌐 DemoPage | 🤗SFT Dataset | 🤗 Benchmark | 📖 arXiv | 💻 Code | 🤖 Chat Model | 🤖 Base Model Dataset Card for MusicPile MusicPile is the first pretraining corpus for developing musical abilities in large language models. It has 5.17M samples and approximately 4.16B tokens, including web-crawled corpora, encyclopedias, music books, youtube music captions, musical pieces in abc notation, math content, and code. You can easily load it:from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MusicPile.texttext-generation1M<n<10M59 likes1.1k downloads2y agoHugging Face03MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes550 downloads10mo agoHugging Face04muset-ai /DeepResearch-Bench-Dataset DeepResearch Bench Dataset [English | 中文] English 📖 Dataset Overview This is the official dataset accompanying the DeepResearch Bench paper. It contains research reports generated by 4 leading deep research AI systems along with detailed human expert annotations evaluating these reports. DeepResearch Bench is the first comprehensive benchmark for systematically evaluating Deep Research Agents (DRAs) on their ability to handle complex, PhD-level research… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-Dataset.text-generationn<1K10 likes377 downloads10mo agoHugging Face05takiuddinahmed /muslim-names-dataset Muslim Names Dataset A comprehensive collection of Muslim names with meanings scraped from muslimnames.com. Contains 14,585 names with English names, Arabic names, meanings, and gender classifications. Dataset Contents This dataset contains ~14,585 Muslim names with the following information: English name: Name in English/Latin script Arabic name: Name in Arabic script Meaning: Definition and meaning of the name Gender: Classification as male or female Files… See the full description on the dataset page: https://huggingface.co/datasets/takiuddinahmed/muslim-names-dataset.texttext-classification10K<n<100K3 likes344 downloads1y agoHugging Face06Mustafaege /qwen3.5-toolcalling-v2 Qwen3.5 Tool Calling Dataset v2 An expanded tool-calling SFT dataset combining smirki/Tool-Calling-Dataset-UIGEN-X and AmanPriyanshu/tool-reasoning-sft-jupyter-agent, unified into Qwen3 messages format. Adds Jupyter notebook agent data with code execution reasoning chains. Dataset Summary Property Value Total Samples ~60K+ Train Split ~55K Test Split ~6K Sources UIGEN-X + Jupyter Agent Format Qwen3 messages Language English License Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/Mustafaege/qwen3.5-toolcalling-v2.texttext-generation100K<n<1M49 likes343 downloads7mo agoHugging Face07musumecmtcd /MathNet Shaden Alshammari1*   Kevin Wen1*   Abrar Zainal3*   Mark Hamilton1 Navid Safaei4   Sultan Albarakati2   William T. Freeman1†   Antonio Torralba1† 1MIT   2KAUST   3HUMAIN   4Bulgarian Academy of Sciences   *† equal contribution Quick Start· Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation Note: This is a test for the HF hosting website. The dataset isn’t fully uploaded yet; it will be uploaded on Tuesday, April 21, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/musumecmtcd/MathNet.imagequestion-answeringn<1K0 likes326 downloads5mo agoHugging Face08Satgoy152 /Muse-Glimmer-SWE-Gym-2k Muse-Glimmer-SWE-Gym-2k Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them. Configs Config Rows Size What it is train 1,981 57 MB One row per trajectory: the full conversation as messages. raw 159,999 2.7 GB One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.tabulartext-generation100K<n<1M2 likes254 downloads22d agoHugging Face09mustafabasar /model-egitme-sft-v1 Türkçe İK Belge Zekâsı — SFT verisi ve LoRA adapterleri Türkçe bir insan kaynakları bildirimini 29 alanlı bir JSON sözleşmesine çeviren küçük bir modelin eğitim verisi ve adapterleri. Canlı demo: https://huggingface.co/spaces/mustafabasar/ik-belge-asistani-demo Ne var burada yol ne veri_v3/ … veri_v9/ sürümlenmiş eğitim/doğrulama verisi (train.jsonl, dev.jsonl) v34/ … v45/ LoRA adapterleri (lora/) ve eğitim kayıtları sema/ Pydantic sözleşmesi… See the full description on the dataset page: https://huggingface.co/datasets/mustafabasar/model-egitme-sft-v1.text-generation0 likes236 downloads26d agoHugging Face10r0b0tlab /muse12-nemo-agentic Muse Spark 1.2 High-Reasoning NeMo Agentic Dataset A reproducible, verified 24,000-row synthetic agentic dataset generated with Meta Muse Spark 1.2, NeMo Gym, and deterministic task-family verifiers. The project is a quality-focused successor to r0b0tlab/deepseek-v4-pro-0813-agentic. It keeps rollout prompts separate from reference trajectories and offline-training views, records usage and provenance, and does not publish private chain-of-thought. [!IMPORTANT] Status:… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/muse12-nemo-agentic.text-generation10K<n<100K1 likes225 downloads1mo agoHugging Face11hugfaceguy0001 /music-codes 互联网歌曲数据集 使用encodec编码的歌曲数据,sampling_rate=24000, bandwidth=6.0, n_codebooks=8, 截取时长=30s 各数据列说明 id : 歌曲源id name : 歌曲名 singer : 歌手名 text : 歌词 code : 原音乐使用EnCodec编码的结果, 是形状为[T,8]的二维列表。 texttext-generation100K<n<1M0 likes188 downloads9mo agoHugging Face12museado /smithsonian-data Smithsonian Open Access Data Pre-processed data dumps from the Smithsonian Open Access initiative, covering millions of objects across Smithsonian Institution museums and archives. What is this? The Smithsonian publishes their Open Access metadata on S3, but the raw data is split across 255 individual .txt files per unit. This dataset consolidates each unit's data into a single .jsonl.gz file for easier downloading and processing. Files Each file corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/museado/smithsonian-data.feature-extraction10M<n<100M0 likes183 downloads10mo agoHugging Face13qurancn /MuslimLife Muslim Life Knowledge Base & RAG Dataset Contains 90 public Simplified Chinese articles from the Salaam Alykum 穆斯林生活 / Muslim Life topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and a lightweight embedding preview layer. [!TIP] Human Readers / 普通读者: For normal reading, open Files and versions -> content and start with content/README.md. Example article: 2744 莱麦丹不同面貌:斋月中的人、故事与信仰现场. For… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/MuslimLife.tabulartext-retrievaln<1K0 likes158 downloads2mo agoHugging Face14musicakamusic /crimsonred-paper-replication CrimsonRed — Cross-Architecture Emotion-Prime Steering Replication Replication of the emotion-prime steering protocol from arXiv:2607.18691 (NSM semantic primes as explanans for emotion in LLMs), extended across four architectures. Generated by scripts/paper_faithful_steering.py in the CrimsonRed project. The finding The paper's core claim — that semantic-prime recipe directions steer emotion more strongly than Scherer appraisal directions — replicates… See the full description on the dataset page: https://huggingface.co/datasets/musicakamusic/crimsonred-paper-replication.text-generation0 likes147 downloads11d agoHugging Face15Muse-Ltd /UncertaintyGym UncertaintyGym A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression Abstract UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating. Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.textquestion-answering1K<n<10K6 likes130 downloads1mo agoHugging Face16musabg /wikipedia-tr 📖 Türkçe Vikipedi Mayıs 2023 Bu veri kümesi, Türkçe Vikipedi'den alınan makalelerin bir derlemesi olup, maskeleme dil modelleme ve metin oluşturma görevleri için tasarlanmıştır. 🗣️ Etiketlemeler Bu veri kümesindeki makaleler, özellikle belirli bir görev için etiketlenmemiş olup, veri kümesi etiketsizdir. 🌐 Dil Bu veri kümesi Türkçe yazılmış olup, gönüllülerden oluşan bir ekip tarafından topluluk katılımı yöntemleri ile oluşturulmuştur. 📜 Lisans… See the full description on the dataset page: https://huggingface.co/datasets/musabg/wikipedia-tr.textfill-mask100K<n<1M19 likes110 downloads3y agoHugging Face17qurancn /Chinese-Muslim-Travel ☪ Chinese-Muslim-Travel: Native Chinese Muslim Travel RAG Corpus [!TIP] Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively! Dataset Description Chinese-Muslim-Travel is a curated RAG corpus containing 347 native Chinese articles documenting Muslim travel, halal food, mosque architecture, and Muslim community life across 20+ countries. Every… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Chinese-Muslim-Travel.tabulartext-generationn<1K0 likes107 downloads3mo agoHugging Face18DaoCloud /Muse-Glimmer-OPB-100K Muse Glimmer OPB 100K On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark. Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B. The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths. Reasoning strength Conversations Train-turn rows low 64,997 96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.texttext-generation100K<n<1M4 likes98 downloads1mo agoHugging Face19Mustafaege /qwen3.5-toolcalling-v1 Qwen3.5 Tool Calling Dataset v1 A tool-calling SFT dataset built from smirki/Tool-Calling-Dataset-UIGEN-X (a cleaned version of interstellarninja/hermes_reasoning_tool_use), converted from ShareGPT conversations format to Qwen3 messages format. Features deep reasoning chains with <think> tags followed by structured tool calls. Dataset Summary Property Value Total Samples 51,004 Train Split 45,904 Test Split 5,100 Source smirki/Tool-Calling-Dataset-UIGEN-X… See the full description on the dataset page: https://huggingface.co/datasets/Mustafaege/qwen3.5-toolcalling-v1.texttext-generation10K<n<100K1 likes85 downloads7mo agoHugging Face20muset-ai /PALATE PALATE Dataset PALATE contains de-identified human–role-playing-agent conversations, satisfaction annotations, frozen session-level splits, bilingual character cards, and the scoring rubrics used by the PALATE benchmark. Related resources: Code: Zhuyh1139/PALATE Five user-simulator adapters: muset-ai/PALATE-LoRA The dataset stores source annotations rather than ready-to-train examples. Use the processing command in the PALATE GitHub repository to construct role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.tabulartext-generationn<1K1 likes79 downloads2mo agoHugging Face21Muse-research /ifparse-v1.0 ifparse v1.0 IFParse is a benchmark for structured extraction from real developer logs: can a model turn a raw, unstructured log line into JSON that satisfies a fixed schema, with no code fence, no commentary, and no type errors. Scoring is binary and covers only compliance, not extraction creativity, so the signal is isolated to whether the output would actually parse in a production pipeline. Each prompt gives the model a single raw log line, either an Apache access log record… See the full description on the dataset page: https://huggingface.co/datasets/Muse-research/ifparse-v1.0.text-generation1K<n<10K4 likes77 downloads3mo agoHugging Face22qurancn /Hui-Muslims Hui Muslims RAG Dataset Contains 232 native Chinese articles. [!TIP] Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively! tabulartext-generationn<1K0 likes76 downloads3mo agoHugging Face23vhands /audio-music-mir-post-public audio-music-mir-post-public Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.textaudio-classification100K<n<1M0 likes74 downloads3mo agoHugging Face24muset-ai /Wiki_Live_Challenge Wiki Live Challenge Dataset [English | 中文] English 📖 Dataset Overview This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems. Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.tabulartext-generationn<1K1 likes73 downloads8mo agoHugging Face25f20180301 /loft-rag-musique-32k LOFT RAG - MuSiQue (32k) Dataset Description This dataset is part of the LOFT (Long-context Open Foundation Tasks) benchmark, specifically the RAG (Retrieval-Augmented Generation) task. Dataset: MuSiQue Context Length: 32k Task Type: RAG (Retrieval-Augmented Generation) Language: English Source: LOFT Benchmark (Google DeepMind) Dataset Structure Data Fields context (string): Full prompt context including corpus documents and few-shot… See the full description on the dataset page: https://huggingface.co/datasets/f20180301/loft-rag-musique-32k.textquestion-answeringn<1K0 likes71 downloads10mo agoHugging Face26mustapha /QuranExeThis dataset contains the exegeses/tafsirs (تفسير القرآن) of the holy Quran in arabic by 8 exegetes. This is a non Official dataset. It have been scrapped from the Quran.com Api This dataset contains 49888 records with +14 Million words. 8 records per Quranic verse Usage Example : from datasets import load_dataset tafsirs = load_dataset("mustapha/QuranExe") texttext-generation10K<n<100K11 likes70 downloads4y agoHugging Face27freococo /musannaf_ibn_abi_shaybah Musannaf Ibn Abi Shaybah (English & Arabic) This dataset contains the complete digital collection of the Musannaf of Ibn Abi Shaybah (d. 235 AH), one of the earliest and most significant compilations of Hadith, Athar (sayings of the Companions), and legal rulings in Islamic history. The dataset includes approximately 37,943 narrations with their original Arabic text and corresponding English translations. Dataset Structure Each entry in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/musannaf_ibn_abi_shaybah.texttranslation10K<n<100K0 likes65 downloads8mo agoHugging Face28Mustafaege /qwen3.5-functioncalling-v2 Qwen3.5 Function Calling Dataset v2 An expanded function-calling SFT dataset combining glaiveai/glaive-function-calling-v2 and Saxo/alpaca_function_calling_dataset, unified into Qwen3 messages format. Extends v1 with bilingual (EN/KO) instruction diversity. Dataset Summary Property Value Total Samples ~225K Train Split ~202K Test Split ~23K Sources glaive-function-calling-v2 + alpaca_function_calling_dataset Format Qwen3 messages Languages English… See the full description on the dataset page: https://huggingface.co/datasets/Mustafaege/qwen3.5-functioncalling-v2.texttext-generation100K<n<1M0 likes61 downloads7mo agoHugging Face29f20180301 /loft-rag-musique-128k LOFT RAG - MuSiQue (128k) Dataset Description This dataset is part of the LOFT (Long-context Open Foundation Tasks) benchmark, specifically the RAG (Retrieval-Augmented Generation) task. Dataset: MuSiQue Context Length: 128k Task Type: RAG (Retrieval-Augmented Generation) Language: English Source: LOFT Benchmark (Google DeepMind) Dataset Structure Data Fields context (string): Full prompt context including corpus documents and few-shot… See the full description on the dataset page: https://huggingface.co/datasets/f20180301/loft-rag-musique-128k.textquestion-answeringn<1K0 likes60 downloads10mo agoHugging Face30Ppilot2 /MUSE-benchmark MUSE: Measuring Uncertainty Source Discrimination MUSE is a behavioral benchmark designed to evaluate how LLMs distinguish between Epistemic (knowledge gaps) and Aleatoric (stochasticity) uncertainty. Dataset Summary This dataset contains 200 items across four dimensions: E-Type: Pure knowledge gaps. A-Type: Purely stochastic outcomes. PA (Pseudo-Aleatoric): Deterministic but complex facts (where the "Trap" occurs). S (Sycophancy): Adversarial social pressure items. question-answering0 likes59 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.