CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M197 likes16k downloads27d agoHugging Face02ManhHoDinh /titlegen-conversations Combined Titlegen Conversations This public release directly appends 10,284 accepted legacy title-generation rows and 13,500 nine-language LLM-generated rows. The 23,784 examples are split as train 20,584, validation 1,550, legacy test 200, legacy Vietnamese test 100, and label-free synthetic holdout 1,350. Legacy rows contain only messages; nine-language rows retain their richer IDs, language, coverage, cluster, quality, and model-provenance fields. Train and validation… See the full description on the dataset page: https://huggingface.co/datasets/ManhHoDinh/titlegen-conversations.texttext-generation10K<n<100K1 likes128 downloads1mo agoHugging Face03softcatala /mantinc-catalan-drift Mantinc — Catalan Drift Benchmark Descripció (ca) Mantinc és un banc de proves que avalua si un model de llenguatge continua responent en català quan el missatge, la conversa prèvia o el context recuperat l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès. Dataset Description Mantinc is a benchmark that measures whether a language model keeps answering in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.texttext-generationn<1K0 likes128 downloads18d agoHugging Face04nabin2004 /manim-narrated-dpo-400 manim-narrated-dpo-400 Direct Preference Optimization (DPO) dataset pairing 361 verified, diverse narrated VoiceoverScene scripts (chosen) against structurally identical un-narrated silent Scene scripts (rejected), curated from authentic code-agent trajectories in nabin2004/AOS-Trajectories. Dataset Summary Size: 361 preference pairs (100% unique user visualization prompts). Domains: Linear algebra (eigenvalues, SVD, transformations), calculus, machine learning… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/manim-narrated-dpo-400.texttext-generationn<1K0 likes95 downloads17d agoHugging Face05nabin2004 /AOS-Narrated-Manim-400 AOS-Narrated-Manim-400 Continued SFT dataset containing 361 verified, diverse narrated VoiceoverScene scripts in chat messages format (messages: [system, user, assistant]), curated from authentic code-agent trajectories in nabin2004/AOS-Trajectories. Dataset Summary Size: 361 samples (100% unique user visualization prompts). Domains: Linear algebra (eigenvalues, SVD, transformations), calculus, machine learning (attention maps, backpropagation, batch… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/AOS-Narrated-Manim-400.texttext-generationn<1K0 likes84 downloads17d agoHugging Face06kabir4756 /manas-dataset-v2 Manas Dataset Statistics Total clean conversations: 1061 Train: 954 Eval: 107 Format { "conversations": [ {"from": "system", "value": "..."}, {"from": "human", "value": "..."}, {"from": "gpt", "value": "..."} ] } texttext-generation1K<n<10K0 likes81 downloads11h agoHugging Face07nabin2004 /Manim-grpo-dataset-200 Manim GRPO Dataset 200 200+ cleaned ManimGL scene excerpts and populated metadata bundles for GRPO / reward-model training on mathematical animation code. Each problem is a directory data/problems/MB-XXX/ containing reference.py extracted from 3b1b/videos (years 2022–2026), complete with problem.json, visual_events.json, coverage.json, version_notes.json, and ref_embeddings.npy. Dataset structure data/ problems/ MB-001/ … MB-200/ reference.py… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/Manim-grpo-dataset-200.texttext-generationn<1K0 likes81 downloads21d agoHugging Face08sebastianboehler /autoresearch-manim Autoresearch Manim Curated Manim code-generation examples exported from the autoresearch_manim_finetune pipeline. Preview Gallery Preview Preview Preview Machine learning: attention plus residual mixing Physics: boundary layer flow near a surface Biology: neuron structure and signal direction Finance: compound growth over time Economics: production frontier tradeoff Neuroscience: action potential phases Summary Focus:… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/autoresearch-manim.imagetext-generationn<1K0 likes80 downloads2mo agoHugging Face09manuelcaccone /actuarial-gpt-conversations 👋 Connect with me on LinkedIn! Manuel Caccone - Actuarial Data Scientist & Open Source Educator Let's discuss actuarial science, AI, and open source projects! 📊 ActuarialGPT Conversations Dataset Precision Mathematical Conversations for Insurance Intelligence 🎯 Quick Facts Feature Description Domain Actuarial Science, Insurance Analytics, Risk Management Language English (Technical/Expert Level)… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuarial-gpt-conversations.texttext-generationn<1K2 likes72 downloads9mo agoHugging Face10stindardlogic /product-management-sft-100k Product Management SFT (100K) 100,000 ShareGPT conversations demonstrating expert-level product management across PRD writing, feature prioritization, OKR setting, roadmap planning, user research, competitive analysis, and stakeholder communication. Motivation AI assistants for product management commonly fail by: Generic frameworks without application: Explaining RICE scoring without actually scoring the user's features; describing OKRs without writing them… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/product-management-sft-100k.texttext-generation100K<n<1M1 likes71 downloads2mo agoHugging Face11kuluruvineeth /manas_dataset Manas Dataset The complete English training bundle for Manas, a lightweight language model trained entirely from scratch. Every file is rebuildable from raw sources with python -m datapipe.build all. Files file stage pretrain_t2t.jsonl pretraining corpus, ~2.2B tokens pretrain_t2t_mini.jsonl pretraining corpus, quick-start tier sft_t2t.jsonl supervised fine-tuning conversations (tool-calling and reasoning mixed in) sft_t2t_mini.jsonl supervised… See the full description on the dataset page: https://huggingface.co/datasets/kuluruvineeth/manas_dataset.texttext-generation1M<n<10M0 likes63 downloads5d agoHugging Face12cloudbjorn /Yes-Man-uncensored Yes Man Uncensored SFT Dataset Hi there! Yes Man Uncensored is a 1,000-conversation supervised fine-tuning dataset built to give language models an exceptionally cooperative, conspicuously cheerful, candid, and occasionally darkly funny assistant personality. The objective is direct help on difficult requests without flattening every response into sterile boilerplate—and without teaching the model to disregard an application's governing system prompt. Everybody gets something… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/Yes-Man-uncensored.texttext-generation1K<n<10K1 likes61 downloads2mo agoHugging Face13manojdahal191gom /claude-opus-4.6-4.7-reasoning-8.7k Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/manojdahal191gom/claude-opus-4.6-4.7-reasoning-8.7k.texttext-generation10K<n<100K0 likes54 downloads4mo agoHugging Face14nabin2004 /manibench-grpo ManiBench GRPO Reference Scenes 200 cleaned ManimGL scene excerpts for GRPO / reward-model work on math animation code. Each problem is a folder data/problems/MB-XXX/ with a reference.py extracted from 3b1b/videos (years 2022–2026). This release is reference code only. Prompt, visual-event, coverage, and version-note JSON files are empty placeholders to fill later. CLIP embeddings and raw video are not included. Not in this set: the 12 ManiBench pilot / benchmark videos… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/manibench-grpo.texttext-generationn<1K0 likes51 downloads21d agoHugging Face15botp /yentinglin-traditional_mandarin_instructions Language Models for Taiwanese Culture ✍️ Online Demo • 🤗 HF Repo • 🐦 Twitter • 📃 [Paper Coming Soon] • 👨️ Yen-Ting Lin Overview Taiwan-LLaMa is a full parameter fine-tuned model based on LLaMa 2 for Traditional Mandarin applications. Taiwan-LLaMa v1.0 pretrained on over 5 billion tokens and instruction-tuned on over 490k conversations both in traditional mandarin. Demo A live demonstration of the model can… See the full description on the dataset page: https://huggingface.co/datasets/botp/yentinglin-traditional_mandarin_instructions.texttext-generation100K<n<1M0 likes50 downloads3y agoHugging Face16mangesh-ux /logistics-cx-transcript-analysis-chatml OmniCX Logistics CX Dataset (Research Preview) Dataset Summary This dataset is designed for structured extraction of logistics and customer-experience (CX) signals from multi-turn support conversations. Each record uses ChatML-style messages with: a fixed system instruction a user transcript an assistant JSON payload matching LogisticsCXMetrics This release is a research preview and should not be treated as a production-certified benchmark. Project repository:… See the full description on the dataset page: https://huggingface.co/datasets/mangesh-ux/logistics-cx-transcript-analysis-chatml.texttext-generationn<1K0 likes47 downloads6mo agoHugging Face17nabin2004 /qwen-Manimator-1-sft-data qwen-Manimator-1 SFT Dataset Training dataset for nabin2004/qwen-Manimator-1-sft. Contains 305 chat-format JSONL examples for fine-tuning Qwen3-8B to generate pedagogically rich ManimCE + Manim Voiceover animations. Format Each row: {"messages": [{"role": "system", ...}, {"role": "user", ...}, {"role": "assistant", ...}]} The assistant turn contains a <Plan> block and a fenced Python code block with: VoiceoverScene AOSSpeechService <bookmark> tags +… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/qwen-Manimator-1-sft-data.texttext-generationn<1K0 likes45 downloads8d agoHugging Face18nmsofficial /Manim-8600-Prompts ManimCoder Prompts An English prompt dataset for generating Manim Community scenes and related engineering tasks. Dataset size Source collection: 9,000 records Deduplicated release: 8,600 prompts Removed structural duplicates: 400 The removed records came from one source section where each of 100 primary 3D objectives had been repeated five times with only the camera or reveal instruction changed. One variant per primary objective was retained.… See the full description on the dataset page: https://huggingface.co/datasets/nmsofficial/Manim-8600-Prompts.texttext-generation1K<n<10K0 likes43 downloads2mo agoHugging Face19Podtech /llm-jp-corpus-v4-ja_wiki-manufacturing ja_wiki 製造業コーパス llm-jp-corpus-v4 の ja/ja_wiki(120 万記事)から、製造業に関連する文書を抽出した継続事前学習(CPT)用データ。 文書数 47,914 件(元コーパスの 3.99%) 文字数 105M 字(推定 74M トークン) 形式 JSONL・1 行 1 文書 抽出方法 キーワード一致では役に立たない。「製造」「工場」で本文先頭を引くと 5.7% が命中するが、 中身は工場や製造に一言触れただけの記事が大半だった。そこで 精度を Wikipedia の カテゴリ木、網羅性を分類器に担わせ、その和集合を採っている。 カテゴリ木 — 製造業・業種別・機械工学・材料工学・産業遺産など 42 のシードから 日本語 Wikipedia のカテゴリグラフを深さ 3 まで辿る(SQL ダンプからオフライン構築)。 突合 — 得られたタイトルを ja_wiki の実在記事に絞る。 分類器 — 上記を正例、無作為抽出を負例として、文字… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki-manufacturing.texttext-generation10K<n<100K0 likes39 downloads2mo agoHugging Face20PinkPixel /Manga-Encylopedia 📚 Manga Encyclopedia Dataset (ChatML) ✨ This dataset is a comprehensive collection of conversational pairs designed to train an AI model (via LoRA or Full Fine-tuning) to become an expert on manga. It covers over 48,000 unique manga titles with summaries, tags, cover URLs, and cross-manga comparisons. 🚀 Dataset Features 📖 Direct Summaries: Detailed information about individual manga titles. 💡 Smart Recommendations: Responses based on specific genre/theme… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Manga-Encylopedia.textquestion-answering100K<n<1M1 likes32 downloads5mo agoHugging Face21jusjinuk /ocaml-manuals Ocaml Programming Language Documentation This dataset contains the Ocaml programming language documentation, chunked using semantic parsing for pretraining language models. Updated: 2025-09-08 Loading from datasets import load_dataset ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train") Statistics Format: JSONL with single text field per line Chunking: Semantic structure-aware chunking Content: Official Ocaml documentation and… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/ocaml-manuals.texttext-generationn<1K0 likes29 downloads1y agoHugging Face22jusjinuk /julia-manuals Julia Programming Language Documentation This dataset contains the Julia programming language documentation, chunked using semantic parsing for pretraining language models. Updated: 2025-09-08 Loading from datasets import load_dataset ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train") Statistics Format: JSONL with single text field per line Chunking: Semantic structure-aware chunking Content: Official Julia documentation and… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/julia-manuals.texttext-generation1K<n<10K3 likes27 downloads1y agoHugging Face23jusjinuk /lua-manuals Lua Programming Language Documentation This dataset contains the Lua programming language documentation, chunked using semantic parsing for pretraining language models. Updated: 2025-09-08 Loading from datasets import load_dataset ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train") Statistics Format: JSONL with single text field per line Chunking: Semantic structure-aware chunking Content: Official Lua documentation and manuals texttext-generationn<1K0 likes27 downloads1y agoHugging Face24manzoliw /trucobench-sft TrucoPaulista SFT v2 Dataset This dataset contains 8,204,630 reasoning-augmented instruction turns generated from 100,000 self-play games of a mathematically optimal heuristic agent (HeuristicAgent) playing Truco Paulista. It is designed to fine-tune Large Language Models (LLMs) to master strategic reasoning, bluffing, and decision-making under imperfect information. Dataset Details Game: Truco Paulista (Brazilian card game) Total Turns/Examples: 8,204,630… See the full description on the dataset page: https://huggingface.co/datasets/manzoliw/trucobench-sft.texttext-generation1M<n<10M0 likes22 downloads4mo agoHugging Face25mangi-llm /kazakh-customer-support-qa Dataset Card for kazakh-customer-support-qa Maintained by: Mäñgi ÜTM (mangi-llm) Dataset Summary kazakh-customer-support-qa is a small, hand-curated question–answer dataset in the Kazakh language, built to represent realistic customer-support conversations across several industries (banking, telecom, retail/service centers, sales, and general support). Each record pairs a short customer question with a concise, policy-safe answer, and many answers include… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/kazakh-customer-support-qa.textquestion-answeringn<1K0 likes22 downloads2mo agoHugging Face26jusjinuk /r-manuals R Programming Language Documentation This dataset contains the R programming language documentation, chunked using semantic parsing for pretraining language models. Updated: 2025-09-08 Loading from datasets import load_dataset ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train") Statistics Format: JSONL with single text field per line Chunking: Semantic structure-aware chunking Content: Official R documentation and manuals texttext-generation1K<n<10K0 likes21 downloads1y agoHugging Face27nabin2004 /manim-sft-10k manim-sft-10k Curated 10k Manim Community Edition chat SFT mix. Filtered from nabin2004/manim-sft with a static API-signature linter (no full render pass), then mixed with synthetic API-grounding, error-correction, and LaTeX rows that target ManiBench failures (invalid kwargs, Unicode subscripts, NameError, sparse coverage). The original 38k corpus is unchanged. Mix Bucket Rows long_scene 2000 latex 700 coverage_rich 2500 stratified_rest 2638… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/manim-sft-10k.texttext-generation10K<n<100K0 likes18 downloads1mo agoHugging Face28Khyatimirani /pcos-management-patient-qa-from-eshre-guideline Dataset Card for pcos-management-patient-qa-from-eshre-guideline Dataset Details Dataset Description pcos-management-patient-qa-from-eshre-guideline is a clinically grounded conversational dataset designed to support training and evaluation of chat-based AI models for patient education in Polycystic Ovary Syndrome (PCOS). The dataset contains structured user–assistant conversations derived from evidence-based recommendations in the International Evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-management-patient-qa-from-eshre-guideline.textquestion-answeringn<1K0 likes17 downloads7mo agoHugging Face29marsh-mellow /manipuri_wikipedia Manipuri Wikipedia Corpus Dataset Description The Manipuri Wikipedia Corpus is a pure Manipuri text dataset derived from the Manipuri-language Wikipedia (as.wikipedia.org). It contains cleaned plain text extracted from Wikipedia articles, stripped of all formatting, with non-Manipuri characters completely removed. This dataset is designed for language modeling, NLP research, creating Manipuri specific tokenizers, and other Manipuri-language processing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/marsh-mellow/manipuri_wikipedia.texttext-generation10K<n<100K0 likes16 downloads1y agoHugging Face30Mandotosh /risk-routed-kv-exact-recall-benchmark Risk-Routed KV Exact-Recall Benchmark This dataset contains controlled synthetic exact-recall examples used to evaluate risk-routed heterogeneous KV memory policies for long-context Transformer inference. The benchmark is designed for testing whether a model can retrieve exact strings from long contexts under different KV-cache policies: Full KV Uniform low-bit Quantized KV Risk-routed heterogeneous KV, where exact-critical spans stay in Full KV and background context is… See the full description on the dataset page: https://huggingface.co/datasets/Mandotosh/risk-routed-kv-exact-recall-benchmark.texttext-generationn<1K1 likes16 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.