datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.titlegen-conversations
Combined Titlegen Conversations
This public release directly appends 10,284 accepted legacy title-generation
rows and 13,500 nine-language LLM-generated rows. The 23,784 examples are split
as train 20,584, validation 1,550, legacy test 200, legacy Vietnamese test 100,
and label-free synthetic holdout 1,350. Legacy rows contain only messages;
nine-language rows retain their richer IDs, language, coverage, cluster,
quality, and model-provenance fields. Train and validation… See the full description on the dataset page: https://huggingface.co/datasets/ManhHoDinh/titlegen-conversations.mantinc-catalan-drift
Mantinc — Catalan Drift Benchmark
Descripció (ca)
Mantinc és un banc de proves que avalua si un model de llenguatge continua
responent en català quan el missatge, la conversa prèvia o el context recuperat
l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès.
Dataset Description
Mantinc is a benchmark that measures whether a language model keeps answering
in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.manim-narrated-dpo-400
manim-narrated-dpo-400
Direct Preference Optimization (DPO) dataset pairing 361 verified, diverse narrated VoiceoverScene scripts (chosen) against structurally identical un-narrated silent Scene scripts (rejected), curated from authentic code-agent trajectories in nabin2004/AOS-Trajectories.
Dataset Summary
Size: 361 preference pairs (100% unique user visualization prompts).
Domains: Linear algebra (eigenvalues, SVD, transformations), calculus, machine learning… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/manim-narrated-dpo-400.AOS-Narrated-Manim-400
AOS-Narrated-Manim-400
Continued SFT dataset containing 361 verified, diverse narrated VoiceoverScene scripts in chat messages format (messages: [system, user, assistant]), curated from authentic code-agent trajectories in nabin2004/AOS-Trajectories.
Dataset Summary
Size: 361 samples (100% unique user visualization prompts).
Domains: Linear algebra (eigenvalues, SVD, transformations), calculus, machine learning (attention maps, backpropagation, batch… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/AOS-Narrated-Manim-400.manas-dataset-v2
Manas Dataset
Statistics
Total clean conversations: 1061
Train: 954
Eval: 107
Format
{
"conversations": [
{"from": "system", "value": "..."},
{"from": "human", "value": "..."},
{"from": "gpt", "value": "..."}
]
}
Manim-grpo-dataset-200
Manim GRPO Dataset 200
200+ cleaned ManimGL scene excerpts and populated metadata bundles for GRPO / reward-model training on mathematical animation code. Each problem is a directory data/problems/MB-XXX/ containing reference.py extracted from 3b1b/videos (years 2022–2026), complete with problem.json, visual_events.json, coverage.json, version_notes.json, and ref_embeddings.npy.
Dataset structure
data/
problems/
MB-001/ … MB-200/
reference.py… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/Manim-grpo-dataset-200.autoresearch-manim
Autoresearch Manim
Curated Manim code-generation examples exported from the autoresearch_manim_finetune pipeline.
Preview Gallery
Preview
Preview
Preview
Machine learning: attention plus residual mixing
Physics: boundary layer flow near a surface
Biology: neuron structure and signal direction
Finance: compound growth over time
Economics: production frontier tradeoff
Neuroscience: action potential phases
Summary
Focus:… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/autoresearch-manim.actuarial-gpt-conversations
👋 Connect with me on LinkedIn!
Manuel Caccone - Actuarial Data Scientist & Open Source Educator
Let's discuss actuarial science, AI, and open source projects!
📊 ActuarialGPT Conversations Dataset
Precision Mathematical Conversations for Insurance Intelligence
🎯 Quick Facts
Feature
Description
Domain
Actuarial Science, Insurance Analytics, Risk Management
Language
English (Technical/Expert Level)… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuarial-gpt-conversations.product-management-sft-100k
Product Management SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level product management across PRD writing, feature prioritization, OKR setting, roadmap planning, user research, competitive analysis, and stakeholder communication.
Motivation
AI assistants for product management commonly fail by:
Generic frameworks without application: Explaining RICE scoring without actually scoring the user's features; describing OKRs without writing them… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/product-management-sft-100k.manas_dataset
Manas Dataset
The complete English training bundle for Manas, a lightweight language model trained entirely from scratch. Every file is rebuildable from raw sources with python -m datapipe.build all.
Files
file
stage
pretrain_t2t.jsonl
pretraining corpus, ~2.2B tokens
pretrain_t2t_mini.jsonl
pretraining corpus, quick-start tier
sft_t2t.jsonl
supervised fine-tuning conversations (tool-calling and reasoning mixed in)
sft_t2t_mini.jsonl
supervised… See the full description on the dataset page: https://huggingface.co/datasets/kuluruvineeth/manas_dataset.Yes-Man-uncensored
Yes Man Uncensored SFT Dataset
Hi there! Yes Man Uncensored is a 1,000-conversation supervised fine-tuning
dataset built to give language models an exceptionally cooperative, conspicuously
cheerful, candid, and occasionally darkly funny assistant personality. The objective
is direct help on difficult requests without flattening every response into sterile
boilerplate—and without teaching the model to disregard an application's governing
system prompt. Everybody gets something… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/Yes-Man-uncensored.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/manojdahal191gom/claude-opus-4.6-4.7-reasoning-8.7k.manibench-grpo
ManiBench GRPO Reference Scenes
200 cleaned ManimGL scene excerpts for GRPO / reward-model work on math animation code. Each problem is a folder data/problems/MB-XXX/ with a reference.py extracted from 3b1b/videos (years 2022–2026).
This release is reference code only. Prompt, visual-event, coverage, and version-note JSON files are empty placeholders to fill later. CLIP embeddings and raw video are not included.
Not in this set: the 12 ManiBench pilot / benchmark videos… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/manibench-grpo.yentinglin-traditional_mandarin_instructions
Language Models for Taiwanese Culture
✍️ Online Demo
•
🤗 HF Repo • 🐦 Twitter • 📃 [Paper Coming Soon]
• 👨️ Yen-Ting Lin
Overview
Taiwan-LLaMa is a full parameter fine-tuned model based on LLaMa 2 for Traditional Mandarin applications.
Taiwan-LLaMa v1.0 pretrained on over 5 billion tokens and instruction-tuned on over 490k conversations both in traditional mandarin.
Demo
A live demonstration of the model can… See the full description on the dataset page: https://huggingface.co/datasets/botp/yentinglin-traditional_mandarin_instructions.logistics-cx-transcript-analysis-chatml
OmniCX Logistics CX Dataset (Research Preview)
Dataset Summary
This dataset is designed for structured extraction of logistics and customer-experience (CX) signals from multi-turn support conversations.
Each record uses ChatML-style messages with:
a fixed system instruction
a user transcript
an assistant JSON payload matching LogisticsCXMetrics
This release is a research preview and should not be treated as a production-certified benchmark.
Project repository:… See the full description on the dataset page: https://huggingface.co/datasets/mangesh-ux/logistics-cx-transcript-analysis-chatml.qwen-Manimator-1-sft-data
qwen-Manimator-1 SFT Dataset
Training dataset for nabin2004/qwen-Manimator-1-sft.
Contains 305 chat-format JSONL examples for fine-tuning Qwen3-8B
to generate pedagogically rich ManimCE + Manim Voiceover animations.
Format
Each row: {"messages": [{"role": "system", ...}, {"role": "user", ...}, {"role": "assistant", ...}]}
The assistant turn contains a <Plan> block and a fenced Python code block with:
VoiceoverScene
AOSSpeechService
<bookmark> tags +… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/qwen-Manimator-1-sft-data.Manim-8600-Prompts
ManimCoder Prompts
An English prompt dataset for generating Manim Community scenes and related engineering tasks.
Dataset size
Source collection: 9,000 records
Deduplicated release: 8,600 prompts
Removed structural duplicates: 400
The removed records came from one source section where each of 100 primary 3D objectives had been repeated five times with only the camera or reveal instruction changed. One variant per primary objective was retained.… See the full description on the dataset page: https://huggingface.co/datasets/nmsofficial/Manim-8600-Prompts.llm-jp-corpus-v4-ja_wiki-manufacturing
ja_wiki 製造業コーパス
llm-jp-corpus-v4 の
ja/ja_wiki(120 万記事)から、製造業に関連する文書を抽出した継続事前学習(CPT)用データ。
文書数
47,914 件(元コーパスの 3.99%)
文字数
105M 字(推定 74M トークン)
形式
JSONL・1 行 1 文書
抽出方法
キーワード一致では役に立たない。「製造」「工場」で本文先頭を引くと 5.7% が命中するが、
中身は工場や製造に一言触れただけの記事が大半だった。そこで 精度を Wikipedia の
カテゴリ木、網羅性を分類器に担わせ、その和集合を採っている。
カテゴリ木 — 製造業・業種別・機械工学・材料工学・産業遺産など 42 のシードから
日本語 Wikipedia のカテゴリグラフを深さ 3 まで辿る(SQL ダンプからオフライン構築)。
突合 — 得られたタイトルを ja_wiki の実在記事に絞る。
分類器 — 上記を正例、無作為抽出を負例として、文字… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki-manufacturing.Manga-Encylopedia
📚 Manga Encyclopedia Dataset (ChatML) ✨
This dataset is a comprehensive collection of conversational pairs designed to train an AI model (via LoRA or Full Fine-tuning) to become an expert on manga. It covers over 48,000 unique manga titles with summaries, tags, cover URLs, and cross-manga comparisons.
🚀 Dataset Features
📖 Direct Summaries: Detailed information about individual manga titles.
💡 Smart Recommendations: Responses based on specific genre/theme… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Manga-Encylopedia.ocaml-manuals
Ocaml Programming Language Documentation
This dataset contains the Ocaml programming language documentation,
chunked using semantic parsing for pretraining language models.
Updated: 2025-09-08
Loading
from datasets import load_dataset
ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train")
Statistics
Format: JSONL with single text field per line
Chunking: Semantic structure-aware chunking
Content: Official Ocaml documentation and… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/ocaml-manuals.julia-manuals
Julia Programming Language Documentation
This dataset contains the Julia programming language documentation,
chunked using semantic parsing for pretraining language models.
Updated: 2025-09-08
Loading
from datasets import load_dataset
ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train")
Statistics
Format: JSONL with single text field per line
Chunking: Semantic structure-aware chunking
Content: Official Julia documentation and… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/julia-manuals.lua-manuals
Lua Programming Language Documentation
This dataset contains the Lua programming language documentation,
chunked using semantic parsing for pretraining language models.
Updated: 2025-09-08
Loading
from datasets import load_dataset
ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train")
Statistics
Format: JSONL with single text field per line
Chunking: Semantic structure-aware chunking
Content: Official Lua documentation and manuals
trucobench-sft
TrucoPaulista SFT v2 Dataset
This dataset contains 8,204,630 reasoning-augmented instruction turns generated from 100,000 self-play games of a mathematically optimal heuristic agent (HeuristicAgent) playing Truco Paulista. It is designed to fine-tune Large Language Models (LLMs) to master strategic reasoning, bluffing, and decision-making under imperfect information.
Dataset Details
Game: Truco Paulista (Brazilian card game)
Total Turns/Examples: 8,204,630… See the full description on the dataset page: https://huggingface.co/datasets/manzoliw/trucobench-sft.kazakh-customer-support-qa
Dataset Card for kazakh-customer-support-qa
Maintained by: Mäñgi ÜTM (mangi-llm)
Dataset Summary
kazakh-customer-support-qa is a small, hand-curated question–answer dataset in the Kazakh language, built to represent realistic customer-support conversations across several industries (banking, telecom, retail/service centers, sales, and general support). Each record pairs a short customer question with a concise, policy-safe answer, and many answers include… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/kazakh-customer-support-qa.r-manuals
R Programming Language Documentation
This dataset contains the R programming language documentation,
chunked using semantic parsing for pretraining language models.
Updated: 2025-09-08
Loading
from datasets import load_dataset
ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train")
Statistics
Format: JSONL with single text field per line
Chunking: Semantic structure-aware chunking
Content: Official R documentation and manuals
manim-sft-10k
manim-sft-10k
Curated 10k Manim Community Edition chat SFT mix. Filtered from nabin2004/manim-sft with a static API-signature linter (no full render pass), then mixed with synthetic API-grounding, error-correction, and LaTeX rows that target ManiBench failures (invalid kwargs, Unicode subscripts, NameError, sparse coverage).
The original 38k corpus is unchanged.
Mix
Bucket
Rows
long_scene
2000
latex
700
coverage_rich
2500
stratified_rest
2638… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/manim-sft-10k.pcos-management-patient-qa-from-eshre-guideline
Dataset Card for pcos-management-patient-qa-from-eshre-guideline
Dataset Details
Dataset Description
pcos-management-patient-qa-from-eshre-guideline is a clinically grounded conversational dataset designed to support training and evaluation of chat-based AI models for patient education in Polycystic Ovary Syndrome (PCOS).
The dataset contains structured user–assistant conversations derived from evidence-based recommendations in the International Evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-management-patient-qa-from-eshre-guideline.manipuri_wikipedia
Manipuri Wikipedia Corpus
Dataset Description
The Manipuri Wikipedia Corpus is a pure Manipuri text dataset derived from the Manipuri-language Wikipedia (as.wikipedia.org).
It contains cleaned plain text extracted from Wikipedia articles, stripped of all formatting, with non-Manipuri characters completely removed.
This dataset is designed for language modeling, NLP research, creating Manipuri specific tokenizers, and other Manipuri-language processing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/marsh-mellow/manipuri_wikipedia.risk-routed-kv-exact-recall-benchmark
Risk-Routed KV Exact-Recall Benchmark
This dataset contains controlled synthetic exact-recall examples used to evaluate risk-routed heterogeneous KV memory policies for long-context Transformer inference.
The benchmark is designed for testing whether a model can retrieve exact strings from long contexts under different KV-cache policies:
Full KV
Uniform low-bit Quantized KV
Risk-routed heterogeneous KV, where exact-critical spans stay in Full KV and background context is… See the full description on the dataset page: https://huggingface.co/datasets/Mandotosh/risk-routed-kv-exact-recall-benchmark.
