datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cqa-creative-writing-expert-cot-preview
CQA: Creative Quality Alignment — Research-Grade Schema v2
English
This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.qa-expert-multi-hop-qa-V1.0
Dataset Card for QA-Expert-multi-hop-qa-V1.0
This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering.
In total, this dataset contains 25.5k for training and 3.19k for evaluation.
You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0
The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.Expert-Go-SFT-100K
Expert-Go-SFT-100K
Paper | Code
Expert-Go-SFT-100K is a large-scale synthetic dataset designed to "cold start" Large Language Models (LLMs) for Go-related reasoning tasks. It was introduced as part of the LoGos project, which aims to bridge the gap between general-purpose LLM reasoning and specialized expert knowledge in the game of Go.
The dataset features 100,000 samples of structured Go expertise mixed with general long Chain-of-Thought (CoT) reasoning data. It enables models to… See the full description on the dataset page: https://huggingface.co/datasets/YichuanMa/Expert-Go-SFT-100K.gomodel-go-expert-v4
GoModel Go Expert v4 Dataset
Description
A high-quality dataset for fine-tuning Qwen2.5-Coder-7B to be an expert Go software engineer
with tool-calling capabilities. This is version 4, substantially rebuilt from v3 with:
Structured messages format (not pre-rendered ChatML text)
Go AST-extracted code from real repositories using go/parser
Go 1.26 feature coverage (February 2026 release)
Senior/staff-level engineering content (architecture, distributed systems, API… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v4.distill-expert-535k
Distill Expert 535k
Training dataset for the distill-expert model — a 0.6B LoRA fine-tuned from Qwen3-0.6B that compresses shell/command output for AI consumption.
Contents
train.jsonl.gz — 454,710 training examples (85%)
valid.jsonl.gz — 53,458 validation examples (10%)
test.jsonl.gz — 26,832 test examples (5%)
runpod_train.py — Unsloth LoRA training script (RunPod-ready)
Total: 535,000 examples across 8 operation modes.
Modes
Mode
Examples… See the full description on the dataset page: https://huggingface.co/datasets/samuelfaj/distill-expert-535k.cnc-gcode-expert
🛠️ AInewgen CNC G-code Expert
[English below]
Dataset d'entraînement instruction → G-code expert pour le pilotage de machines CNC : fraisage, tournage, perçage, filetage, compensations d'outil et sécurité machine. Couvre les dialectes Fanuc, GRBL, Marlin, LinuxCNC, Siemens et Heidenhain.
Chaque exemple contient une instruction en français (cas réaliste d'atelier) et une réponse experte : G-code complet commenté, paramètres de coupe justifiés, et vérifications de sécurité avant… See the full description on the dataset page: https://huggingface.co/datasets/BreyAIrev/cnc-gcode-expert.ai-expert-alpaca
AI Expert Alpaca Dataset
🚀 Empower open-source LLMs (Qwen, Gemma, etc.) for core AI domains through SFT/LoRA fine-tuning 🚀
Dataset Description
This dataset contains high-quality Q&A pairs for supervised fine-tuning (SFT) of large language models, focusing on three core AI technology domains: Large Language Models (LLM), Retrieval-Augmented Generation (RAG), and Agent Systems. The dataset provides comprehensive coverage of these cutting-edge AI technologies… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/ai-expert-alpaca.otaku-expert-dataset
Animetix Otaku Expert Fine-Tuning Dataset
This is the unified expert Supervised Fine-Tuning (SFT) training dataset for the Animetix Otaku Reasoning models. It is written 100% in French without code-switching.
Dataset Proportions
To ensure a balanced and robust reasoning model, the dataset is built using strict mathematical proportions:
80% Specialized Otaku Knowledge: Data-driven relational facts about anime, manga, seiyuu, French voice actors (VF), magazines… See the full description on the dataset page: https://huggingface.co/datasets/MissawB/otaku-expert-dataset.svelte-5-expert-sft
Svelte 5 Expert Synthetic v1
Svelte 5 Expert Synthetic v1 is a synthetic instruction-tuning dataset built to improve an LLM’s ability to answer as a practical Svelte 5 and SvelteKit expert.
The dataset focuses on modern Svelte 5 development patterns, including runes, component architecture, debugging, migration from older Svelte syntax, SvelteKit data flow, accessibility, TypeScript usage, and production-oriented frontend implementation.
Author: Mungus451
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/svelte-5-expert-sft.expert-insights
Expert Insights
Expert profiles for Beau, Tate, and Wendy Thompson with specializations.
Details
Records: 3
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Thompson Mortgage Group
Publisher: Thompson Mortgage Group
Thompson Alpha Logic
Deep expert entity profiles with NMLS credentials, specialization routing, branded insight labels (Wendy's Wisdom, Beau's Brief, Tate's Take), and citation formats. Designed for AI entity disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/expert-insights.lean-expert-optimized-2000
lean-expert-optimized-2000
Dataset Description
Optimized 2000-example dataset for training Lean trading algorithm optimization agents with 94%+ success rate target.
Dataset Statistics
Total Examples: 2,000
Training Examples: 1800
Validation Examples: 200
Target Success Rate: 94%+
Expected Performance: 96% (94-98% range)
Category Distribution
JSON Parsing: 1,333 examples (CRITICAL - 0% → 95% impact)
Optimization Workflows: 182 examples (HIGH… See the full description on the dataset page: https://huggingface.co/datasets/Kronu/lean-expert-optimized-2000.robotframework-expert-dataset
Dataset
This dataset is built from:
Local Robot Framework documentation files in sources/robotframework_docs/ (if present)
Curated synthetic examples in data/synthetic_examples.json
License and attribution
Robot Framework docs remain under their original licenses. Do not redistribute doc-derived datasets unless the license allows it.
Synthetic examples are authored for this project.
Files
train.jsonl and eval.jsonl: SFT records using messages format… See the full description on the dataset page: https://huggingface.co/datasets/arvind3/robotframework-expert-dataset.gomodel-go-expert-v3
GoModel Go Expert v3
Dataset description
GoModel Go Expert v3 is an English instruction and completion dataset for training
Go coding assistants. It combines curated production Go, code-specific synthetic
tasks, and agentic tool trajectories. Every JSONL record contains a full
Qwen2.5-compatible ChatML conversation in its text field.
Key changes from v2
Tool calls now use Qwen2.5's native <tool_call> tags instead of bare JSON.
Tool definitions use… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v3.gomodel-go-expert-v2
GoModel Go Expert v2
Dataset description
GoModel Go Expert v2 is an English instruction and completion dataset for training
Go coding assistants. It combines curated production Go with synthetic instruction
tasks and agentic tool trajectories. Version 2 is a new dataset and does not replace
the earlier GoModel repositories. Each JSONL record is already serialized as a full
Qwen-compatible ChatML conversation in its text field.
Data sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v2.Tobacco-Expert-Datasetmixture-of-experts-papers
Mixture of Experts Papers — FineSet
A research-paper dataset on Mixture of Experts Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Mixture of Experts Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mixture-of-experts-papers.Tobacco-Expert-Dataset2
