datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kaidol-character-dataset
KAIdol Character Chat Dataset
한국어 캐릭터 롤플레이 대화 데이터셋
📋 목차
개요
데이터셋 통계
데이터 형식
캐릭터 목록
품질 지표
사용 방법
학습 가이드
제한사항
라이선스
🎯 개요
KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다.
주요 특징
특징
설명
🎭 41개 캐릭터
다양한 성격, 배경, 말투를 가진 캐릭터
🗣️ 음성 프로필
시그니처 표현, 종결어미, 금지 표현 정의
📊 3가지 형식
SFT, DPO, Multiturn 학습 지원
✅ 품질 검증
A등급 음성 프로필 일치율 (0.805)
🇰🇷 100% 한국어
자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.claude-fable-5-claude-code-etheroi
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/developerjeremylive/claude-fable-5-claude-code-etheroi.korean-character-roleplay-sft
Korean Character Roleplay SFT Dataset
Character-based Korean roleplay conversation dataset for fine-tuning language models.
Dataset Description
This dataset contains high-quality Korean roleplay conversations between users and AI characters. Each conversation follows a specific character's personality, speech patterns, and voice profile.
Dataset Statistics
Split
Samples
Train
965
Test
108
Total
1,073
Quality Metrics
Overall… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/korean-character-roleplay-sft.kaidol-phase2-rp-base-v0.3
KAIDOL Phase 2 RP Base Dataset v0.3
Dataset Description
KAIDOL Phase 2 RP Base v0.3 is a Korean-English bilingual conversational dataset designed for fine-tuning large language models (LLMs) for roleplay and character-based dialogue systems. This version includes GPT-Slop filtering to remove AI-sounding patterns and improve response quality.
What's New in v0.3
GPT-Slop Filtering: Removed 1,529 samples containing AI-sounding patterns
Cleaner Responses: Filtered… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-phase2-rp-base-v0.3.OncoAgent-Clinical-266K
🧬 OncoAgent Clinical Dataset — 266K
Curated Multi-Source Oncology Training Dataset
AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0
Dataset Description
This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks.
Composition
Source
Samples
Description
PMC-Patients
~100,000
Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/lablab-ai-amd-developer-hackathon/OncoAgent-Clinical-266K.developers-high-quality-mozgach
developers-high-quality-mozgach
Описание
Высококачественные примеры для разработчиков, сгенерированные mozgach108.
Датасет содержит отборные примеры для различных задач программирования:
Написание кода
Отладка
Рефакторинг
Архитектурные решения
Code review
Тестирование
Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108.
Сгенерировано через Ollama (mozgach108:latest).
Статистика
Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.developers-108-perfect
developers-108-perfect
Описание
Датасет для обучения AI помощников разработчиков (continue.ai).
Включает 7 специализированных сфер × 108 примеров:
073: Developer - написание кода
074: Code Reviewer - проверка и ревью кода
075: Architect - проектирование систем
076: DevOps Engineer - CI/CD и инфраструктура
077: QA Tester - тестирование и качество
078: Technical Writer - документация
Плюс духовная сфера 001 + 1080 примеров Alpaca.
Сгенерировано через Ollama… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-108-perfect.Chronos-Thinking-v1-mini
English:
🌌 Chronos-Thinking-v1-mini: The Genesis of Structured Reasoning
Chronos-Thinking-v1-mini is a fundamental, high—density dataset designed to initialize deep reasoning processes in large language models (LLM). This dataset is the first step in the Chronos Super-AI project. Unlike mass datasets generated automatically, v1-mini relies on absolute quality and density of knowledge. He trains the model not just to answer questions, but to think like a system… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Thinking-v1-mini.Chronos-Reasoning-v1
🌌 Chronos Omega Reasoning v1 (Alpha)
📝 Описание
Chronos Omega Reasoning v1 — это высококачественный синтетический датасет, разработанный командой KZ Media Developers. Он предназначен для обучения языковых моделей глубокому логическому рассуждению (Chain-of-Thought) и формированию осознанного внутреннего монолога перед выдачей ответа.
Датасет сфокусирован на сложных задачах в области математики, программирования, физики и лингвистического анализа… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Reasoning-v1.developer-marketing-playbook
Developer Marketing Playbook
Complete developer marketing playbook covering DevRel programs, documentation as marketing, API developer experience, community building, and hackathon strat...
📦 Install on ClawHub
clawhub install developer-marketing-playbook
Then ask your AI agent:
"Plan a 90-day marketing strategy for our B2B SaaS"
Installs the full Developer Marketing Playbook playbook — battle-tested with 30+ Product Hunt #1 wins, AFFiNE 60K+ GitHub stars… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/developer-marketing-playbook.
