datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kaidol-character-dataset
KAIdol Character Chat Dataset
한국어 캐릭터 롤플레이 대화 데이터셋
📋 목차
개요
데이터셋 통계
데이터 형식
캐릭터 목록
품질 지표
사용 방법
학습 가이드
제한사항
라이선스
🎯 개요
KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다.
주요 특징
특징
설명
🎭 41개 캐릭터
다양한 성격, 배경, 말투를 가진 캐릭터
🗣️ 음성 프로필
시그니처 표현, 종결어미, 금지 표현 정의
📊 3가지 형식
SFT, DPO, Multiturn 학습 지원
✅ 품질 검증
A등급 음성 프로필 일치율 (0.805)
🇰🇷 100% 한국어
자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/diamond-in/github-top-developers.DBPedia_ClassesAbout Dataset
DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in Wikipedia.
This is an extract of the data (after cleaning, kernel included) that provides taxonomic, hierarchical categories ("classes") for 342,782 wikipedia articles. There are 3 levels, with 9, 70 and 219 classes respectively.
A version of this dataset is a popular baseline for NLP/text classification tasks. This version of the dataset is much tougher… See the full description on the dataset page: https://huggingface.co/datasets/DeveloperOats/DBPedia_Classes.CMAPSS_Jet_Engine_Simulated_Datastack-overflow-developer-surveyclaude-fable-5-claude-code-etheroi
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/developerjeremylive/claude-fable-5-claude-code-etheroi.kimi-idol-dataset
Kimi K2 Idol Character Dataset
한국어 아이돌 캐릭터 대화 데이터셋 - Kimi K2 Teacher-Student Distillation 용
Overview
이 데이터셋은 moonshotai/Kimi-K2-Instruct-0905 모델을 Teacher로 사용하여 생성한 고품질 한국어 아이돌 캐릭터 대화 데이터셋입니다. 5명의 독특한 캐릭터가 6단계 thinking (CoT) 과정을 통해 썸남/썸녀 관계의 밀당 대화를 생성합니다.
Dataset Details
Total Samples: 942 (Train: 847, Eval: 95)
Characters: 5명 (강율, 서이안, 이지후, 차도하, 최민)
Teacher Model: moonshotai/Kimi-K2-Instruct-0905
Format: Messages format with </think> tags for thinking
Quality:… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kimi-idol-dataset.cyber-cve2cwe-extension
cyber-cve2cwe-extension
Overview
The CVE-to-CWE classification task suffers from low macro-averaged F1 scores because many CWE categories appear only a handful of times in the training data. This dataset supplies additional examples for 36 low-frequency (tail) CWE classes with the aim of improving model performance on those categories and providing a reproducible record of how the training data for the companion model was extended. It is intended as a transparency… See the full description on the dataset page: https://huggingface.co/datasets/luca-software-developer/cyber-cve2cwe-extension.aicoolies-developer-tools-knowledge-graph
aicoolies-developer-tools-knowledge-graph
Public catalog dump from aicoolies.com: tools, comparisons, and scored reviews as JSON.
This dataset is not a coding agent. It does not edit repositories, run tools, or execute code. It is a machine-readable snapshot of the public aicoolies Developer Tools Knowledge Graph so humans and research agents can reuse the catalog without scraping HTML.
Canonical open-data page: https://aicoolies.com/data
Homepage: https://aicoolies.com… See the full description on the dataset page: https://huggingface.co/datasets/rasitakyol/aicoolies-developer-tools-knowledge-graph.kerala.datasetgithub-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-developers.aws-lambda-developer-guide-docskorean-character-roleplay-sft
Korean Character Roleplay SFT Dataset
Character-based Korean roleplay conversation dataset for fine-tuning language models.
Dataset Description
This dataset contains high-quality Korean roleplay conversations between users and AI characters. Each conversation follows a specific character's personality, speech patterns, and voice profile.
Dataset Statistics
Split
Samples
Train
965
Test
108
Total
1,073
Quality Metrics
Overall… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/korean-character-roleplay-sft.kaidol-phase2-rp-base-v0.3
KAIDOL Phase 2 RP Base Dataset v0.3
Dataset Description
KAIDOL Phase 2 RP Base v0.3 is a Korean-English bilingual conversational dataset designed for fine-tuning large language models (LLMs) for roleplay and character-based dialogue systems. This version includes GPT-Slop filtering to remove AI-sounding patterns and improve response quality.
What's New in v0.3
GPT-Slop Filtering: Removed 1,529 samples containing AI-sounding patterns
Cleaner Responses: Filtered… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-phase2-rp-base-v0.3.developer-reference-datasets
Developer Reference Datasets
Open, reproducible lookup tables that web and app developers reach for constantly — computed from first principles, not scraped, so every value is exact and re-runnable. CC BY 4.0.
Quick answers (straight from the data)
What is 16:9 in pixels? 1920×1080, 1280×720, 3840×2160. 9:16 (Stories, Reels, TikTok) is those flipped. → aspect-ratios, resolutions
What contrast ratio does WCAG require? 4.5:1 for normal text (AA), 3:1 for large… See the full description on the dataset page: https://huggingface.co/datasets/cleanorlabs/developer-reference-datasets.Million_News_HeadlinesAbout Dataset
Context
This contains data of news headlines published over a period of nineteen years.
Sourced from the reputable Australian news source ABC (Australian Broadcasting Corporation)
Agency Site: (http://www.abc.net.au)
Content
Format: CSV ; Single File
publish_date: Date of publishing for the article in yyyyMMdd format
headline_text: Text of the headline in Ascii , English , lowercase
Start Date: 2003-02-19 ; End Date: 2021-12-31
Inspiration
I look at this news dataset as a… See the full description on the dataset page: https://huggingface.co/datasets/DeveloperOats/Million_News_Headlines.software-developer-hourly-rates-2026
Software Developer Hourly Rate Benchmarks 2026 (by Platform, Region & Tier)
Open benchmark of software/web developer hourly rates in 2026, segmented by technology platform, geographic region, and delivery tier (independent freelancer vs. team-based agency). Two files:
developer_hourly_rates_2026.csv — 120 rows: platform, region, tier, hourly_low_usd, hourly_median_usd, hourly_high_usd. Platforms include WordPress, Shopify, Magento, WooCommerce, custom (React/Next/Node), and… See the full description on the dataset page: https://huggingface.co/datasets/zimzum1984/software-developer-hourly-rates-2026.Developer-Documentation-QAafrica-synth-realestate-developer-profiles-nigeria
Africa Synth Realestate Developer Profiles Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-realestate-developer-profiles-nigeria.qiniu-developer-faq_ShareGPTdataset_23github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/github-top-developers.test_el_talar
Dataset Card for "test_el_talar"
More Information needed
developer-portfolio-ragdevelopers-high-quality-mozgach
developers-high-quality-mozgach
Описание
Высококачественные примеры для разработчиков, сгенерированные mozgach108.
Датасет содержит отборные примеры для различных задач программирования:
Написание кода
Отладка
Рефакторинг
Архитектурные решения
Code review
Тестирование
Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108.
Сгенерировано через Ollama (mozgach108:latest).
Статистика
Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.cowdbqiniu-developerdeveloper-salaries-norway-2024
Developer Salaries 2024 in Norway
This dataset is from kode24's 2024 salary survey.It includes salary information for 2,682 developers working in Norway who reported their salaries for 2024.
Columns:
kjønn: Gender of the respondent (e.g., "mann" for male, "kvinne" for female).
utdanning: Level of education, represented numerically:
4: Bachelor's degree
5: Master's degree
Other values represent different educational levels.
erfaring: Years of professional… See the full description on the dataset page: https://huggingface.co/datasets/habedi/developer-salaries-norway-2024.synthetic_AI_course_support_ticket_and_labelZoyeMedical_medical_data
