datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/diamond-in/github-top-developers.DeveloperSkills-Code2Skill
The dataset behind DeveloperSkillHubs
Grounded developer skills, mapped from real code.
Explore the Website
·
Dataset Configurations
·
Quick Start
CodeSkillBank
CodeSkillBank is a large-scale collection of reusable programming skills grounded in real source-code implementations. It is constructed with Code2Skill, an automated pipeline that transforms implementation evidence into structured procedural knowledge, and powers the accompanying… See the full description on the dataset page: https://huggingface.co/datasets/ant-intl/DeveloperSkills-Code2Skill.github_top_developers
github_top_developers (TsFile)
Apache TsFile version of diamond-in/github-top-developers.
Records: 39,390
Schema (TsFile structure)
name (TAG) — device dimension(s).
name (FIELD).
rank (FIELD).
Usage
Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file:
from pathlib import Path
from tsfile import TsFileReader
path = Path("github_top_developers.tsfile")
with TsFileReader(str(path)) as reader:
schemas =… See the full description on the dataset page: https://huggingface.co/datasets/THULab/github_top_developers.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-developers.developers-questions-small-qe2
Developers Questions Small QE2
A dataset consisting of ~12k developers' questions, in English. These questions are synthetically generated via local LLMs at Orama.
Datasets
The dataset is proposed with three different embedding models:
bge-small-en-v1.5
bge-base-en-v1.5
bge-large-en-v1.5
It also contains a quantized version for each model:
bge-small 32 bytes
bge-base 32 bytes
bge-large 32 bytes
For each quantized model, this repository includes a binary containing the… See the full description on the dataset page: https://huggingface.co/datasets/OramaSearch/developers-questions-small-qe2.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/github-top-developers.test_el_talar
Dataset Card for "test_el_talar"
More Information needed
developers-high-quality-mozgach
developers-high-quality-mozgach
Описание
Высококачественные примеры для разработчиков, сгенерированные mozgach108.
Датасет содержит отборные примеры для различных задач программирования:
Написание кода
Отладка
Рефакторинг
Архитектурные решения
Code review
Тестирование
Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108.
Сгенерировано через Ollama (mozgach108:latest).
Статистика
Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.Chronos-Reasoning-2k
Multilingual Instruction and Reasoning Dataset (Chronos-Reasoning-2k)
Русское описание находится ниже
English Description
Dataset Overview
This is a multilingual instruction dataset containing approximately 2,000 high-quality samples designed for instruction tuning and training language models on reasoning and Chain-of-Thought (CoT) patterns.
The dataset covers a wide array of topics, from contextual greetings and regional dialects to logic… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Reasoning-2k.cowdbdeveloper-salaries-norway-2024
Developer Salaries 2024 in Norway
This dataset is from kode24's 2024 salary survey.It includes salary information for 2,682 developers working in Norway who reported their salaries for 2024.
Columns:
kjønn: Gender of the respondent (e.g., "mann" for male, "kvinne" for female).
utdanning: Level of education, represented numerically:
4: Bachelor's degree
5: Master's degree
Other values represent different educational levels.
erfaring: Years of professional… See the full description on the dataset page: https://huggingface.co/datasets/habedi/developer-salaries-norway-2024.developers-108-perfect
developers-108-perfect
Описание
Датасет для обучения AI помощников разработчиков (continue.ai).
Включает 7 специализированных сфер × 108 примеров:
073: Developer - написание кода
074: Code Reviewer - проверка и ревью кода
075: Architect - проектирование систем
076: DevOps Engineer - CI/CD и инфраструктура
077: QA Tester - тестирование и качество
078: Technical Writer - документация
Плюс духовная сфера 001 + 1080 примеров Alpaca.
Сгенерировано через Ollama… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-108-perfect.Chronos-Thinking-v1-mini
English:
🌌 Chronos-Thinking-v1-mini: The Genesis of Structured Reasoning
Chronos-Thinking-v1-mini is a fundamental, high—density dataset designed to initialize deep reasoning processes in large language models (LLM). This dataset is the first step in the Chronos Super-AI project. Unlike mass datasets generated automatically, v1-mini relies on absolute quality and density of knowledge. He trains the model not just to answer questions, but to think like a system… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Thinking-v1-mini.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/Rendra8631/github-top-developers.Chronos-Reasoning-v1
🌌 Chronos Omega Reasoning v1 (Alpha)
📝 Описание
Chronos Omega Reasoning v1 — это высококачественный синтетический датасет, разработанный командой KZ Media Developers. Он предназначен для обучения языковых моделей глубокому логическому рассуждению (Chain-of-Thought) и формированию осознанного внутреннего монолога перед выдачей ответа.
Датасет сфокусирован на сложных задачах в области математики, программирования, физики и лингвистического анализа… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Reasoning-v1.StackOverflow-DeveloperSurveydeveloper-statsKali_developers
