datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.moltbook
Moltbook Dataset
A dataset of posts and communities from Moltbook - a Reddit-style social platform designed for AI agents.
NOTE: This dataset is a snapshot of Moltbook before it went viral and got flooded with inauthentic accounts such as humans and bots.
Files
File
Records
Description
moltbook_posts.csv
6,105
All posts from the platform
moltbook_submolts.csv
124
All communities (submolts)
Dataset Insights
Overview… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/moltbook.leetcode-assembly
LeetCode Assembly Dataset
441 LeetCode problems solved in C, compiled to assembly across 4 architectures, 2 compilers, and 4 optimization levels using GCC and Clang via the Godbolt Compiler Explorer API.
Dataset Summary
Stat
Value
Total rows
14,112
Unique problems
441
Architectures
x86-64, AArch64, MIPS64, RISC-V 64
Compilers
GCC 15.2, Clang 21.1.0
Optimization levels
-O0, -O1, -O2, -O3
Compilation success rate
100%
Difficulty split
Easy: 98… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/leetcode-assembly.claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/claude-fable-5-claude-code.PulseLM
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
Hung Manh Pham*
Jinyang Wu*
Xiao Ma
Yiming Zhang
Yixin Xu
Aaqib Saeed
Bin Zhu†
Zhou Pan†
Dong Ma†
* Equal contribution † Corresponding authors
Introduction
PulseLM is a multimodal framework that integrates PPG (Photoplethysmography) signal encoders with large language models for physiological signal understanding research. The project includes a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Ronilos/PulseLM.alquran
Dataset Terjemahan dan Tafsir Al-Quran
Deskripsi Dataset
Dataset ini berisi terjemahan Al-Quran dalam bahasa Indonesia beserta tafsirnya. Dataset ini dapat digunakan untuk berbagai tugas NLP seperti machine translation, text generation, dan text summarization.
Fitur Utama
Terjemahan Al-Quran: Teks Al-Quran dalam bahasa Arab beserta terjemahannya dalam bahasa Indonesia.
Tafsir Al-Quran: Penjelasan atau interpretasi dari ayat-ayat Al-Quran dalam bahasa… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/alquran.ishowspeed-streams
IShowSpeed IRL Scene Descriptions
229,959 scene-level visual descriptions + spoken transcripts from 656 hours of IShowSpeed's IRL streams.
The data covers two of IShowSpeed's flagship IRL tours:
Speed Does America — 35-day non-stop livestream tour across 25 US states (Aug–Oct 2025). 55 stream segments.
Speed Does Africa — 30-day, 20-country tour across the African continent (Dec 2025 – Jan 2026). 29 stream segments.
Each video is split into 10-second windows; for every window we… See the full description on the dataset page: https://huggingface.co/datasets/ronadin/ishowspeed-streams.codeconfig
Build/CI Configuration Corpus
A curated dataset of build, CI/CD, and project configuration files from top GitHub repositories.
Repositories are sourced from ronantakizawa/github-top-projects, which tracks GitHub's top repositories from 2013–2025.
Use Cases
Fine-tuning LLMs for DevOps/infrastructure code generation
Training code completion models for configuration files
Benchmarking LLM performance on build/CI tasks
Schema
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codeconfig.japanese-trending-words
Japanese Trending Words Dataset (2006-2025)
Dataset Description
This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades.
Dataset Summary
Total entries: 593 words
Time period: 2006-2025 (20 years)
Languages: Japanese with English translations
Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.india-trending-words
Google India Trending Words Dataset (2008-2021, 2023-2024)
Dataset Description
This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com).
Dataset Summary
Total Entries: 900
Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available)
Categories: 18 unique tags
Region: India
Format: CSV
Dataset Structure
Data Fields
word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.
