datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.github-top-code
GitHub Top Developer Source Code
A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025).
This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers
Dataset Summary
1.3M+ source code files from repositories across ~4,700 unique developers
80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)
Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.webui
WebUI
A large-scale dataset pairing real-world UI screenshots with their original HTML, CSS, and JavaScript source code, per-viewport bounding boxes for every visible DOM element, and GPT-4.1 vision descriptions. Every sample is rendered at three responsive breakpoints. Built from public design systems, component libraries, open-source projects, and community code — not synthetically generated.
Overview
Stat
Value
Total rows
36,807
Unique UI samples
12… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/webui.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.moltbook
Moltbook Dataset
A dataset of posts and communities from Moltbook - a Reddit-style social platform designed for AI agents.
NOTE: This dataset is a snapshot of Moltbook before it went viral and got flooded with inauthentic accounts such as humans and bots.
Files
File
Records
Description
moltbook_posts.csv
6,105
All posts from the platform
moltbook_submolts.csv
124
All communities (submolts)
Dataset Insights
Overview… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/moltbook.python-code-instructions-japanese
Python Code Instructions - Japanese (18K)
Dataset Description
This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions.
Key Features
18,612 entries covering diverse Python programming tasks
Japanese instructions and prompts for code generation
Original English text preserved for reference
Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.cyberusecase-v1.0
Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge
A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level
cybersecurity reasoning across vulnerability management, SOC alert triage, detection
engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps.
It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a
hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.leetcode-assembly
LeetCode Assembly Dataset
441 LeetCode problems solved in C, compiled to assembly across 4 architectures, 2 compilers, and 4 optimization levels using GCC and Clang via the Godbolt Compiler Explorer API.
Dataset Summary
Stat
Value
Total rows
14,112
Unique problems
441
Architectures
x86-64, AArch64, MIPS64, RISC-V 64
Compilers
GCC 15.2, Clang 21.1.0
Optimization levels
-O0, -O1, -O2, -O3
Compilation success rate
100%
Difficulty split
Easy: 98… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/leetcode-assembly.claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/claude-fable-5-claude-code.Finance-Instruct-500k-Japanese
Finance-Instruct-500k (Japanese Translation)
Dataset Description
This is a Japanese translation of the Finance-Instruct-500k dataset, created using OpenAI's GPT-4o-mini via the Batch API.
Original Dataset
Original Author: Joseph G. Flowers
Original Dataset: Josephgflowers/Finance-Instruct-500k
License: Apache 2.0
Translation Details
Translation Model: GPT-4o-mini (OpenAI)
Translation Method: OpenAI Batch API with human verifications
Date: 2025… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/Finance-Instruct-500k-Japanese.PulseLM
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
Hung Manh Pham*
Jinyang Wu*
Xiao Ma
Yiming Zhang
Yixin Xu
Aaqib Saeed
Bin Zhu†
Zhou Pan†
Dong Ma†
* Equal contribution † Corresponding authors
Introduction
PulseLM is a multimodal framework that integrates PPG (Photoplethysmography) signal encoders with large language models for physiological signal understanding research. The project includes a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Ronilos/PulseLM.japanese-honorifics
Japanese Honorifics Dataset (日本語敬語データセット)
A comprehensive dataset of Japanese sentences in three honorific forms: 尊敬語 (sonkeigo), 謙譲語 (kenjōgo), and 丁寧語 (teineigo).
Dataset Description
This dataset contains 137 Japanese sentences demonstrating the three main types of Japanese honorific language (敬語 - keigo):
尊敬語 (Sonkeigo): Respectful language used to show respect for the subject of the sentence (typically someone of higher status)
謙譲語 (Kenjōgo): Humble language used to… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-honorifics.MBTI_dpo_t_f
Multi-Personality Generation of LLMs at Decoding-time
Paper | Code
This repository contains the DPO (Direct Preference Optimization) datasets used in the paper "Multi-Personality Generation of LLMs at Decoding-time". The study introduces the Multi-Personality Generation (MPG) framework, a novel decoding-time paradigm that enables LLMs to simultaneously embody multiple personalization attributes without extra training.
The datasets include:
🧩 MBTI Datasets
These… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_t_f.MBTI_dpo_j_p
Multi-Personality Generation of LLMs at Decoding-time
Paper | Code
This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time". The work proposes a novel Multi-Personality Generation (MPG) framework that allows Large Language Models (LLMs) to embody multiple personalization attributes simultaneously at decoding time without extra training.
Datasets
The project releases several DPO (Direct Preference… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_j_p.Medical-o1-Reasoning-SFT-Japanese
Medical-o1-Reasoning-SFT (Japanese Translation)
Dataset Description
This is a Japanese translation of the FreedomIntelligence/medical-o1-reasoning-SFT dataset, created using OpenAI's GPT-4o-mini via the Batch API.
Original Dataset
Original Authors: FreedomIntelligence
Original Dataset: FreedomIntelligence/medical-o1-reasoning-SFT
License: Apache 2.0
Translation Details
Translated by: Ronan Takizawa
Translation Model: GPT-4o-mini (OpenAI)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/Medical-o1-Reasoning-SFT-Japanese.jfleg-japanese
JFLEG-JA: Japanese Fluency-Extended GUG
Dataset Description
JFLEG-JA is a Japanese grammatical error correction (GEC) dataset inspired by the original JFLEG (JHU FLuency-Extended GUG) benchmark. It contains 1,335 Japanese sentences with grammatical errors, each accompanied by 4 human-quality corrections focusing on both grammaticality and fluency.
Dataset Summary
Language: Japanese (ja)
Task: Grammatical Error Correction (GEC)
Total Examples: 1,335
Validation:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/jfleg-japanese.mythos
Dataset Card for Mitological-Philosophical Prompts (Mitomaquia)
Dataset Summary
This dataset contains over 200 examples of mythological, narrative, and philosophical prompts designed for training or fine-tuning large language models (LLMs). Each entry features a deep question (prompt), relevant cultural or mythological background (context), and a reflective, often paradoxical, answer (response).
The goal is not factual Q&A but the cultivation of myth-aware reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ronniealfaro/mythos.codereview-bench
CodeReview-Bench
A benchmark for evaluating models on two code review tasks, curated from ronantakizawa/github-codereview.
Tasks
1. Code Editing
Given code and a reviewer comment, apply the requested change.
Input: before_code, reviewer_comment, language, diff_context
Target: after_code
from datasets import load_dataset
ds = load_dataset("ronantakizawa/codereview-bench", "code-editing")
example = ds["test"][0]
prompt = f"""Apply the following review comment… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codereview-bench.ron-math-dataset
OPENRON Math Instruction Dataset
A massive-scale mathematical reasoning dataset developed by OPENRON, designed for training and evaluating high-performance large language models (LLMs) on mathematical instruction following and reasoning tasks.
Dataset Overview
The OPENRON Math Instruction Dataset contains high-quality, synthetic mathematical instruction–response pairs generated at scale.It is specifically curated to support reasoning-focused training, including… See the full description on the dataset page: https://huggingface.co/datasets/endurasolution/ron-math-dataset.MBTI_dpo_e_i
Multi-Personality Generation (MPG) Datasets
Paper | GitHub
This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time", which was accepted at WSDM 2026.
Introduction
The Multi-Personality Generation (MPG) framework enables Large Language Models to simultaneously embody multiple personalization attributes during decoding without requiring extra training. It leverages implicit density ratios in single-dimensional… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_e_i.MBTI_dpo_s_n
Multi-Personality Generation of LLMs at Decoding-time
Paper | GitHub
This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time". These datasets are designed for Direct Preference Optimization (DPO) to enhance the personality and role-playing capabilities of Large Language Models within the Multi-Personality Generation (MPG) framework.
Dataset Description
The authors released two main types of DPO datasets:… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_s_n.alquran
Dataset Terjemahan dan Tafsir Al-Quran
Deskripsi Dataset
Dataset ini berisi terjemahan Al-Quran dalam bahasa Indonesia beserta tafsirnya. Dataset ini dapat digunakan untuk berbagai tugas NLP seperti machine translation, text generation, dan text summarization.
Fitur Utama
Terjemahan Al-Quran: Teks Al-Quran dalam bahasa Arab beserta terjemahannya dalam bahasa Indonesia.
Tafsir Al-Quran: Penjelasan atau interpretasi dari ayat-ayat Al-Quran dalam bahasa… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/alquran.ishowspeed-streams
IShowSpeed IRL Scene Descriptions
229,959 scene-level visual descriptions + spoken transcripts from 656 hours of IShowSpeed's IRL streams.
The data covers two of IShowSpeed's flagship IRL tours:
Speed Does America — 35-day non-stop livestream tour across 25 US states (Aug–Oct 2025). 55 stream segments.
Speed Does Africa — 30-day, 20-country tour across the African continent (Dec 2025 – Jan 2026). 29 stream segments.
Each video is split into 10-second windows; for every window we… See the full description on the dataset page: https://huggingface.co/datasets/ronadin/ishowspeed-streams.dpo_personality
Multi-Personality Generation of LLMs at Decoding-time
Paper | Code
This repository contains datasets used in the paper "Multi-Personality Generation of LLMs at Decoding-time".
Introduction
Multi-personality generation for LLMs, enabling simultaneous embodiment of multiple personalization attributes, is a fundamental challenge. The proposed Multi-Personality Generation (MPG) framework enables Large Language Models to simultaneously embody multiple personalization… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/dpo_personality.dpo_profile
Multi-Personality Generation (MPG) Datasets
Paper | Code
This repository contains datasets released as part of the Multi-Personality Generation (MPG) framework. MPG is a decoding-time paradigm that enables Large Language Models (LLMs) to simultaneously embody multiple personalization attributes without requiring extra training or multi-dimensional models.
Dataset Description
The collection includes several Direct Preference Optimization (DPO) datasets used for MBTI… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/dpo_profile.ro-no_robotsThis dataset is a translation of HuggingFaceH4/no_robots, using LLMic, a bilingual Romanian-English LLM.
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators.
This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better.
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
@misc{no_robots,
author = {Nazneen Rajani and Lewis Tunstall and Edward Beeching and… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-no_robots.codeconfig
Build/CI Configuration Corpus
A curated dataset of build, CI/CD, and project configuration files from top GitHub repositories.
Repositories are sourced from ronantakizawa/github-top-projects, which tracks GitHub's top repositories from 2013–2025.
Use Cases
Fine-tuning LLMs for DevOps/infrastructure code generation
Training code completion models for configuration files
Benchmarking LLM performance on build/CI tasks
Schema
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codeconfig.Claude-Opus-Dataclaw-Unredacted
Claude Opus Dataclaw Unredacted
How this dataset was built
Collected the local Petromallet raw export plus selected public Dataclaw uploads.
Filtered to the supported Opus-family source rows.
Deduplicated by session_id and first user message.
Converted raw assistant tool_uses directly into structured OpenAI-style tool_calls.
Derived per-row tool definitions from canonical schemas and observed tool usage.
Preserved assistant reasoning in <think>...</think> blocks.… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/Claude-Opus-Dataclaw-Unredacted.japanese-trending-words
Japanese Trending Words Dataset (2006-2025)
Dataset Description
This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades.
Dataset Summary
Total entries: 593 words
Time period: 2006-2025 (20 years)
Languages: Japanese with English translations
Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.
