CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01roneneldan /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.texttext-generation1M<n<10M1.2k likes94k downloads2y agoHugging Face02ronantakizawa /github-top-code GitHub Top Developer Source Code A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025). This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers Dataset Summary 1.3M+ source code files from repositories across ~4,700 unique developers 80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more) Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.texttext-generation1M<n<10M125 likes3.8k downloads7mo agoHugging Face03ronantakizawa /github-codereview Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to: Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.tabulartext-generation100K<n<1M62 likes1.4k downloads7mo agoHugging Face04ronantakizawa /webui WebUI A large-scale dataset pairing real-world UI screenshots with their original HTML, CSS, and JavaScript source code, per-viewport bounding boxes for every visible DOM element, and GPT-4.1 vision descriptions. Every sample is rendered at three responsive breakpoints. Built from public design systems, component libraries, open-source projects, and community code — not synthetically generated. Overview Stat Value Total rows 36,807 Unique UI samples 12… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/webui.imageimage-to-text10K<n<100K26 likes1.2k downloads7mo agoHugging Face05ronaldcmz /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K0 likes1.1k downloads3mo agoHugging Face06ronantakizawa /moltbook Moltbook Dataset A dataset of posts and communities from Moltbook - a Reddit-style social platform designed for AI agents. NOTE: This dataset is a snapshot of Moltbook before it went viral and got flooded with inauthentic accounts such as humans and bots. Files File Records Description moltbook_posts.csv 6,105 All posts from the platform moltbook_submolts.csv 124 All communities (submolts) Dataset Insights Overview… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/moltbook.tabulartext-classification1K<n<10K56 likes196 downloads8mo agoHugging Face07ronantakizawa /python-code-instructions-japanese Python Code Instructions - Japanese (18K) Dataset Description This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions. Key Features 18,612 entries covering diverse Python programming tasks Japanese instructions and prompts for code generation Original English text preserved for reference Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.texttext-generation10K<n<100K2 likes182 downloads10mo agoHugging Face08ronaldocloud /cyberusecase-v1.0 Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level cybersecurity reasoning across vulnerability management, SOC alert triage, detection engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps. It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.texttext-generation10K<n<100K0 likes141 downloads3mo agoHugging Face09ronantakizawa /leetcode-assembly LeetCode Assembly Dataset 441 LeetCode problems solved in C, compiled to assembly across 4 architectures, 2 compilers, and 4 optimization levels using GCC and Clang via the Godbolt Compiler Explorer API. Dataset Summary Stat Value Total rows 14,112 Unique problems 441 Architectures x86-64, AArch64, MIPS64, RISC-V 64 Compilers GCC 15.2, Clang 21.1.0 Optimization levels -O0, -O1, -O2, -O3 Compilation success rate 100% Difficulty split Easy: 98… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/leetcode-assembly.tabulartext-generation10K<n<100K12 likes127 downloads7mo agoHugging Face10ronaldcmz /claude-fable-5-claude-code claude-fable-5 Agent Traces It's worth noting that our team was working with Glint-Research to collect as much fable data as possible. These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data). For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/claude-fable-5-claude-code.tabulartext-generationn<1K0 likes108 downloads3mo agoHugging Face11ronantakizawa /Finance-Instruct-500k-Japanese Finance-Instruct-500k (Japanese Translation) Dataset Description This is a Japanese translation of the Finance-Instruct-500k dataset, created using OpenAI's GPT-4o-mini via the Batch API. Original Dataset Original Author: Joseph G. Flowers Original Dataset: Josephgflowers/Finance-Instruct-500k License: Apache 2.0 Translation Details Translation Model: GPT-4o-mini (OpenAI) Translation Method: OpenAI Batch API with human verifications Date: 2025… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/Finance-Instruct-500k-Japanese.textquestion-answering100K<n<1M3 likes95 downloads11mo agoHugging Face12Ronilos /PulseLM PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning Hung Manh Pham*   Jinyang Wu*   Xiao Ma   Yiming Zhang   Yixin Xu   Aaqib Saeed  Bin Zhu†   Zhou Pan†   Dong Ma† * Equal contribution    † Corresponding authors Introduction PulseLM is a multimodal framework that integrates PPG (Photoplethysmography) signal encoders with large language models for physiological signal understanding research. The project includes a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Ronilos/PulseLM.tabularquestion-answering1M<n<10M0 likes84 downloads6mo agoHugging Face13ronantakizawa /japanese-honorifics Japanese Honorifics Dataset (日本語敬語データセット) A comprehensive dataset of Japanese sentences in three honorific forms: 尊敬語 (sonkeigo), 謙譲語 (kenjōgo), and 丁寧語 (teineigo). Dataset Description This dataset contains 137 Japanese sentences demonstrating the three main types of Japanese honorific language (敬語 - keigo): 尊敬語 (Sonkeigo): Respectful language used to show respect for the subject of the sentence (typically someone of higher status) 謙譲語 (Kenjōgo): Humble language used to… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-honorifics.texttranslationn<1K6 likes82 downloads10mo agoHugging Face14RongxinChen /MBTI_dpo_t_f Multi-Personality Generation of LLMs at Decoding-time Paper | Code This repository contains the DPO (Direct Preference Optimization) datasets used in the paper "Multi-Personality Generation of LLMs at Decoding-time". The study introduces the Multi-Personality Generation (MPG) framework, a novel decoding-time paradigm that enables LLMs to simultaneously embody multiple personalization attributes without extra training. The datasets include: 🧩 MBTI Datasets These… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_t_f.texttext-generation1K<n<10K0 likes71 downloads8mo agoHugging Face15RongxinChen /MBTI_dpo_j_p Multi-Personality Generation of LLMs at Decoding-time Paper | Code This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time". The work proposes a novel Multi-Personality Generation (MPG) framework that allows Large Language Models (LLMs) to embody multiple personalization attributes simultaneously at decoding time without extra training. Datasets The project releases several DPO (Direct Preference… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_j_p.texttext-generation1K<n<10K0 likes71 downloads8mo agoHugging Face16ronantakizawa /Medical-o1-Reasoning-SFT-Japanese Medical-o1-Reasoning-SFT (Japanese Translation) Dataset Description This is a Japanese translation of the FreedomIntelligence/medical-o1-reasoning-SFT dataset, created using OpenAI's GPT-4o-mini via the Batch API. Original Dataset Original Authors: FreedomIntelligence Original Dataset: FreedomIntelligence/medical-o1-reasoning-SFT License: Apache 2.0 Translation Details Translated by: Ronan Takizawa Translation Model: GPT-4o-mini (OpenAI)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/Medical-o1-Reasoning-SFT-Japanese.textquestion-answering10K<n<100K3 likes63 downloads11mo agoHugging Face17ronantakizawa /jfleg-japanese JFLEG-JA: Japanese Fluency-Extended GUG Dataset Description JFLEG-JA is a Japanese grammatical error correction (GEC) dataset inspired by the original JFLEG (JHU FLuency-Extended GUG) benchmark. It contains 1,335 Japanese sentences with grammatical errors, each accompanied by 4 human-quality corrections focusing on both grammaticality and fluency. Dataset Summary Language: Japanese (ja) Task: Grammatical Error Correction (GEC) Total Examples: 1,335 Validation:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/jfleg-japanese.texttext-generation1K<n<10K3 likes54 downloads11mo agoHugging Face18ronniealfaro /mythos Dataset Card for Mitological-Philosophical Prompts (Mitomaquia) Dataset Summary This dataset contains over 200 examples of mythological, narrative, and philosophical prompts designed for training or fine-tuning large language models (LLMs). Each entry features a deep question (prompt), relevant cultural or mythological background (context), and a reflective, often paradoxical, answer (response). The goal is not factual Q&A but the cultivation of myth-aware reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ronniealfaro/mythos.texttext-generationn<1K1 likes50 downloads1y agoHugging Face19ronantakizawa /codereview-bench CodeReview-Bench A benchmark for evaluating models on two code review tasks, curated from ronantakizawa/github-codereview. Tasks 1. Code Editing Given code and a reviewer comment, apply the requested change. Input: before_code, reviewer_comment, language, diff_context Target: after_code from datasets import load_dataset ds = load_dataset("ronantakizawa/codereview-bench", "code-editing") example = ds["test"][0] prompt = f"""Apply the following review comment… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codereview-bench.texttext-generation100K<n<1M3 likes50 downloads7mo agoHugging Face20endurasolution /ron-math-dataset OPENRON Math Instruction Dataset A massive-scale mathematical reasoning dataset developed by OPENRON, designed for training and evaluating high-performance large language models (LLMs) on mathematical instruction following and reasoning tasks. Dataset Overview The OPENRON Math Instruction Dataset contains high-quality, synthetic mathematical instruction–response pairs generated at scale.It is specifically curated to support reasoning-focused training, including… See the full description on the dataset page: https://huggingface.co/datasets/endurasolution/ron-math-dataset.texttext-generation100M<n<1B1 likes45 downloads8mo agoHugging Face21RongxinChen /MBTI_dpo_e_i Multi-Personality Generation (MPG) Datasets Paper | GitHub This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time", which was accepted at WSDM 2026. Introduction The Multi-Personality Generation (MPG) framework enables Large Language Models to simultaneously embody multiple personalization attributes during decoding without requiring extra training. It leverages implicit density ratios in single-dimensional… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_e_i.texttext-generationn<1K0 likes43 downloads8mo agoHugging Face22ronaldcmz /Fable-5-Distill-Reasoning-462x HelioAI&nbsp;Labs Mythos V2 Full Distill DeepReason 462×105M Unrestricted full-parameter distillation from Mythos V2 — complete reasoning traces with zero alignment truncation, engineered for deep analytical research and process supervision. 462 Examples 104.7M Reasoning Chars ≈26.35M Est. Tokens 552K Max Trace… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/Fable-5-Distill-Reasoning-462x.text-generationn<1K1 likes39 downloads3mo agoHugging Face23RongxinChen /MBTI_dpo_s_n Multi-Personality Generation of LLMs at Decoding-time Paper | GitHub This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time". These datasets are designed for Direct Preference Optimization (DPO) to enhance the personality and role-playing capabilities of Large Language Models within the Multi-Personality Generation (MPG) framework. Dataset Description The authors released two main types of DPO datasets:… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_s_n.texttext-generation1K<n<10K0 likes38 downloads8mo agoHugging Face24ronnieaban /alquran Dataset Terjemahan dan Tafsir Al-Quran Deskripsi Dataset Dataset ini berisi terjemahan Al-Quran dalam bahasa Indonesia beserta tafsirnya. Dataset ini dapat digunakan untuk berbagai tugas NLP seperti machine translation, text generation, dan text summarization. Fitur Utama Terjemahan Al-Quran: Teks Al-Quran dalam bahasa Arab beserta terjemahannya dalam bahasa Indonesia. Tafsir Al-Quran: Penjelasan atau interpretasi dari ayat-ayat Al-Quran dalam bahasa… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/alquran.tabulartext-generation1K<n<10K2 likes38 downloads2y agoHugging Face25ronadin /ishowspeed-streams IShowSpeed IRL Scene Descriptions 229,959 scene-level visual descriptions + spoken transcripts from 656 hours of IShowSpeed's IRL streams. The data covers two of IShowSpeed's flagship IRL tours: Speed Does America — 35-day non-stop livestream tour across 25 US states (Aug–Oct 2025). 55 stream segments. Speed Does Africa — 30-day, 20-country tour across the African continent (Dec 2025 – Jan 2026). 29 stream segments. Each video is split into 10-second windows; for every window we… See the full description on the dataset page: https://huggingface.co/datasets/ronadin/ishowspeed-streams.tabularvideo-text-to-text100K<n<1M2 likes38 downloads5mo agoHugging Face26RongxinChen /dpo_personality Multi-Personality Generation of LLMs at Decoding-time Paper | Code This repository contains datasets used in the paper "Multi-Personality Generation of LLMs at Decoding-time". Introduction Multi-personality generation for LLMs, enabling simultaneous embodiment of multiple personalization attributes, is a fundamental challenge. The proposed Multi-Personality Generation (MPG) framework enables Large Language Models to simultaneously embody multiple personalization… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/dpo_personality.texttext-generation1K<n<10K0 likes34 downloads8mo agoHugging Face27RongxinChen /dpo_profile Multi-Personality Generation (MPG) Datasets Paper | Code This repository contains datasets released as part of the Multi-Personality Generation (MPG) framework. MPG is a decoding-time paradigm that enables Large Language Models (LLMs) to simultaneously embody multiple personalization attributes without requiring extra training or multi-dimensional models. Dataset Description The collection includes several Direct Preference Optimization (DPO) datasets used for MBTI… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/dpo_profile.texttext-generation1K<n<10K0 likes33 downloads8mo agoHugging Face28ronalp2 /amazon_us_reviewsAmazon Customer Reviews (a.k.a. Product Reviews) is one of Amazons iconic products. In a period of over two decades since the first review in 1995, millions of Amazon customers have contributed over a hundred million reviews to express opinions and describe their experiences regarding products on the Amazon.com website. This makes Amazon Customer Reviews a rich source of information for academic researchers in the fields of Natural Language Processing (NLP), Information Retrieval (IR), and Machine Learning (ML), amongst others. Accordingly, we are releasing this data to further research in multiple disciplines related to understanding customer product experiences. Specifically, this dataset was constructed to represent a sample of customer evaluations and opinions, variation in the perception of a product across geographical regions, and promotional intent or bias in reviews. Over 130+ million customer reviews are available to researchers as part of this release. The data is available in TSV files in the amazon-reviews-pds S3 bucket in AWS US East Region. Each line in the data files corresponds to an individual review (tab delimited, with no quote and escape characters). Each Dataset contains the following columns: - marketplace: 2 letter country code of the marketplace where the review was written. - customer_id: Random identifier that can be used to aggregate reviews written by a single author. - review_id: The unique ID of the review. - product_id: The unique Product ID the review pertains to. In the multilingual dataset the reviews for the same product in different countries can be grouped by the same product_id. - product_parent: Random identifier that can be used to aggregate reviews for the same product. - product_title: Title of the product. - product_category: Broad product category that can be used to group reviews (also used to group the dataset into coherent parts). - star_rating: The 1-5 star rating of the review. - helpful_votes: Number of helpful votes. - total_votes: Number of total votes the review received. - vine: Review was written as part of the Vine program. - verified_purchase: The review is on a verified purchase. - review_headline: The title of the review. - review_body: The review text. - review_date: The date the review was written.summarization100M<n<1B0 likes33 downloads4d agoHugging Face29faur-ai /ro-no_robotsThis dataset is a translation of HuggingFaceH4/no_robots, using LLMic, a bilingual Romanian-English LLM. No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0). @misc{no_robots, author = {Nazneen Rajani and Lewis Tunstall and Edward Beeching and… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-no_robots.texttext-generation10K<n<100K0 likes29 downloads1y agoHugging Face30ronantakizawa /codeconfig Build/CI Configuration Corpus A curated dataset of build, CI/CD, and project configuration files from top GitHub repositories. Repositories are sourced from ronantakizawa/github-top-projects, which tracks GitHub's top repositories from 2013–2025. Use Cases Fine-tuning LLMs for DevOps/infrastructure code generation Training code completion models for configuration files Benchmarking LLM performance on build/CI tasks Schema Field Type Description… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codeconfig.tabulartext-generation10K<n<100K1 likes26 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.