CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ronantakizawa /github-codereview Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to: Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.tabulartext-generation100K<n<1M62 likes1.4k downloads7mo agoHugging Face02ronaldcmz /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K0 likes1.1k downloads3mo agoHugging Face03ronantakizawa /moltbook Moltbook Dataset A dataset of posts and communities from Moltbook - a Reddit-style social platform designed for AI agents. NOTE: This dataset is a snapshot of Moltbook before it went viral and got flooded with inauthentic accounts such as humans and bots. Files File Records Description moltbook_posts.csv 6,105 All posts from the platform moltbook_submolts.csv 124 All communities (submolts) Dataset Insights Overview… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/moltbook.tabulartext-classification1K<n<10K56 likes196 downloads8mo agoHugging Face04ronantakizawa /leetcode-assembly LeetCode Assembly Dataset 441 LeetCode problems solved in C, compiled to assembly across 4 architectures, 2 compilers, and 4 optimization levels using GCC and Clang via the Godbolt Compiler Explorer API. Dataset Summary Stat Value Total rows 14,112 Unique problems 441 Architectures x86-64, AArch64, MIPS64, RISC-V 64 Compilers GCC 15.2, Clang 21.1.0 Optimization levels -O0, -O1, -O2, -O3 Compilation success rate 100% Difficulty split Easy: 98… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/leetcode-assembly.tabulartext-generation10K<n<100K12 likes127 downloads7mo agoHugging Face05ronaldcmz /claude-fable-5-claude-code claude-fable-5 Agent Traces It's worth noting that our team was working with Glint-Research to collect as much fable data as possible. These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data). For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/claude-fable-5-claude-code.tabulartext-generationn<1K0 likes108 downloads3mo agoHugging Face06Ronilos /PulseLM PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning Hung Manh Pham*   Jinyang Wu*   Xiao Ma   Yiming Zhang   Yixin Xu   Aaqib Saeed  Bin Zhu†   Zhou Pan†   Dong Ma† * Equal contribution    † Corresponding authors Introduction PulseLM is a multimodal framework that integrates PPG (Photoplethysmography) signal encoders with large language models for physiological signal understanding research. The project includes a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Ronilos/PulseLM.tabularquestion-answering1M<n<10M0 likes84 downloads6mo agoHugging Face07ronnieaban /alquran Dataset Terjemahan dan Tafsir Al-Quran Deskripsi Dataset Dataset ini berisi terjemahan Al-Quran dalam bahasa Indonesia beserta tafsirnya. Dataset ini dapat digunakan untuk berbagai tugas NLP seperti machine translation, text generation, dan text summarization. Fitur Utama Terjemahan Al-Quran: Teks Al-Quran dalam bahasa Arab beserta terjemahannya dalam bahasa Indonesia. Tafsir Al-Quran: Penjelasan atau interpretasi dari ayat-ayat Al-Quran dalam bahasa… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/alquran.tabulartext-generation1K<n<10K2 likes38 downloads2y agoHugging Face08ronadin /ishowspeed-streams IShowSpeed IRL Scene Descriptions 229,959 scene-level visual descriptions + spoken transcripts from 656 hours of IShowSpeed's IRL streams. The data covers two of IShowSpeed's flagship IRL tours: Speed Does America — 35-day non-stop livestream tour across 25 US states (Aug–Oct 2025). 55 stream segments. Speed Does Africa — 30-day, 20-country tour across the African continent (Dec 2025 – Jan 2026). 29 stream segments. Each video is split into 10-second windows; for every window we… See the full description on the dataset page: https://huggingface.co/datasets/ronadin/ishowspeed-streams.tabularvideo-text-to-text100K<n<1M2 likes38 downloads5mo agoHugging Face09ronantakizawa /codeconfig Build/CI Configuration Corpus A curated dataset of build, CI/CD, and project configuration files from top GitHub repositories. Repositories are sourced from ronantakizawa/github-top-projects, which tracks GitHub's top repositories from 2013–2025. Use Cases Fine-tuning LLMs for DevOps/infrastructure code generation Training code completion models for configuration files Benchmarking LLM performance on build/CI tasks Schema Field Type Description… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codeconfig.tabulartext-generation10K<n<100K1 likes26 downloads7mo agoHugging Face10ronantakizawa /japanese-trending-words Japanese Trending Words Dataset (2006-2025) Dataset Description This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades. Dataset Summary Total entries: 593 words Time period: 2006-2025 (20 years) Languages: Japanese with English translations Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.tabulartext-generationn<1K4 likes24 downloads10mo agoHugging Face11ronantakizawa /india-trending-words Google India Trending Words Dataset (2008-2021, 2023-2024) Dataset Description This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com). Dataset Summary Total Entries: 900 Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available) Categories: 18 unique tags Region: India Format: CSV Dataset Structure Data Fields word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.tabulartext-classificationn<1K2 likes12 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.