CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes2.4k downloads6mo agoHugging Face02codesignal /tsla-historic-pricestabular1K<n<10K2 likes1.6k downloads3y agoHugging Face03liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes435 downloads6mo agoHugging Face04av-codes /harbor-benchtabular10K<n<100K1 likes373 downloads2y agoHugging Face05p-doom /crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.tabular100K<n<1M5 likes301 downloads8mo agoHugging Face06APProjects /warn-act-notice-type-codes-crosswalk WARN Act notice-type codes — the crosswalk Every US state publishes WARN Act layoff notices with a free-text column saying what kind of event it is. The statute recognises two: a plant closing and a mass layoff. Across 48 states that column contains 552 distinct exact strings (531 once you fold case). This dataset is the crosswalk: one row per raw string, how many notices carry it, which states emit it, and what it normalizes to. The finding that matters 521 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.tabulartabular-classificationn<1K1 likes267 downloads15h agoHugging Face07code-rider /spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d tabular1K<n<10K2 likes244 downloads9mo agoHugging Face08Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes160 downloads4mo agoHugging Face09Ichlibitiche /appliancedb-error-codes-repair-database ApplianceDB: Home Appliance Error Codes & Ranked Repairs Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.tabular1K<n<10K0 likes141 downloads3d agoHugging Face10BambusControl /AI-Code-Optimization-for-Sustainability-Dataset AI Code Optimization for Sustainability: Dataset Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury 📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency. The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.tabular10K<n<100K0 likes121 downloads7mo agoHugging Face11vinsblack /CodeReality CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset ⚠️ Important Limitations ⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use. Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.tabulartext-generation1K<n<10K1 likes109 downloads1y agoHugging Face12code-switching /naturalnesstabularn<1K0 likes109 downloads21d agoHugging Face13codezerro /test-dataset-v1image100K<n<1M0 likes99 downloads1y agoHugging Face14CooperBench /qwen9b-coop-claude-code qwen9b-coop-claude-code Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo (single-agent) baseline is at CooperBench/qwen9b-solo-claude-code. Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.tabulartext-generationn<1K0 likes72 downloads4mo agoHugging Face15prokelly /neuromoyo-sahara-codeswitch-benchmark NEUROMOYO — Sahara CodeSwitch Africa Benchmark 🔗 Live Benchmark Results Interactive benchmark: https://www.neuromoyo.app/benchmark This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech. 🚀 Live NEUROMOYO Demo Live application: https://www.neuromoyo.app The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.tabularn<1K0 likes71 downloads20h agoHugging Face16CooperBench /qwen9b-solo-claude-code qwen9b-solo-claude-code Single-agent coding trajectories generated by running CooperBench in solo mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. One agent implements both features in each task. The matched coop (two-agent) version is at CooperBench/qwen9b-coop-claude-code. Same task corpus, same model, same agent — only the coordination differs, so together they isolate the cooperation deficit. At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.tabulartext-generationn<1K0 likes70 downloads4mo agoHugging Face17codewithdark /online-retail-refined-datasettabular100K<n<1M1 likes64 downloads1y agoHugging Face18SDAIANCAI /Saudilang-Code-Switch-Corpus SCC - Saudilang Code-Switch Corpus The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”. This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.tabularautomatic-speech-recognition1K<n<10K3 likes60 downloads2y agoHugging Face19codealchemist01 /letterboxd-movies Letterboxd Movies Dataset Dataset Description A comprehensive dataset of movies scraped from Letterboxd, including genres, ratings, runtime, countries, and detailed movie characteristics. This dataset contains 16246 movies with 28 features each, scraped from Letterboxd. It's perfect for: 🎬 Movie recommendation systems 📊 Film industry analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Movie discovery algorithms Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/letterboxd-movies.tabulartext-classification10K<n<100K1 likes50 downloads11mo agoHugging Face20data-sci-project /buildings_and_incomes_by_postal_codetabularn<1K0 likes44 downloads5d agoHugging Face21ronnieaban /hs-codetabulartext-classification1K<n<10K1 likes43 downloads2y agoHugging Face22anthony-code /gentle-meadow-412d55 gentle-meadow-412d55 Synthetic products test data: 33 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/anthony-code/gentle-meadow-412d55.tabularn<1K0 likes43 downloads12d agoHugging Face23armixz /2020_BRFSS_Codebook_CDCtabular100K<n<1M0 likes37 downloads4y agoHugging Face24codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes35 downloads11mo agoHugging Face25SciCode /SciCode-Domain-Codegated DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes34 downloads7mo agoHugging Face26ron-the-code /Telangana_time_series_2023-2025The dataset was retrieved from Open Data Telangana, from February 1, 2023, to January 31, 2025 with daily granularity. The dataset contains various fields such as District, Mandal, Date, rainfall (in millimeters), minimum and maximum temperature (in Celsius), minimum and maximum wind speed, and humidity. It provides a District and Mandal wise distribution as well. Total Rows - 4,45,213 Total Columns - 10 tabular100K<n<1M2 likes31 downloads2y agoHugging Face27ryen-stuff /Deepseek-code DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.tabulartext-generation1K<n<10K0 likes31 downloads1mo agoHugging Face28lucsaint /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face29CodeSoftHF /bs-mapsThe data about the built in maps in Beat Saber. Contains all OST, Camellia, and Extra songs. A couple of songpacks are added. tabularn<1K0 likes27 downloads3y agoHugging Face30hackerrank /example_annotated_code_repo_dataA description of the fields: Column What it captures Typical values id Row identifier 1-100 repo_name Example repository label repo_14 file_path Path + filename with extension src/utils/parsefile.py language Programming language Python, Java… function_name Target symbol that was reviewed validateSession annotation_summary Free-text note written by the annotator “Added input validation…” potential_bug Did the annotator flag a likely bug? (Yes/No)… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/example_annotated_code_repo_data.tabularn<1K0 likes25 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.