CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Fraser /dream-coder Program Synthesis Data Generated program synthesis datasets used to train dreamcoder. Currently just supports text & list data. text1K<n<10K6 likes684 downloads4y agoHugging Face02code-rider /spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d tabular1K<n<10K2 likes244 downloads9mo agoHugging Face03Trelis /openassistant-deepseek-coder Chat Fine-tuning Dataset - OpenAssistant DeepSeek Coder This dataset allows for fine-tuning chat models using: B_INST = '\n### Instruction:\n' E_INST = '\n### Response:\n' BOS = '<|begin▁of▁sentence|>' EOS = '\n<|EOT|>\n' Sample Preparation: The dataset is cloned from TimDettmers, which itself is a subset of the Open Assistant dataset, which you can find here. This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples. The… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/openassistant-deepseek-coder.text10K<n<100K10 likes128 downloads3y agoHugging Face04vinsblack /CodeReality CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset ⚠️ Important Limitations ⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use. Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.tabulartext-generation1K<n<10K1 likes109 downloads1y agoHugging Face05Coder-Dragon /wikipedia-movies Wikipedia Movie Plots with Images. 30,000+ movies plot descriptions and images. Plot summary descriptions of movies scrapped from Wikipedia. Dataset is subset of this dataset. Content The dataset contains descriptions of 34,886 movies from around the world. Column descriptions are listed below: Release Year - Year in which the movie was released Title - Movie title Origin/Ethnicity - Origin of movie (i.e. American, Bollywood, Tamil, etc.) Director - Director(s) Genre -… See the full description on the dataset page: https://huggingface.co/datasets/Coder-Dragon/wikipedia-movies.textfeature-extraction10K<n<100K1 likes77 downloads3y agoHugging Face06Coder-Dragon /Indian-IPO-2006-2025This dataset contains information about Initial Public Offer (IPO) released in India from 2006-2025. Content Open DateClose DateListing DateFace ValueIssue PriceIssue SizeLot SizePrice Listing OnTotal Shares OfferedAnchor Investors Shared OfferedNII Shares OfferedQIB Shares OfferedOther Shares OfferedRII Shares OfferedMinimum InvestmentTotal SubscriptionQIB SubscriptionRII SubscriptionNII SubscriptionMarket Maker Shares Offered text1K<n<10K0 likes34 downloads1y agoHugging Face07XythicK /CodeReasoningPro CodeReasoningPro Dataset Summary CodeReasoningPro is a large-scale synthetic dataset comprising 1,785,725 competitive programming problems in Python, created by XythicK, an MLOps Engineer. Designed for supervised fine-tuning (SFT) of machine learning models for coding tasks, it draws inspiration from datasets like OpenCodeReasoning. The dataset includes problem statements, Python solutions, and reasoning explanations, covering algorithmic topics such as arrays, subarrays… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/CodeReasoningPro.texttext-generation1M<n<10M4 likes33 downloads1y agoHugging Face08ksjpswaroop /openassistant-deepseek-codertext10K<n<100K0 likes24 downloads3y agoHugging Face09npc-worldwide /enpisi-coder-data enpisi-coder RL dataset Judge-rated npcsh agent traces and derived RL training data for the enpisi-coder model family. Produced by scripts/rate_traces.py (LLM-as-judge) and scripts/analyze_ratings.py; built into SFT/DPO/GRPO/PPO splits by scripts/train_from_csv.py. Splits Split Rows Description rated_traces 900 Per-trace judge scores (correctness, tool_selection, efficiency, clarity, partial_credit, composite) tasks 100 Benchmark task definitions… See the full description on the dataset page: https://huggingface.co/datasets/npc-worldwide/enpisi-coder-data.tabular1K<n<10K0 likes23 downloads2mo agoHugging Face10coder242 /osworld_tasks_filestextn<1K0 likes20 downloads11mo agoHugging Face1177xiaoyuanzi8 /code_reviewertext10K<n<100K5 likes17 downloads3y agoHugging Face12codergautam /rebus-dataset |🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/codergautam/rebus-dataset.image1K<n<10K0 likes17 downloads7mo agoHugging Face1377xiaoyuanzi8 /code_reviewer_demotext1K<n<10K0 likes16 downloads3y agoHugging Face14coderGit /Eng-PidginBioData Eng-PidginBioData: English–Nigerian Pidgin Biology Translation Dataset This dataset is archived on Zenodo with DOI: https://doi.org/10.5281/zenodo.18888857 Dataset Summary Eng-PidginBioData is a domain-specific parallel corpus for English ↔ Nigerian Pidgin machine translation focused on biological and scientific texts. The dataset contains 2,300 sentence pairs extracted from open-source biological research papers and manually translated into Nigerian Pidgin. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/coderGit/Eng-PidginBioData.texttranslation1K<n<10K0 likes16 downloads7mo agoHugging Face15Zakariae-drabech /code-route-maroc-dataset 🚗 Code de la Route Marocain Dataset (Loi 52-05) Ce jeu de données regroupe des questions, réponses et textes juridiques formalisés pour l'entraînement de modèles de langage (LLM) sur la réglementation routière au Maroc. 📊 Origine des données Pipeline NLP Source : Récupéré depuis le projet Kaggle medaymanelkajdouhi/code-route-maroc-nlp. Format d'export : Fichier export_final.csv converti en data.csv. 🎯 Utilisation Ce dataset sert de support direct pour le… See the full description on the dataset page: https://huggingface.co/datasets/Zakariae-drabech/code-route-maroc-dataset.tabularn<1K0 likes16 downloads4mo agoHugging Face16dappalumbo91 /hermes-coder-surge-datasettext1K<n<10K0 likes14 downloads8mo agoHugging Face17VaraprasadTarunkumar /gitdiff_codereviewtextn<1K2 likes9 downloads2y agoHugging Face18Stepan4545 /code_risktabular1K<n<10K0 likes9 downloads5mo agoHugging Face19Nutanix /coderag-evaltabularn<1K0 likes8 downloads2y agoHugging Face20saad2002 /code-review-tone-classificationtext1K<n<10K0 likes7 downloads6mo agoHugging Face21timmaythetoolmann /code-refusal-for-abliterationgated code-refusal-for-abliteration Takes datasets of responses / refusals used for abliteration, and filters these down to programming-specific tasks for code models to be abliterated. Sources: https://github.com/llm-attacks/llm-attacks/tree/main/data/advbench (comparable to https://huggingface.co/datasets/mlabonne/harmful_behaviors ) Also see: https://github.com/AI-secure/RedCode/tree/main/dataset / https://huggingface.co/datasets/monsoon-nlp/redcode-hf for samples using Python… See the full description on the dataset page: https://huggingface.co/datasets/timmaythetoolmann/code-refusal-for-abliteration.textn<1K0 likes6 downloads4mo agoHugging Face22SnehaPriyaaMP /CodeResponsetextn<1K0 likes5 downloads2y agoHugging Face23VaraprasadTarunkumar /codereview_gitdifftextn<1K1 likes2 downloads2y agoHugging Face24coder2520 /EnglishKoreanTranslationtextn<1K0 likes2 downloads1y agoHugging Face25cbcodershop /codershoptext100K<n<1M0 likes2 downloads3mo agoHugging Face26AASHIK-CODER /Voyxa_finetunetext10K<n<100K0 likes1 downloads2y agoHugging Face27Coder19 /AI_interviewtextn<1K1 likes1 downloads2y agoHugging Face28Coder19 /new_mocktextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.