datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dream-coder
Program Synthesis Data
Generated program synthesis datasets used to train dreamcoder.
Currently just supports text & list data.
spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
openassistant-deepseek-coder
Chat Fine-tuning Dataset - OpenAssistant DeepSeek Coder
This dataset allows for fine-tuning chat models using:
B_INST = '\n### Instruction:\n'
E_INST = '\n### Response:\n'
BOS = '<|begin▁of▁sentence|>'
EOS = '\n<|EOT|>\n'
Sample Preparation:
The dataset is cloned from TimDettmers, which itself is a subset of the Open Assistant dataset, which you can find here. This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples.
The… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/openassistant-deepseek-coder.CodeReality
CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset
⚠️ Important Limitations
⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use.
Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.wikipedia-movies
Wikipedia Movie Plots with Images.
30,000+ movies plot descriptions and images.
Plot summary descriptions of movies scrapped from Wikipedia.
Dataset is subset of this dataset.
Content
The dataset contains descriptions of 34,886 movies from around the world. Column descriptions are listed below:
Release Year - Year in which the movie was released
Title - Movie title
Origin/Ethnicity - Origin of movie (i.e. American, Bollywood, Tamil, etc.)
Director - Director(s)
Genre -… See the full description on the dataset page: https://huggingface.co/datasets/Coder-Dragon/wikipedia-movies.Indian-IPO-2006-2025This dataset contains information about Initial Public Offer (IPO) released in India from 2006-2025.
Content
Open DateClose DateListing DateFace ValueIssue PriceIssue SizeLot SizePrice Listing OnTotal Shares OfferedAnchor Investors Shared OfferedNII Shares OfferedQIB Shares OfferedOther Shares OfferedRII Shares OfferedMinimum InvestmentTotal SubscriptionQIB SubscriptionRII SubscriptionNII SubscriptionMarket Maker Shares Offered
CodeReasoningPro
CodeReasoningPro
Dataset Summary
CodeReasoningPro is a large-scale synthetic dataset comprising 1,785,725 competitive programming problems in Python, created by XythicK, an MLOps Engineer. Designed for supervised fine-tuning (SFT) of machine learning models for coding tasks, it draws inspiration from datasets like OpenCodeReasoning. The dataset includes problem statements, Python solutions, and reasoning explanations, covering algorithmic topics such as arrays, subarrays… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/CodeReasoningPro.openassistant-deepseek-coderenpisi-coder-data
enpisi-coder RL dataset
Judge-rated npcsh agent traces and derived RL training data for the
enpisi-coder model family.
Produced by scripts/rate_traces.py (LLM-as-judge) and
scripts/analyze_ratings.py; built into SFT/DPO/GRPO/PPO splits by
scripts/train_from_csv.py.
Splits
Split
Rows
Description
rated_traces
900
Per-trace judge scores (correctness, tool_selection, efficiency, clarity, partial_credit, composite)
tasks
100
Benchmark task definitions… See the full description on the dataset page: https://huggingface.co/datasets/npc-worldwide/enpisi-coder-data.osworld_tasks_filescode_reviewerrebus-dataset
|🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/codergautam/rebus-dataset.code_reviewer_demoEng-PidginBioData
Eng-PidginBioData: English–Nigerian Pidgin Biology Translation Dataset
This dataset is archived on Zenodo with DOI:
https://doi.org/10.5281/zenodo.18888857
Dataset Summary
Eng-PidginBioData is a domain-specific parallel corpus for English ↔ Nigerian Pidgin machine translation focused on biological and scientific texts. The dataset contains 2,300 sentence pairs extracted from open-source biological research papers and manually translated into Nigerian Pidgin.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/coderGit/Eng-PidginBioData.code-route-maroc-dataset
🚗 Code de la Route Marocain Dataset (Loi 52-05)
Ce jeu de données regroupe des questions, réponses et textes juridiques formalisés pour l'entraînement de modèles de langage (LLM) sur la réglementation routière au Maroc.
📊 Origine des données
Pipeline NLP Source : Récupéré depuis le projet Kaggle medaymanelkajdouhi/code-route-maroc-nlp.
Format d'export : Fichier export_final.csv converti en data.csv.
🎯 Utilisation
Ce dataset sert de support direct pour le… See the full description on the dataset page: https://huggingface.co/datasets/Zakariae-drabech/code-route-maroc-dataset.hermes-coder-surge-datasetgitdiff_codereviewcode_riskcoderag-evalcode-review-tone-classificationcode-refusal-for-abliteration
code-refusal-for-abliteration
Takes datasets of responses / refusals used for abliteration,
and filters these down to programming-specific tasks for code models to be abliterated.
Sources:
https://github.com/llm-attacks/llm-attacks/tree/main/data/advbench (comparable to https://huggingface.co/datasets/mlabonne/harmful_behaviors )
Also see: https://github.com/AI-secure/RedCode/tree/main/dataset / https://huggingface.co/datasets/monsoon-nlp/redcode-hf for samples using Python… See the full description on the dataset page: https://huggingface.co/datasets/timmaythetoolmann/code-refusal-for-abliteration.CodeResponsecodereview_gitdiffEnglishKoreanTranslationcodershopVoyxa_finetuneAI_interviewnew_mock
