datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TempleOS-Source-CodePython-source-codeopen_parallel_think_code_source
open_parallel_think_code_source
A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems.
Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.Algorithm_and_Python_Source_CodeAlgorithm_and_Python_Source_Code
This dataset provides different algorithms and their corresponding source code in Python.
credits: Source codes given here are taken from "iamtarun/python_code_instructions_18k_alpaca" dataset in Hugging Face.
github-source-code-dataset
Github Source Code Dataset
Complete source code from Agnuxo projects.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
source_code纯文本数据,内容:高质量编程源代码,包括Python,Java,CPP源代码buddi-asr-source-code
Buddi-ASR
A privacy-aware training and deployment pipeline for a small Arabic/Gulf-Arabic + English child-speech model.
فارسی · GPU training · Data governance · Experiment plan · Raspberry Pi
Buddi-ASR turns multilingual Whisper Tiny or Base into a domain model for short child utterances, then exports the merged checkpoint to a quantized whisper.cpp artifact for Raspberry Pi, mobile, or server inference.
This repository contains the reproducible pipeline, not fabricated model… See the full description on the dataset page: https://huggingface.co/datasets/NabuxAi/buddi-asr-source-code.python-algorithm-sourcecode
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset provides algorithms and corresponding Python source code which can be leveraged for any type of code conversion applications.
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/annawleo/python-algorithm-sourcecode.synthetic-sensitive-data-in-source-code-n300
Synthetic Sensitive Data in Source Code (N=300)
Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII.Every sample includes at least one sensitive finding (no clean negatives).
Designed for evaluating local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure scenarios in AI-assisted coding workflows.
Version 1.2: multi_secret (and related) samples label every secret present in code_text (complete ground truth).
All… See the full description on the dataset page: https://huggingface.co/datasets/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300.sourcecode-detectionc-java-source-code
🧩 Cross-Project Defect Prediction (CPDP) Dataset — C & Java Projects
This repository hosts a custom dataset for Cross-Project Defect Prediction (CPDP) research, curated from a diverse collection of real-world open-source projects written in C (441 projects) and Java (98 projects).The dataset aims to support research on software defect prediction, transfer learning, and imbalanced data handling across heterogeneous programming environments.
📘 Overview
Language… See the full description on the dataset page: https://huggingface.co/datasets/SuraviAkhter/c-java-source-code.RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.vuln-with-source-codeDoom-source-codeedusite-source-code
☀️ Sun Shine School — Narasapuram | Since 1982
A full-stack educational website for Sun Shine School, Narasapuram (West Godavari District, Andhra Pradesh), built with React (frontend) and Python/FastAPI (backend).
Content migrated from www.sunshineschoolnsp.com.
Pages
🏠 Home — About Sun Shine School, Education Excellence Award, academic programs overview
🌟 What We Offer — Smart Classes, Sports, Labs, Language Development, Hostel
🎨 Activities — Singing, Dancing… See the full description on the dataset page: https://huggingface.co/datasets/christopherlancei/edusite-source-code.source-code-Review-vulnProduct-Source-Code-DatasetDataset Description:
This dataset is a large-scale collection of coding and data, designed to support the development of advanced AI systems for code generation, program understanding, software intelligence, debugging assistance, and next-generation developer AI applications.
Additionally, this dataset can be integrated into pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, helping improve AI performance in code completion, automated… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Product-Source-Code-Dataset.SourceCodellama-inst-filtered-1k-source-code
Description
Slightly modified and formatted version of a subset of the original dataset for my own purpose.
Original Dataset
togethercomputer/llama-instruct · Datasets at Hugging Face
LICENSE
LLAMA 2 COMMUNITY LICENSE
Download Llama
Purr-Data_example_source_codesPurr-Data Patch Source Code Dataset:
This dataset is designed for training language models to generate source code for Purr-Data patches. It focuses specifically on patches that output a particular message when a "bang" object is clicked.
Dataset Creation:
The dataset was created with the goal of evaluating the ability of large language models like Google's 2B GEMMA to be fine-tuned for Purr-Data source code generation.
Dataset Characteristics:
Content: Each data point consists of two… See the full description on the dataset page: https://huggingface.co/datasets/ParZiVal04/Purr-Data_example_source_codes.LiteCoder_SourceCode
LiteCoder Experiment Reproducing package
To run the pre-train objective use the following scripts:
Reproduce LiteCoder with all objectives:
Navigate the folder Pre-training containing the LiteCoder.py file
Then, run Python LiteCoder.py --train-tt --train-cs --train-pd
The pretrained model is released on hugging face, therefore it automatically loads.
To run the ablation studies:
Ablation 1: Python LiteCoder.py --train-tt
Ablation 2: Python LiteCoder.py --train-tt… See the full description on the dataset page: https://huggingface.co/datasets/LiteCoder/LiteCoder_SourceCode.Javascript-source-codetigle-source-code
TIGLE
The interface is built as a prototype based on Dzogchen, Atiyoga teachings available in English and sourced, compiled by a practitioner exploring how Dharma language and current global AI could intersect. The architecture, the pipeline works. The answers are useful for orientation — learning key terms, lineages, main practices, understanding the view.
This is why it is accessible as repository rather than a product:
Digital Bardo - is the current state of samsara.… See the full description on the dataset page: https://huggingface.co/datasets/Tigle/tigle-source-code.EDUrocks-3.0-frontend-source-codeCode_Geass_Lelouch_of_the_Rebellion_Source_Videos
test-smells-with-source-codeOpen-Source-Modeling-of-Syngas-Production-from-Biomass-A-CodeDriven-Approachcode-source-des-taxes-foncieres-tf
Code source des taxes foncières (TF)
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Code source des taxes foncières (TF) qui est disponible à l'adresse https://www.data.gouv.fr/datasets/5d5f9fc96f4441094179a94c
Description
Le code source des taxes foncières est produit par la Direction Générale des Finances publiques. Ce code source est sous licence CeCILL v2.1(détail dans le fichier LICENSE.txt). Cette… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/code-source-des-taxes-foncieres-tf.hollow-knight-silk-song-sourcecode
