datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TempleOS-Source-Codeopen_parallel_think_code_source
open_parallel_think_code_source
A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems.
Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.Algorithm_and_Python_Source_CodeAlgorithm_and_Python_Source_Code
This dataset provides different algorithms and their corresponding source code in Python.
credits: Source codes given here are taken from "iamtarun/python_code_instructions_18k_alpaca" dataset in Hugging Face.
github-source-code-dataset
Github Source Code Dataset
Complete source code from Agnuxo projects.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
source_code纯文本数据,内容:高质量编程源代码,包括Python,Java,CPP源代码python-algorithm-sourcecode
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset provides algorithms and corresponding Python source code which can be leveraged for any type of code conversion applications.
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/annawleo/python-algorithm-sourcecode.sourcecode-detectionRLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.vuln-with-source-codeDoom-source-codeedusite-source-code
☀️ Sun Shine School — Narasapuram | Since 1982
A full-stack educational website for Sun Shine School, Narasapuram (West Godavari District, Andhra Pradesh), built with React (frontend) and Python/FastAPI (backend).
Content migrated from www.sunshineschoolnsp.com.
Pages
🏠 Home — About Sun Shine School, Education Excellence Award, academic programs overview
🌟 What We Offer — Smart Classes, Sports, Labs, Language Development, Hostel
🎨 Activities — Singing, Dancing… See the full description on the dataset page: https://huggingface.co/datasets/christopherlancei/edusite-source-code.source-code-Review-vulnProduct-Source-Code-DatasetDataset Description:
This dataset is a large-scale collection of coding and data, designed to support the development of advanced AI systems for code generation, program understanding, software intelligence, debugging assistance, and next-generation developer AI applications.
Additionally, this dataset can be integrated into pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, helping improve AI performance in code completion, automated… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Product-Source-Code-Dataset.SourceCodellama-inst-filtered-1k-source-code
Description
Slightly modified and formatted version of a subset of the original dataset for my own purpose.
Original Dataset
togethercomputer/llama-instruct · Datasets at Hugging Face
LICENSE
LLAMA 2 COMMUNITY LICENSE
Download Llama
Purr-Data_example_source_codesPurr-Data Patch Source Code Dataset:
This dataset is designed for training language models to generate source code for Purr-Data patches. It focuses specifically on patches that output a particular message when a "bang" object is clicked.
Dataset Creation:
The dataset was created with the goal of evaluating the ability of large language models like Google's 2B GEMMA to be fine-tuned for Purr-Data source code generation.
Dataset Characteristics:
Content: Each data point consists of two… See the full description on the dataset page: https://huggingface.co/datasets/ParZiVal04/Purr-Data_example_source_codes.LiteCoder_SourceCode
LiteCoder Experiment Reproducing package
To run the pre-train objective use the following scripts:
Reproduce LiteCoder with all objectives:
Navigate the folder Pre-training containing the LiteCoder.py file
Then, run Python LiteCoder.py --train-tt --train-cs --train-pd
The pretrained model is released on hugging face, therefore it automatically loads.
To run the ablation studies:
Ablation 1: Python LiteCoder.py --train-tt
Ablation 2: Python LiteCoder.py --train-tt… See the full description on the dataset page: https://huggingface.co/datasets/LiteCoder/LiteCoder_SourceCode.tigle-source-code
TIGLE
The interface is built as a prototype based on Dzogchen, Atiyoga teachings available in English and sourced, compiled by a practitioner exploring how Dharma language and current global AI could intersect. The architecture, the pipeline works. The answers are useful for orientation — learning key terms, lineages, main practices, understanding the view.
This is why it is accessible as repository rather than a product:
Digital Bardo - is the current state of samsara.… See the full description on the dataset page: https://huggingface.co/datasets/Tigle/tigle-source-code.test-smells-with-source-codeFEEDBACK_BASED_SOURCE_CODE_GENERATION
