datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv_deep_learning_python_research_code_functions_summaries
Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries
Dataset Summary
AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.python-github-codearxiv_python_research_code
Dataset Card for "ArtifactAI/arxiv_python_research_code"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code
Dataset Summary
AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (4.13GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.rlvr-code-data-python-r1-format-filteredPython-React-Code-Datasetarxiv_deep_learning_python_research_code
ArXiv Deep Learning Python Research Code
A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.
Dataset Summary
Statistic
Value
Total files
391,496
Total size
1.49 GB
Source repos
34,099
Time span
ArXiv inception through July 2023
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.Dr-Zeon-Github-Python-Code-Dataset
Luck Spark 1B - High Quality Code Dataset
The first quality-scored, star-agnostic code dataset for training 1B MoE code models.
Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot.
Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.rlvr-code-data-python-r1github-python-code-fim
github python code fim
Generated from tomekkorbak/python-github-code, limiting context length to 8192.
python-code-instructions-18k-alpaca-standardized
Dataset Card for "python-code-instructions-18k-alpaca-standardized"
More Information needed
Python-Security-Code-Datasetd1_code_pythonthe_stack_dedup_python_hits_1_qsc_code_cate_autogend1_code_python_3k_eval_636d
mlfoundations-dev/d1_code_python_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
23.3
58.2
73.0
0.4
43.5
41.1
17.2
8.8
12.1
AIME24
Average Accuracy: 23.33% ± 2.00%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
33.33%
10
30
2
23.33%
7
30
3
23.33%
7
30
4
26.67%
8
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_python_3k_eval_636d.rlvr-code-data-python-r1-format-filteredcode_vulnerability_pythoncode.evol.instruct.wiz.oss_python.jsonpython-code-instructions-85k-mypo-qaqc
joshuasundance/python-code-instructions-85k-mypo QA/QC artifact
This dataset repo is a QA/QC derivative generated by myponline.
What is included
Root-level train.parquet / validation.parquet / test.parquet with full QA/QC annotations.
filtered_basic/ with rows that pass structural QA/QC checks.
filtered_strict/ with rows whose chosen side passes structural QA/QC plus standalone ruff and mypy --strict.
summary.json with aggregate counts and provenance.… See the full description on the dataset page: https://huggingface.co/datasets/joshuasundance/python-code-instructions-85k-mypo-qaqc.code_search_net_python_filtered_top50k
Dataset Card for "code_search_net_python_filtered_top50k"
More Information needed
python_code_instructions_18k_alpaca-standardized
Dataset Card for "python_code_instructions_18k_alpaca-standardized"
More Information needed
python_comment_code_ratio_08
Dataset Card for "python_comment_code_ratio_08"
More Information needed
SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
python-code-instructions-18k-alpaca-standardized_cluster_2
Dataset Card for "python-code-instructions-18k-alpaca-standardized_cluster_2"
More Information needed
code_search_net_python_processed_400k
Dataset Card for "code_search_net_python_processed_400k"
More Information needed
python-code-instructions-18k-alpaca-standardized_cluster_2_std
Dataset Card for "python-code-instructions-18k-alpaca-standardized_cluster_2_std"
More Information needed
python-code-instructions-18k-alpaca-standardized_cluster_7_std
Dataset Card for "python-code-instructions-18k-alpaca-standardized_cluster_7_std"
More Information needed
hero-run-4-code-sdg-prompts-python-fenced-n16
Hero Run 4 Code SDG Prompts (16x)
Python fenced-code prompt source for Hero Run 4 code SDG.
This dataset is a prompt-only source dataset for Marin synthetic data generation jobs. Each row is intended to receive exactly one model response in the generated_text column. Because each unique prompt is repeated in the source data, generation scripts do not need a loop over response indices.
Provenance
Source dataset: mlfoundations-dev/hero_run_4_code
Source revision:… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/hero-run-4-code-sdg-prompts-python-fenced-n16.python-code-instructions-18k-alpaca-standardized_cluster_5_std
Dataset Card for "python-code-instructions-18k-alpaca-standardized_cluster_5_std"
More Information needed
python-code-instructions-18k-alpaca-standardized_cluster_8
Dataset Card for "python-code-instructions-18k-alpaca-standardized_cluster_8"
More Information needed
code-contest-python-cleaned
