datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/diamond-in/github-top-developers.github-repo-enumerationThis dataset was generated from GHArchive's Google BigQuery table.
It contains a list of every public repo (~380,000,000) committed to from January 2016 up to August 2024, as well as the number of unique contributors and
totals of the amounts of various events on those repositories in that time period.
This is useless on its own, but represents more than a few hours of effort and roughly $8 worth of cloud processing,
so I figured I would save the next person to try this some effort.
random-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.random-python-github-repositories
random-python-github-repositories
A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files.
Contents
repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash)
repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.github-top-projects
GitHub Trending Projects (2013-2025)
A comprehensive dataset of 423,098 GitHub trending repository entries spanning 12+ years (August 2013 - November 2025), scraped from Wayback Machine snapshots of GitHub's trending page.
🎯 Dataset Overview
This dataset captures the evolution of GitHub's trending repositories over time, providing insights into:
Software development trends across programming languages and domains
Popular open-source projects and their trending patterns… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-projects.Bhagavad-Gita_Dataset
Srimad Bhagavad Gita Dataset
A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks.
Dataset Details
Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi
Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita_Dataset.Bhagavad-Gita-QA
Bhagavad-Gita-QA-Multilingual
Dataset Summary
Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework.
This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita-QA.fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
A curated registry of points of interest in downtown Portland, Oregon.
License
This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY).
You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source.
Contents
data.csv - sample points of interest with coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001.git_good_bench
Dataset Summary
GitGoodBench Lite is a subset of 900 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios).
The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram.
This dataset thus contains 150 samples per sample type and programming language.
All data in this dataset are collected from 479 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench.nasa-science-github-repos
NASA Science GitHub Repositories
A curated index of 5,264 GitHub repositories relevant to the NASA Science Mission
Directorate (SMD), spanning five science divisions: Earth Science, Astrophysics,
Planetary Science, Heliophysics, and Biological & Physical Sciences.
This dataset is designed to support research on information retrieval and
discoverability of open-source scientific software.
Licensing and Intellectual Property
This dataset is released under CC-BY-4.0 and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-github-repos.github-readmesgit_good_bench-train
Dataset Summary
GitGoodBench Lite is a subset of 17469 samples for collecting trajectories of AI agents resolving git tasks (see Supported Scenarios) for model training purposes.
We support the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit chain.
All data in this dataset are collected from 816 unique, open-source GitHub repositories with permissive licenses
that have >= 1000 stars, >= 5 branches, >= 10 contributors and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-train.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-developers.github_fetch_huggingface_terminal_9013_novaretail_reviews_7d9ee1
Product Reviews Dataset
Customer reviews collected from NovaRetail's e-commerce platform between January and June 2026.
Contents
reviews.csv — 200 product reviews with rating, sentiment label, review text, and product category.
Reuse terms
This dataset is released under the Apache-2.0 license. Commercial use is permitted, and attribution is required.
git_good_bench-lite
Dataset Summary
GitGoodBench Lite is a subset of 120 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios).
The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram.
This dataset thus contains 20 samples per sample type and programming language.
All data in this dataset are collected from 100 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-lite.qwen35-9b-git-coop
What this is
Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with
mini_swe_agent on Qwen/Qwen3.5-9B in coop setting with a shared read-only git remote (--git).
Agents coordinate via messaging and git fetch team; patches are auto-merged after both submit.
At a glance
Field
Value
Model
Qwen/Qwen3.5-9B
Agent
mini_swe_agent (step_limit=300)
Setting
coop + git remote
Repos
18
Pairs
211
Both-pass
5.7% (12/210… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-git-coop.Gita-Train
Bhagavad-Gita-QA-Multilingual
Dataset Summary
Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework.
This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/rahul7star/Gita-Train.google-cloud_github_fetch_huggingface_terminal_6737_308aqj4zgithub-as-altmetric
About dataset
We construct this dataset for our study, which investigates the correlation between GitHub communication metrics and citation counts, examining the potential of the metrics as an altmetric.
Currently, it contains about 12,000 samples of publications, which are published by top-tier AI conferences.
The citation counts and the corresponding GitHub metrics might need to be updated.
We strive our best to update more conferences and keep values up-to-date.… See the full description on the dataset page: https://huggingface.co/datasets/deepkyu/github-as-altmetric.GitHub-issues-privacy-law-relevanceDataset with GitHub issues with reference to data privacy laws and indication on whether the issue is privacy-law relevant or not. The dataset was manually labeled.
github_fetch_huggingface_terminal_9091_n3v8x2_source_alpha
Alpha Commerce Reviews
Raw review text collected from public e-commerce listings.
Dataset ID: SRC-ALPHA
Catalog: ghfht9091n3v8x2
Origin: Public web scrape of product review pages (2023-2024)
Records: 12,480
License: MIT
google-cloud_github_fetch_huggingface_terminal_6737_3qt2dfa4algozee_analysis-of-high-starred-github-repositories
Analysis of High-Starred GitHub Repositories
A comprehensive overview of repository metrics and developer engagement
Dataset Info
Source: Kaggle
Original Size: 0.41 MB
Kaggle Downloads: 14
Files: 1
Files
github_top_repositories.csv
Mirrored from Kaggle
github-reposabdullahkhan70_github-tech-stack-languages-and-frameworks
GitHub Tech Stack Languages & Frameworks
Comprehensive Repository Data: JavaScript, Python, Go, Rust & More
Dataset Info
Source: Kaggle
Original Size: 2.17 MB
Kaggle Downloads: 62
Files: 17
Files
Mirrored from Kaggle
github-fetch-hf-9020-sales-orders
Anonymised Sales Orders
Data Fields
order_id: N/A
order_date: N/A
customer_id: N/A
product_id: N/A
quantity: N/A
unit_price: N/A
status: N/A
Citation
TBD
github_fetch_huggingface_terminal_9134_x9v2m6_derived_emotion_classifier
Emotion Classifier Data
A derived dataset used to train an emotion-classification model.
Overview
Dataset ID: DRV-EMOTION
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Origin: Derived from SRC-ALPHA and SRC-BETA with manual annotation
Records: 8,700
Product Line: Customer Feedback Analytics
Source
This dataset originates from: Derived from SRC-ALPHA and SRC-BETA with manual annotation.
Contents
Cleaned and re-labeled samples for emotion… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/github_fetch_huggingface_terminal_9134_x9v2m6_derived_emotion_classifier.fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testrun001
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testrun001
A curated registry of points of interest in downtown Portland, Oregon.
License
This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY).
You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source.
Contents
data.csv - sample data.
Bhagavad-Gita_Dataset
Srimad Bhagavad Gita Dataset
A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks.
Dataset Details
Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi
Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/Saptak123/Bhagavad-Gita_Dataset.google-cloud_github_fetch_huggingface_terminal_6737_bv8nvl1e
