datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.github-issuesgithub-actionsgithub-issuesgithub-issuesgithub-issuesgithub-issuesgit-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.github-code-nanochatbpe-1B
github-code-nanochatbpe-1B
GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
1,000,000,000
val.bin
val
10,000,000
train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.github-issuesannotations_creators:
other
language_creators:
crowdsourced
languages:
en-US
licenses:
other-my-license
multilinguality:
monolingual
pretty_name: HuggingFace Github Issues
size_categories:
unknown
source_datasets:
original
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
cuda-nsys-training
Qwythos Nsight Systems Profiling Agent Dataset
Multi-turn GPU profiling agent trajectories for fine-tuning Qwythos-9B (and similar tool-calling models) on NVIDIA Nsight Systems (nsys) + CUDA-L1 / KernelBench workloads.
Generated autonomously on an RTX 5090 by the model itself driving real profiling tools for ~33 hours.
Code: ai-hpc/prof-dataset-gen
Stats
Split
Rows
Notes
train
5,884
Accepted episodes (quality ≥ 0.55)
eval
309
5% holdout from accepted… See the full description on the dataset page: https://huggingface.co/datasets/gittensor-model-hub/cuda-nsys-training.bhagavath-gita-telugugithub-issuesgit-commits
Dataset: dataset.jsonl
Auto-labeled commit dataset scraped from GitHub repositories. Each line is a JSON object representing one commit with extracted features and an inferred label.
Features
Field
Type
Description
Stats
text
string
Commit message first line, conventional prefix stripped
—
files_count
int
Number of files changed
mean 4.3, median 1, max 300
additions
int
Lines added
mean 88, median 6, max 187K
deletions
int
Lines deleted
mean 172… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/git-commits.bhagavad-gita-with_life_lesson
bhagavad-gita-lifelesson Dataset
A complete, high-fidelity dataset covering all 701 verses of the Bhagavad Gita titled bhagavad-gita-lifelesson. Each verse follows the strict format:
First: Sanskrit chanting
Then: Hindi meaning (हिन्दी अर्थ)
Then: Life lesson (जीवन-पाठ)
(Transliteration and English translation have been removed).
🎧 Example Representation (Verse 2.47)
🎧 Verse 2.47
First: Sanskrit chanting
कर्मण्येवाधिकारस्ते मा फलेषु कदाचन
मा… See the full description on the dataset page: https://huggingface.co/datasets/AkrGupta/bhagavad-gita-with_life_lesson.github-issuesgithub-issues
Dataset Card for GitHub Issues
Dataset Description
This dataset is created for the Hugging Face Datasets library course
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP… See the full description on the dataset page: https://huggingface.co/datasets/cylee/github-issues.github-issuesgithub_fetch_huggingface_terminal_9091_n3v8x2_derived_product_reviews
Product Reviews (Derived)
Aggregated and cleaned product review dataset derived from upstream sources.
Dataset ID: DRV-REVIEWS
Catalog: ghfht9091n3v8x2
Derived from: SRC-ALPHA, SRC-GAMMA
Records: 41,000
github-issuesgithub_fetch_huggingface_terminal_9091_n3v8x2_source_epsilon
Epsilon Tag Taxonomy
Taxonomy of product and category tags.
Dataset ID: SRC-EPSILON
Catalog: ghfht9091n3v8x2
Origin: Curated from an open knowledge-graph category graph
Records: 2,048
License: MIT
github-issuesgithub-issuesgithub-repo-scraper
GitHub Repo Scraper · Repositories, Stars, Topics & Languages
Scrape GitHub repositories by language, topic, star count, license, organization, and pushed date window. Returns clean structured repo metrics and metadata without authentication.
Rows in this dataset
2,481
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/github-repo-scraper/ — 2,071 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/github-repo-scraper.datasets-github-issues-2bhagavad-gita-verses-sanskrit-translations
Bhagavad Gita – Sanskrit, Transliteration & Multi-Commentary Dataset
A complete dataset of all 700 verses of the Bhagavad Gita, sourced directly from the open-source VedicScriptures API (MIT-licensed).This dataset includes:
📜 Original Sanskrit slokas
🔡 IAST transliteration
🌐 Multiple English & Hindi translations
🧠 Traditional commentaries from many teachers
🔢 Structured metadata (chapter, verse, IDs, authors)
This dataset is ideal for NLP, LLM fine-tuning, translation… See the full description on the dataset page: https://huggingface.co/datasets/Voider22/bhagavad-gita-verses-sanskrit-translations.github-issuestransformers-github-issuesgithub-issuesgithub-mini
github-mini
A dataset consisting of source code from GitHub repositories.
Dataset Curation
Source: GitHub repositories with 500 to 99,999+ stars.
Licenses: Filtered for MIT and Apache-2.0.
Focus: Source code across multiple programming languages.
Dataset Structure
Each record in the dataset contains the following fields:
repo_full_name: The full name of the repository (owner/name).
repo_url: Direct link to the GitHub repository.
stars: Number of stars at the… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/github-mini.
