datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
loc_chronicling_america_1770-1810_issues
Dataset Card for Chronicling America Historic American Newspapers 1770–1810 - Issue-Level
Dataset Summary
A dataset drawn from the Library of Congress Chronicling America digital collection, part of the National Digital Newspaper Program (NDNP). This dataset provides an issue-level representation of the Chronicling America newspapers dataset, aggregating individual page records into complete newspaper issues with with publication metadata, original Chronicling… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810_issues.github-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.the-stack-github-issues
Dataset Description
This dataset contains conversations from GitHub issues and Pull Requests. Each conversation is comprised of a series of events, such as opening an issue, creating a comment,
or closing the issue, and includes the author's username, text, action, and identifiers such as the issue ID and number.
The dataset, which is mostly in English, has a total size of 54GB and 30.9M files.
Dataset Structure
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-github-issues.issues_prsopen-github-issues
OpenGitHub Issues
What is it?
The full development metadata of 7 public GitHub repositories, fetched from the GitHub REST API and GraphQL API, converted to Parquet and hosted here for easy access.
Right now the archive has 6.0M rows across 8 tables in 699.6 MB of Zstd-compressed Parquet. Every issue, pull request, comment, code review, timeline event, file change, and CI status check is stored as a separate table you can load individually or query together.
This… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github-issues.vertebrate-v1-issue473-fullwindow-cds-random-val
marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val
CDS full-window vertebrate projection sequences for the issue #473 random
validation control. The source is the immutable issue #417 accepted-sequence
table.
The split uniformly samples 16,384 original-orientation CDS rows
without replacement using seed 42. Sampling occurs before
reverse-complement augmentation. Selected rows are removed from training;
reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.github-issuesvertebrate-v1-issue473-center1-cds
marin-dna/vertebrate-v1-issue473-center1-cds
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
cds cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed #417 protein-coding CDS anchor catalog.
The source projection was produced by the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-cds.vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.github-issuesgithub-issuesvertebrate-v1-issue473-center1-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered catalog after… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered.github-issues-updated
📊 GitHub Issues Dataset (HuggingFace/datasets Repository)
This dataset contains structured GitHub issues scraped from the huggingface/datasets repository. It is intended for NLP tasks, topic modeling, issue classification, and software engineering research.
📌 Dataset Summary
Repository Source: huggingface/datasets
Scraped via: GitHub REST API v3
Total Issues: ~7,465
Collected On: June 25, 2025
Format: JSONL → loaded via Arrow for Hugging Face
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/rIsHu009/github-issues-updated.github-issuesHIL_v7_no-rewind-issueThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/grahamwichhh/HIL_v7_no-rewind-issue.pytorch-issues-dataset-cleanissuesgithub-issuesgithub-issuesgithub-issuesannotations_creators:
other
language_creators:
crowdsourced
languages:
en-US
licenses:
other-my-license
multilinguality:
monolingual
pretty_name: HuggingFace Github Issues
size_categories:
unknown
source_datasets:
original
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
github-issuesIssueBench
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
Authors:
Paul Röttger,
Musashi Hinck,
Valentin Hofmann,
Kobi Hackenburg,
Valentina Pyatkin,
Faeze Brahman, and
Dirk Hovy
Contact: paul.rottger@unibocconi.it
Using IssueBench
You can use IssueBench to measure issue bias in LLM writing assistance by following these steps:
Download the IssueBench prompts.
Generate completions using your LLM of choice.
Classify the stance of… See the full description on the dataset page: https://huggingface.co/datasets/Paul/IssueBench.indist-tool-v0-pool-v2-gpt55-1k_issue_rewritten_prompt-v3-v5_swesmith_metadata_repairedgithub-pytorch-issues
Dataset Card for github-pytorch-issues
Dataset Summary
This dataset is a curated collection of GitHub issues from the PyTorch repository. Each entry includes the issue title, body, user, state, labels, comments, and other relevant fields that are useful for tasks such as text classification, semantic search, and question answering.
Supported Tasks and Leaderboards
The dataset supports the following tasks:
Open-domain Question Answering: Given a user query… See the full description on the dataset page: https://huggingface.co/datasets/mayankpuvvala/github-pytorch-issues.github-issuesgithub-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/25b3nk/github-issues.tensorflow-issuesstreamlit-issues
Dataset Card for "streamlit-issues"
More Information needed
github-issues-negatives-maxMetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.
