datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
govdocs1-pdf-source
govdocs1: source PDF files
[!NOTE]
Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details
5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.gaia-dr3-source
Gaia DR3 Source
This dataset mirrors the complete ESA Gaia Data Release 3 gaia_source
bulk-download table. It contains one record for every published Gaia source and
all 152 columns served by ESA, including source identifiers, astrometry,
photometry, observing statistics, quality fields, classifications, and
astrophysical parameters.
Each of ESA's 3,386 compressed ECSV shards is retained as one dataset
configuration. The configuration name is the source filename stem verbatim.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-source.Know-Your-SourcesElephantBench-Source
ElephantBench Source Corpus
This dataset contains the low-quality partition ($D_{\mathrm{low}}$) used to construct ElephantBench. The released benchmark is available separately at panzs19/ElephantBench.
$D_{\mathrm{low}}$ is derived from cx-cmu/repro-organic-data-72B using the RePro fastText quality score. It contains 47,850,862 English web documents in 600 JSONL.zstd shards (approximately 65 GiB compressed). Each record retains the source text, URL, quality score, and original… See the full description on the dataset page: https://huggingface.co/datasets/panzs19/ElephantBench-Source.transformers-dependents
transformers metrics
This dataset contains metrics about the huggingface/transformers package.
Number of repositories in the dataset: 27067
Number of packages in the dataset: 823
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 65 packages that have more than 1000 stars.
There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.psi0-g1-sneaker-205ep-v2-source
Psi0 G1 Sneaker-in-Box — 205 episodes (v2 canonical source)
⚠️ Do not use this dataset directly for training. This is the canonical immutable union of the v1 and v2 collections, kept as a source of truth for reproducibility. For v2 fine-tuning use psi0-g1-sneaker-199ep-v2; for held-out evaluation use psi0-g1-sneaker-6ep-v2-eval. Together these two derivatives reconstruct this canonical dataset exactly: 199 + 6 = 205.
205 teleoperated episodes of a Unitree G1 humanoid (with Inspire… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/psi0-g1-sneaker-205ep-v2-source.gradio-dependents
Dataset Card for "gradio-dependents"
More Information needed
financial-english-source-corpus-gemma4-e2b-1280financial-english-source-corpus
Financial English Source Corpus
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. This
version preserves the final pre-split source rows.
Derived 1280-token split versions are available separately:
financial-english-source-corpus-qwen35-1280
financial-english-source-corpus-gemma4-e2b-1280
Dataset
Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.datasets-dependents
datasets metrics
This dataset contains metrics about the huggingface/datasets package.
Number of repositories in the dataset: 4997
Number of packages in the dataset: 215
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 22 packages that have more than 1000 stars.
There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.source-analysis
NuBerea Source Analysis
Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.evaluate-dependents
evaluate metrics
This dataset contains metrics about the huggingface/evaluate package.
Number of repositories in the dataset: 106
Number of packages in the dataset: 3
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 1 packages that have more than 1000 stars.
There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.accelerate-dependents
accelerate metrics
This dataset contains metrics about the huggingface/accelerate package.
Number of repositories in the dataset: 727
Number of packages in the dataset: 37
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 10 packages that have more than 1000 stars.
There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.secondary-sources
NuBerea/secondary-sources
Second Temple Jewish secondary sources in Greek: the complete extant Greek
corpora of Flavius Josephus (Jewish Antiquities, Jewish War, Vita, Contra
Apionem) and Philo of Alexandria (all 31 works), segmented for scholarly
text-retrieval and lexical-semantic study. These two first-century authors are
the principal non-biblical Jewish witnesses to the Second Temple period and its
milieu, and this repository serves as the Second Temple companion corpus to… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/secondary-sources.source-classifications
NuBerea Source Gold Set
Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists.
This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.diffusers-dependents
diffusers metrics
This dataset contains metrics about the huggingface/diffusers package.
Number of repositories in the dataset: 160
Number of packages in the dataset: 2
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.optimum-dependents
optimum metrics
This dataset contains metrics about the huggingface/optimum package.
Number of repositories in the dataset: 19
Number of packages in the dataset: 6
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.classifier_source
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.qmmit-open-source-agent-commit-index
Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures
Dataset release: 2026-09-18-v3.0Schema: 3.0.0
Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16
Abstract
This dataset contains 2000 repository-level observations from public Git
repositories. Each observation estimates a lower bound on the proportion of
non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.financial-english-source-corpus-qwen35-1280
Financial English Source Corpus Qwen35 1280
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. The
uploaded Parquet files are already prepared with the 1280-token source split
used by the downstream training pipeline.
This split version is derived from the pre-split
Financial English Source Corpus
by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.archon-comm-v1-sources-light
archon-comm-v1-sources-light
ARCHON v1 light sources -- 7 flow-level IDS datasets (CIC-IDS, UNSW-NB15 small, SIMARGL, network anomaly, witfoo SOC, etc.) -- ~68 GB total. For network flow training without raw packets. Use jescy525/archon-comm-v1-sources-packets for packet-level data.
Disclaimer
This bundle aggregates publicly available security research datasets for
defensive purposes only. See SOURCES.md for per-source license attribution.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/jescy525/archon-comm-v1-sources-light.neurips-2025-arxiv-latex-sources
NeurIPS 2025 arXiv LaTeX Source Files
This dataset contains file-level Parquet rows built from extracted raw arXiv
source packages for papers mapped from the NeurIPS 2025 proceedings to arXiv
records.
Each row is one file from one arXiv source package. Use arxiv_id to group files
back into papers.
Columns
arxiv_id: arXiv identifier for the source package.
title: paper title from the mapping CSV.
source_url: arXiv e-print source URL.
paper_index, paper_status… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/neurips-2025-arxiv-latex-sources.proofwriter-source
ProofWriter (The Source)
An unmodified copy of AI2's ProofWriter dataset (release V2020.12.3), re-hosted as datasets
configs for convenient loading. The records are faithful to the upstream release — the id-keyed JSON
is preserved as-is; typing and reasoning-graph extraction happen in later stages.
Each config is a {world}-depth-{n} shelf of the synthetic core (OWA/CWA × depths
0/1/2/3/5), split train/dev/test (dev kept as the corpus names it).
Source:… See the full description on the dataset page: https://huggingface.co/datasets/arqa39/proofwriter-source.news_sources_brazil
news_sources_brazil
17,892 news outlets, one per row, keyed to the IBGE municipality. 16,298 come from atlas da notícia, the survey of local journalism that projor and volt data lab have run since 2017. the other 1,594 are the portals, wire agencies, fact-checkers and international references that truw, my news verification project, was already tracking. the municipality table ships alongside, with 2022 population, HDI, outlet count and the news desert flag, so the whole thing… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/news_sources_brazil.common-corpus-sample-open-sourceopen_parallel_think_code_source
open_parallel_think_code_source
A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems.
Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.search-source-audit
Sources of Truth — AI Search Citations for Mental Health Queries
Which external sources do consumer AI search products actually cite when people ask about mental
health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform,
Multilingual Audit of Citations in AI Mental Health Information Queries."
Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.proofwriter-source
ProofWriter (The Source)
An unmodified copy of AI2's ProofWriter dataset (release V2020.12.3), re-hosted as datasets
configs for convenient loading. The records are faithful to the upstream release — the id-keyed JSON
is preserved as-is; typing and reasoning-graph extraction happen in later stages.
Each config is a {world}-depth-{n} shelf of the synthetic core (OWA/CWA × depths
0/1/2/3/5), split train/dev/test (dev kept as the corpus names it).
Source:… See the full description on the dataset page: https://huggingface.co/datasets/rlhf-and-friends/proofwriter-source.so101_mixed_3_sourcesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Long27/so101_mixed_3_sources.
