datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yodas_fa_nosub_metadatathe-stack-metadata
Dataset Card for The Stack Metadata
Changelog
Release
Description
v1.1
This is the first release of the metadata. It is for The Stack v1.1
v1.2
Metadata dataset matching The Stack v1.2
Dataset Summary
This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories.
Supported Tasks and Leaderboards
The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.c4-en-html-with-metadatadanbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.Amazon_Sample_Metadata_2023
Dataset Card for Dataset Name
Original datasets can be found on: https://amazon-reviews-2023.github.io/
Dataset Details
This dataset was made as sample of several datasets from the link above.
Dataset Description
This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors,
Health and Personal Care, Amazon Clothing Shoes and Jewlery,
Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.fineweb-edu-full-metadata[WIP]
FineWeb-Edu with Metadata
This repo contains 3 versions of the FineWeb-Edu v1 dataset:
fwedu1-metaonly/
fwedu1-text-content-zstd/
fineweb-edu-1.0.0-meta-and-text/
These are all joinable via the hash column, which is xxhash64 in pyspark, calculated on the text column. This hash is unique for all instances in the dataset. For convenience, this join is done for you in the third table
fwedu1-metaonly is just the metadata of the data exactly as it comes from the FineWeb-Edu v1… See the full description on the dataset page: https://huggingface.co/datasets/mmarone/fineweb-edu-full-metadata.arxiv-metadata-2020-2026
arXiv Metadata, enriched (2020–2026)
Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09,
enriched with abstracts, citation counts, and Semantic Scholar identifiers, and
organized as a two-level hierarchy: field of study → year.
Unlike a bare title index, every record here carries the abstract, the
full author list, citation counts, and the Semantic Scholar corpusId,
so you can do retrieval, classification, citation analysis, and corpus building
directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.c4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.c4-en-html-with-training_metadata_allmodel_cards_with_metadatadataset_cards_with_metadatadanbooru-2026-clean-metadatamac-app-store-apps-metadata
Dataset Card for Macappstore Applications Metadata
📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications Metadata sourced by the public API.
Curated by: MacPaw Way Ltd.
Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.Danbooru-2026-parquet-metadataPubMed-Metadata
PubMed Metadata V2
Snapshot of PubMed citation metadata,
augmented with PMC records that were not safely represented in PubMed's current
PMID-to-PMCID mapping.
The PubMed portion was built from the NLM PubMed 2026 annual baseline and daily
update XML files through pubmed26n1592.xml.gz. Updates are applied in numeric
order, revised records replace earlier versions, and citations whose latest event
is a deletion are omitted. The dataset was built on 2026-08-16.
The single train… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/PubMed-Metadata.github-repos-metadata-40M
📊 Metadata for 40 million GitHub repositories
A cleaned, analysis-ready dataset with per-repository statistics aggregated from GH Archive events: stars, forks, pull requests, open issues, visibility, language signals, and more. Column names mirror the GH Archive / GitHub API semantics where possible.
GitHub repo: https://github.com/ibragim-bad/github-repos-metadata-40M
Source: GH Archive (public GitHub event stream).
✅ Projects with complementary ideas
GitHub Repo… See the full description on the dataset page: https://huggingface.co/datasets/ibragim-bad/github-repos-metadata-40M.intermediate-danbooru2025-metadata-prioritized
dataproc5/intermediate-danbooru2025-metadata-prioritized
(Auto-generated summary)
Basic Info:
Shape: 9113291 rows × 61 columns
Total Memory Usage: 35.96 GB
Duplicates: 7 (0.00%)
Column Stats:
Error generating stats table.
Column Summaries:
→ file_url (object)
Unique values: 8836519
→ approver_id (float64)
Min: 1.000, Max: 1215532.000, Mean: 318955.604, Std: 309967.830
→ bit_flags (int64)
Min: 0.000, Max: 3.000, Mean: 0.401… See the full description on the dataset page: https://huggingface.co/datasets/dataproc5/intermediate-danbooru2025-metadata-prioritized.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.gpqa-metadata-blind-answervqav2-full-metadatadatacomp_recap_metadata2pickup-carrot-remove-parquet-metadata-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata-2.galaxies_metadata
Galaxy metadata for pairing with `smith42/galaxies' dataset
Here we have metadata for ~8.5 million galaxies.
This metadata can be paired with galaxy jpg cutouts from the DESI legacy survey DR8,
the cut outs are found here: https://huggingface.co/datasets/Smith42/galaxies.
I've split away 1% of the metadata into a test set, and 1% into a validation set.
The remaining 98% of the metadata comprise the training set.
Useful links
Paper here:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/galaxies_metadata.TalkPlayData-Challenge-Track-Metadatainstitutional-books-hl-metadata
📚 Institutional Books: Harvard Library (metadata-only version)
Institutional Books is a growing corpus of public domain books. This release is comprised of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative. IDI Terms of Use for Early-Access.
983K books, published largely in the 19th and 20th centuries
242B o200k_base tokens
386M pages of text, available in both original… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-metadata.Danbooru-2026-parquet-metadatapickup-carrot-remove-parquet-metadataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata.github-repos-metadata-ge3
GitHub All Repositories Metadata Dataset (>= 3 Stars)
A comprehensive metadata dataset covering 5,289,726 public GitHub repositories with 3 or more stars (>= 3) spanning the history of GitHub from 2008 to 2026.
Data Recency & Snapshot Notice
[!NOTE]
Snapshot Methodology: This dataset combines a comprehensive historical base archive (up to mid-2022) with continuous periodic crawler snapshots (2023 through 2026).
Star Counts & Metrics: Star counts and repository… See the full description on the dataset page: https://huggingface.co/datasets/Mieaz/github-repos-metadata-ge3.amazon2023-item-metadata
Amazon Reviews 2023 — Item Metadata (content features)
Item content features (title, images, price, brand/store, categories,
features, description, …) for five Amazon Reviews 2023 categories, aligned with
the user-interaction splits in
yufan/amazon2023-user-interactions.
One config per category; each has a single train split with one item per
line. Join to the interactions/sequences via parent_asin. Coverage is
100 % of the 5-core items in the companion dataset.
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/yufan/amazon2023-item-metadata.rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
