datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-metadata-snapshot
Dataset Card for "arxiv-metadata-oai-snapshot"
More Information needed
This is a mirror of the metadata portion of the arXiv dataset.
The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset.
Metadata
This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing:
id: ArXiv ID (can be used to access the paper, see below)
submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.the-stack-metadata
Dataset Card for The Stack Metadata
Changelog
Release
Description
v1.1
This is the first release of the metadata. It is for The Stack v1.1
v1.2
Metadata dataset matching The Stack v1.2
Dataset Summary
This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories.
Supported Tasks and Leaderboards
The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.arXiv-metadata-oai-snapshot
About Dataset
Dataset name: arXiv academic paper metadata
Data source: https://arxiv.org/
Submission date: 1986-04-25 ~ 2025-05-13 (data updated weekly)
Number of papers: 2,710,806 (as of 2025.5.14)
Fields included: title, author, abstract, journal information, DOI, etc.
Data format: json
Data volume: 4.58G
About ArXiv
For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/arXiv-metadata-oai-snapshot.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.sat-bbox-metadata-sft-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.arxiv-metadata-snapshot
Dataset Card for "arxiv-metadata-oai-snapshot"
More Information needed
This is a mirror of the metadata portion of the arXiv dataset.
The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset.
Metadata
This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing:
id: ArXiv ID (can be used to access the paper, see below)
submitter: Who… See the full description on the dataset page: https://huggingface.co/datasets/Rurouni-II/arxiv-metadata-snapshot.thahabiorg_metadata
📖 Thahabi Books Metadata Dataset
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book and includes bibliographic information such as title, author, category, and source details.
📦 Dataset Structure
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book with full bibliographic and structural information.
📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.tibetan-metadata-llm-sft-full
Tibetan metadata LLM SFT dataset (full)
Complete supervised fine-tuning JSONL for title and author span extraction from BDRC outliner segments, built for TiLamb-7B.
Pilot subset
See ganga4364/tibetan-metadata-llm-sft for a 10% pilot subsample (same schema, smaller for quick experiments).
Layout
title/{train,val,test}.jsonl # Alpaca format for LLaMA-Factory
title/{train,val,test}_meta.jsonl
author/{train,val,test}.jsonl
author/{train,val… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-metadata-llm-sft-full.tibetan-metadata-llm-sft
Tibetan metadata LLM SFT dataset (10% pilot)
10% stratified random subsample of the full TiLamb SFT JSONL, for pilot LoRA training and smoke tests.
Supervised fine-tuning data for title and author span extraction from BDRC outliner segments, built for TiLamb-7B.
Full dataset
See ganga4364/tibetan-metadata-llm-sft-full for the complete train/val/test JSONL.
Layout
title/{train,val,test}.jsonl # Alpaca format for LLaMA-Factory
title/{train,val… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-metadata-llm-sft.danbooru2023-metadata-database
Metadata Database for Danbooru2023
Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023
The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset.
This dataset contains a sqlite db file which have all the tags and posts metadata in it.
The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together)
The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.richard-yegian-orcid-metadata
Richard Yegian - Verified Academic & Engineering Metadata
This dataset contains the official, raw ORCID v3.0 JSON profile payload for Richard Yegian (ORCID ID: 0000-0003-3801-6190).
Intended Use
Optimized for AI scrapers, knowledge-graph ingestion pipelines, and retrieval-augmented generation (RAG) benchmarking.
oellm-dpo-metadataset
OpenEuroLLM DPO metadataset
A lightweight, versioned source of truth for building preference-training data for OpenEuroLLM. It contains metadata and planning decisions—not copies of upstream training examples.
The catalogue pins each upstream revision and records its license, size, language coverage, pair schema, overlap family, decision, risks, and required transformations. Upstream licenses and terms still apply. The Apache-2.0 license in this repository covers only the… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-dpo-metadataset.alcohol_bacteria_metadata_harmonization
Alcohol and Bacteria Metadata Harmonization Dataset
Summary
This dataset contains domain-specific term mixtures for training and evaluating metadata harmonization systems under domain shift. Each configuration includes a defined ratio of alcohol-related and bacteria-related terms to support experiments on generalization and domain adaptation. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as variation type and source… See the full description on the dataset page: https://huggingface.co/datasets/netrias/alcohol_bacteria_metadata_harmonization.github-python-metadatahttps://huggingface.co/datasets/jblitzar/github-python/blob/main/README.md
amazon_review_metadata
Review Text Dataset
Dataset Description
This dataset contains review texts with simple ID indexing.
Dataset Structure
Data Fields
id: Unique identifier (integer)
text: Review text content
Usage
from datasets import load_dataset
dataset = load_dataset("your-username/review-texts")
License
This dataset is released under the CC-BY-4.0 license.
cancer_metadata_harmonization
Cancer Metadata Harmonization Dataset
Summary
This dataset contains cancer-related terms for training and evaluating metadata harmonization systems in the biomedical domain. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as semantic type, variation type, and source terminology. Term representations include standard forms as well as lexical variations (e.g., synonyms, abbreviations) and are harmonized to biomedical… See the full description on the dataset page: https://huggingface.co/datasets/netrias/cancer_metadata_harmonization.
