CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01librarian-bots /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.texttext-generation1M<n<10M22 likes4.8k downloads2d agoHugging Face02bigcode /the-stack-metadata Dataset Card for The Stack Metadata Changelog Release Description v1.1 This is the first release of the metadata. It is for The Stack v1.1 v1.2 Metadata dataset matching The Stack v1.2 Dataset Summary This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories. Supported Tasks and Leaderboards The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.tabulartext-generation10B<n<100B10 likes4.6k downloads4y agoHugging Face03jackkuo /arXiv-metadata-oai-snapshot About Dataset Dataset name: arXiv academic paper metadata Data source: https://arxiv.org/ Submission date: 1986-04-25 ~ 2025-05-13 (data updated weekly) Number of papers: 2,710,806 (as of 2025.5.14) Fields included: title, author, abstract, journal information, DOI, etc. Data format: json Data volume: 4.58G About ArXiv For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/arXiv-metadata-oai-snapshot.texttext-classification1M<n<10M0 likes559 downloads1y agoHugging Face04ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes388 downloads2mo agoHugging Face05NuTonic /sat-bbox-metadata-sft-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.imagetext-generation100K<n<1M4 likes346 downloads5mo agoHugging Face06Rurouni-II /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter: Who… See the full description on the dataset page: https://huggingface.co/datasets/Rurouni-II/arxiv-metadata-snapshot.texttext-generation1M<n<10M0 likes194 downloads5mo agoHugging Face07freococo /thahabiorg_metadata 📖 Thahabi Books Metadata Dataset This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book and includes bibliographic information such as title, author, category, and source details. 📦 Dataset Structure This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book with full bibliographic and structural information. 📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.tabulartext-generation1M<n<10M0 likes147 downloads3mo agoHugging Face08openpecha /tibetan-metadata-llm-sft-full Tibetan metadata LLM SFT dataset (full) Complete supervised fine-tuning JSONL for title and author span extraction from BDRC outliner segments, built for TiLamb-7B. Pilot subset See ganga4364/tibetan-metadata-llm-sft for a 10% pilot subsample (same schema, smaller for quick experiments). Layout title/{train,val,test}.jsonl # Alpaca format for LLaMA-Factory title/{train,val,test}_meta.jsonl author/{train,val,test}.jsonl author/{train,val… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-metadata-llm-sft-full.text-generation100K<n<1M0 likes137 downloads3mo agoHugging Face09openpecha /tibetan-metadata-llm-sft Tibetan metadata LLM SFT dataset (10% pilot) 10% stratified random subsample of the full TiLamb SFT JSONL, for pilot LoRA training and smoke tests. Supervised fine-tuning data for title and author span extraction from BDRC outliner segments, built for TiLamb-7B. Full dataset See ganga4364/tibetan-metadata-llm-sft-full for the complete train/val/test JSONL. Layout title/{train,val,test}.jsonl # Alpaca format for LLaMA-Factory title/{train,val… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-metadata-llm-sft.text-generation10K<n<100K0 likes98 downloads3mo agoHugging Face10KBlueLeaf /danbooru2023-metadata-databasegated Metadata Database for Danbooru2023 Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023 The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset. This dataset contains a sqlite db file which have all the tags and posts metadata in it. The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together) The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.imageimage-classification1M<n<10M83 likes68 downloads2y agoHugging Face11richard-yegian /richard-yegian-orcid-metadata Richard Yegian - Verified Academic & Engineering Metadata This dataset contains the official, raw ORCID v3.0 JSON profile payload for Richard Yegian (ORCID ID: 0000-0003-3801-6190). Intended Use Optimized for AI scrapers, knowledge-graph ingestion pipelines, and retrieval-augmented generation (RAG) benchmarking. texttext-generationn<1K0 likes55 downloads20d agoHugging Face12birgermoell /oellm-dpo-metadataset OpenEuroLLM DPO metadataset A lightweight, versioned source of truth for building preference-training data for OpenEuroLLM. It contains metadata and planning decisions—not copies of upstream training examples. The catalogue pins each upstream revision and records its license, size, language coverage, pair schema, overlap family, decision, risks, and required transformations. Upstream licenses and terms still apply. The Apache-2.0 license in this repository covers only the… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-dpo-metadataset.tabulartext-generationn<1K0 likes47 downloads9d agoHugging Face13netrias /alcohol_bacteria_metadata_harmonization Alcohol and Bacteria Metadata Harmonization Dataset Summary This dataset contains domain-specific term mixtures for training and evaluating metadata harmonization systems under domain shift. Each configuration includes a defined ratio of alcohol-related and bacteria-related terms to support experiments on generalization and domain adaptation. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as variation type and source… See the full description on the dataset page: https://huggingface.co/datasets/netrias/alcohol_bacteria_metadata_harmonization.texttext-generation1M<n<10M0 likes21 downloads1y agoHugging Face14jblitzar /github-python-metadatahttps://huggingface.co/datasets/jblitzar/github-python/blob/main/README.md texttext-generation10K<n<100K0 likes16 downloads1y agoHugging Face15Jyshen /amazon_review_metadata Review Text Dataset Dataset Description This dataset contains review texts with simple ID indexing. Dataset Structure Data Fields id: Unique identifier (integer) text: Review text content Usage from datasets import load_dataset dataset = load_dataset("your-username/review-texts") License This dataset is released under the CC-BY-4.0 license. texttext-generationn<1K1 likes13 downloads1y agoHugging Face16netrias /cancer_metadata_harmonization Cancer Metadata Harmonization Dataset Summary This dataset contains cancer-related terms for training and evaluating metadata harmonization systems in the biomedical domain. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as semantic type, variation type, and source terminology. Term representations include standard forms as well as lexical variations (e.g., synonyms, abbreviations) and are harmonized to biomedical… See the full description on the dataset page: https://huggingface.co/datasets/netrias/cancer_metadata_harmonization.texttext-generation100K<n<1M0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.