CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saeedzouashkiani /yodas_fa_nosub_metadatatabular100K<n<1M0 likes5.3k downloads28d agoHugging Face02bigcode /the-stack-metadata Dataset Card for The Stack Metadata Changelog Release Description v1.1 This is the first release of the metadata. It is for The Stack v1.1 v1.2 Metadata dataset matching The Stack v1.2 Dataset Summary This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories. Supported Tasks and Leaderboards The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.tabulartext-generation10B<n<100B10 likes4.7k downloads4y agoHugging Face03bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.8k downloads4y agoHugging Face04trojblue /danbooru2025-metadata 🎨 Danbooru 2025 Metadata Latest Post ID: 9,158,800 (as of Apr 16, 2025) 📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork. Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes: More consistent tag history tracking Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.imagetext-to-image1M<n<10M38 likes1.7k downloads1y agoHugging Face05smartcat /Amazon_Sample_Metadata_2023 Dataset Card for Dataset Name Original datasets can be found on: https://amazon-reviews-2023.github.io/ Dataset Details This dataset was made as sample of several datasets from the link above. Dataset Description This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors, Health and Personal Care, Amazon Clothing Shoes and Jewlery, Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.tabular1M<n<10M1 likes1.4k downloads2y agoHugging Face06mmarone /fineweb-edu-full-metadata[WIP] FineWeb-Edu with Metadata This repo contains 3 versions of the FineWeb-Edu v1 dataset: fwedu1-metaonly/ fwedu1-text-content-zstd/ fineweb-edu-1.0.0-meta-and-text/ These are all joinable via the hash column, which is xxhash64 in pyspark, calculated on the text column. This hash is unique for all instances in the dataset. For convenience, this join is done for you in the third table fwedu1-metaonly is just the metadata of the data exactly as it comes from the FineWeb-Edu v1… See the full description on the dataset page: https://huggingface.co/datasets/mmarone/fineweb-edu-full-metadata.tabular100M<n<1B0 likes1.3k downloads1y agoHugging Face07yufan /arxiv-metadata-2020-2026 arXiv Metadata, enriched (2020–2026) Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09, enriched with abstracts, citation counts, and Semantic Scholar identifiers, and organized as a two-level hierarchy: field of study → year. Unlike a bare title index, every record here carries the abstract, the full author list, citation counts, and the Semantic Scholar corpusId, so you can do retrieval, classification, citation analysis, and corpus building directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.tabulartext-retrieval1M<n<10M0 likes1.2k downloads3d agoHugging Face08masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes1k downloads3y agoHugging Face09bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes990 downloads3y agoHugging Face10librarian-bots /model_cards_with_metadatatabular100K<n<1M0 likes767 downloads14h agoHugging Face11librarian-bots /dataset_cards_with_metadatatabular100K<n<1M0 likes759 downloads14h agoHugging Face12ThetaCursed /danbooru-2026-clean-metadatatabular10M<n<100M7 likes549 downloads5mo agoHugging Face13macpaw-research /mac-app-store-apps-metadata Dataset Card for Macappstore Applications Metadata 📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule. Mac App Store Applications Metadata sourced by the public API. Curated by: MacPaw Way Ltd. Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.imagetabular-classification10K<n<100K10 likes449 downloads1mo agoHugging Face14Shio-Koube /Danbooru-2026-parquet-metadataimage10M<n<100M11 likes434 downloads8mo agoHugging Face15haydn-jones /PubMed-Metadata PubMed Metadata V2 Snapshot of PubMed citation metadata, augmented with PMC records that were not safely represented in PubMed's current PMID-to-PMCID mapping. The PubMed portion was built from the NLM PubMed 2026 annual baseline and daily update XML files through pubmed26n1592.xml.gz. Updates are applied in numeric order, revised records replace earlier versions, and citations whose latest event is a deletion are omitted. The dataset was built on 2026-08-16. The single train… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/PubMed-Metadata.tabulartext-retrieval10M<n<100M0 likes409 downloads26d agoHugging Face16ibragim-bad /github-repos-metadata-40M 📊 Metadata for 40 million GitHub repositories A cleaned, analysis-ready dataset with per-repository statistics aggregated from GH Archive events: stars, forks, pull requests, open issues, visibility, language signals, and more. Column names mirror the GH Archive / GitHub API semantics where possible. GitHub repo: https://github.com/ibragim-bad/github-repos-metadata-40M Source: GH Archive (public GitHub event stream). ✅ Projects with complementary ideas GitHub Repo… See the full description on the dataset page: https://huggingface.co/datasets/ibragim-bad/github-repos-metadata-40M.tabular10M<n<100M23 likes398 downloads8mo agoHugging Face17dataproc5 /intermediate-danbooru2025-metadata-prioritized dataproc5/intermediate-danbooru2025-metadata-prioritized (Auto-generated summary) Basic Info: Shape: 9113291 rows × 61 columns Total Memory Usage: 35.96 GB Duplicates: 7 (0.00%) Column Stats: Error generating stats table. Column Summaries: → file_url (object) Unique values: 8836519 → approver_id (float64) Min: 1.000, Max: 1215532.000, Mean: 318955.604, Std: 309967.830 → bit_flags (int64) Min: 0.000, Max: 3.000, Mean: 0.401… See the full description on the dataset page: https://huggingface.co/datasets/dataproc5/intermediate-danbooru2025-metadata-prioritized.image1M<n<10M1 likes395 downloads1y agoHugging Face18ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes391 downloads2mo agoHugging Face19Complementarity /gpqa-metadata-blind-answertabularn<1K0 likes387 downloads29d agoHugging Face20MirandaAbhilash /vqav2-full-metadatatabular100K<n<1M0 likes369 downloads6mo agoHugging Face21umd-vt-nyu /datacomp_recap_metadata2image100M<n<1B2 likes333 downloads2y agoHugging Face22argus-systems /pickup-carrot-remove-parquet-metadata-2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_stationary", "total_episodes": 21, "total_frames": 9383, "total_tasks": 1, "total_videos": 84, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:21" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata-2.tabularrobotics10K<n<100K0 likes281 downloads11mo agoHugging Face23Smith42 /galaxies_metadata Galaxy metadata for pairing with `smith42/galaxies' dataset Here we have metadata for ~8.5 million galaxies. This metadata can be paired with galaxy jpg cutouts from the DESI legacy survey DR8, the cut outs are found here: https://huggingface.co/datasets/Smith42/galaxies. I've split away 1% of the metadata into a test set, and 1% into a validation set. The remaining 98% of the metadata comprise the training set. Useful links Paper here:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/galaxies_metadata.tabular1M<n<10M2 likes274 downloads1y agoHugging Face24talkpl-ai /TalkPlayData-Challenge-Track-Metadatatabular10K<n<100K1 likes270 downloads3mo agoHugging Face25institutional /institutional-books-hl-metadata 📚 Institutional Books: Harvard Library (metadata-only version) Institutional Books is a growing corpus of public domain books. This release is comprised of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative. IDI Terms of Use for Early-Access. 983K books, published largely in the 19th and 20th centuries 242B o200k_base tokens 386M pages of text, available in both original… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-metadata.tabular100K<n<1M17 likes255 downloads1mo agoHugging Face26nick007x /Danbooru-2026-parquet-metadatatabular10M<n<100M0 likes218 downloads6mo agoHugging Face27argus-systems /pickup-carrot-remove-parquet-metadataThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_stationary", "total_episodes": 21, "total_frames": 9383, "total_tasks": 1, "total_videos": 84, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:21" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata.tabularrobotics10K<n<100K0 likes193 downloads11mo agoHugging Face28Mieaz /github-repos-metadata-ge3 GitHub All Repositories Metadata Dataset (>= 3 Stars) A comprehensive metadata dataset covering 5,289,726 public GitHub repositories with 3 or more stars (>= 3) spanning the history of GitHub from 2008 to 2026. Data Recency & Snapshot Notice [!NOTE] Snapshot Methodology: This dataset combines a comprehensive historical base archive (up to mid-2022) with continuous periodic crawler snapshots (2023 through 2026). Star Counts & Metrics: Star counts and repository… See the full description on the dataset page: https://huggingface.co/datasets/Mieaz/github-repos-metadata-ge3.tabulartext-retrieval1M<n<10M0 likes177 downloads26d agoHugging Face29yufan /amazon2023-item-metadata Amazon Reviews 2023 — Item Metadata (content features) Item content features (title, images, price, brand/store, categories, features, description, …) for five Amazon Reviews 2023 categories, aligned with the user-interaction splits in yufan/amazon2023-user-interactions. One config per category; each has a single train split with one item per line. Join to the interactions/sequences via parent_asin. Coverage is 100 % of the 5-core items in the companion dataset. from datasets… See the full description on the dataset page: https://huggingface.co/datasets/yufan/amazon2023-item-metadata.tabularother100K<n<1M0 likes161 downloads3mo agoHugging Face30enelpol /rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset. Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage. Metadata contains six separate categories, each in a dedicated column: Year of the publication (publish_year) Type of the publication (publish_type) Country of the publication - often correlated with the homeland of the authors (country) Number of pages (no_pages) Authors (authors) Keywords (keywords) tabularquestion-answering10K<n<100K2 likes150 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.