CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SKPark1 /ngii-map-full-light ngii-map-full-light Light point/line extract from NGII 1/1000 topographic data for Korea. Not for shipping into GitHub — use this Hugging Face dataset instead. CRS Korea_2000_Central_Belt_2010 projected meters [x, y] Layers (per region under by_region/<region>/) Layer Description C023 poles (전주/통신주) C022 lights (가로등·보안등) A002 roads (도로 중심선) B001_tiny building footprints <25 m² as centroids B002 lines (구분/재질 라인) Also:… See the full description on the dataset page: https://huggingface.co/datasets/SKPark1/ngii-map-full-light.geospatialother1M<n<10M0 likes230k downloads14d agoHugging Face02atokforps /latent_v1_fullrun_alpha3_040 likes81k downloads4y agoHugging Face03atokforps /latent_v1_fullrun_alpha3_060 likes42k downloads4y agoHugging Face04atokforps /latent_v1_fullrun_alpha3_130 likes36k downloads4y agoHugging Face05atokforps /latent_v1_fullrun_alpha3_010 likes35k downloads4y agoHugging Face06atokforps /latent_v1_fullrun_alpha2_040 likes33k downloads4y agoHugging Face07atokforps /latent_v1_fullrun_alpha3_030 likes30k downloads4y agoHugging Face08atokforps /latent_v1_fullrun_alpha2_010 likes30k downloads4y agoHugging Face09McGill-NLP /WebLINX-full WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead: WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models 💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.text10K<n<100K8 likes28k downloads1y agoHugging Face10atokforps /latent_v1_fullrun_alpha3_050 likes28k downloads4y agoHugging Face11atokforps /latent_v1_fullrun_alpha1_080 likes26k downloads4y agoHugging Face12atokforps /latent_v1_fullrun_alpha3_140 likes22k downloads4y agoHugging Face13atokforps /latent_v1_fullrun_alpha3_100 likes21k downloads4y agoHugging Face14atokforps /latent_v1_fullrun_alpha2_030 likes19k downloads4y agoHugging Face15Yelp /yelp_review_full Dataset Card for YelpReviewFull Dataset Summary The Yelp reviews dataset consists of reviews from Yelp. It is extracted from the Yelp Dataset Challenge 2015 data. Supported Tasks and Leaderboards text-classification, sentiment-classification: The dataset is mainly used for text classification: given the text, predict the sentiment. Languages The reviews were mainly written in english. Dataset Structure Data Instances A… See the full description on the dataset page: https://huggingface.co/datasets/Yelp/yelp_review_full.texttext-classification100K<n<1M149 likes19k downloads3y agoHugging Face16AlgorithmicResearchGroup /s2orc_full S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.texttext-generation10M<n<100M2 likes19k downloads5mo agoHugging Face17atokforps /latent_v1_fullrun_alpha3_080 likes19k downloads4y agoHugging Face18atokforps /latent_v1_fullrun_alpha1_030 likes17k downloads4y agoHugging Face19atokforps /latent_v1_fullrun_alpha1_090 likes16k downloads4y agoHugging Face20atokforps /latent_v1_fullrun_alpha1_130 likes15k downloads4y agoHugging Face21atokforps /latent_v1_fullrun_alpha1_060 likes14k downloads4y agoHugging Face22AuthenticIlm /Shamela4_Full_DB Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.text-generation10M<n<100M28 likes14k downloads4mo agoHugging Face23Anthropic /BioMysteryBench-fullgated BioMysteryBench (full set) 90 mystery-bioinformatics problems. Each problem provides anonymized biological data files and asks a question that requires real analysis (alignment, expression, variant calling, motif discovery, structure, etc.) to answer — the source dataset cannot be looked up. v11 (2026-07): 9 problems removed and 24 problems edited after an answer-key audit — see CHANGELOG.md. Contents problems.csv / problems.parquet — one row per problem: id —… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/BioMysteryBench-full.57 likes14k downloads3mo agoHugging Face24atokforps /latent_v1_fullrun_alpha1_040 likes13k downloads4y agoHugging Face25atokforps /latent_v1_fullrun_alpha1_111 likes12k downloads4y agoHugging Face26atokforps /latent_v1_fullrun_alpha1_120 likes12k downloads4y agoHugging Face27atokforps /latent_v1_fullrun_alpha1_070 likes12k downloads4y agoHugging Face28tonyc54 /Total_Editing_Synthetic_Video_Albedo_Full0 likes12k downloads1y agoHugging Face29elefantai /p2p-full-data Open Pixel2Play (P2P) Full Dataset Paper | GitHub | Project Page | Toy Dataset The p2p-full-data dataset contains 8300+ hours of high-quality human annotated data, spanning across more than 40 popular 3D video games. All gameplay is recorded at 20 FPS by experienced players. Each frame is annotated with keyboard and mouse actions, and text instructionsare provided when available. If you found the dataset helpful, please consider upvoting the paper so it can reach more people!… See the full description on the dataset page: https://huggingface.co/datasets/elefantai/p2p-full-data.image-text-to-text1M<n<10M23 likes11k downloads6mo agoHugging Face30deepghs /konachan_fullgated konachan Full Dataset This is the full dataset of konachan.com. And all the original images are maintained here. Information Images There are 319040 images in total. The maximum ID of these images is 391069. Last updated at 2025-07-21 22:19:40 JST. These are the information of recent 50 images: id filename width height mimetype tags file_url 391069 391069.png 4774 2786 image/png animal anthropomorphism azur_lane bird black_hair bondage building car… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/konachan_full.image-classification100K<n<1M15 likes11k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.