datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ngii-map-full-light
ngii-map-full-light
Light point/line extract from NGII 1/1000 topographic data for Korea.
Not for shipping into GitHub — use this Hugging Face dataset instead.
CRS
Korea_2000_Central_Belt_2010 projected meters [x, y]
Layers (per region under by_region/<region>/)
Layer
Description
C023
poles (전주/통신주)
C022
lights (가로등·보안등)
A002
roads (도로 중심선)
B001_tiny
building footprints <25 m² as centroids
B002
lines (구분/재질 라인)
Also:… See the full description on the dataset page: https://huggingface.co/datasets/SKPark1/ngii-map-full-light.latent_v1_fullrun_alpha3_04latent_v1_fullrun_alpha3_06latent_v1_fullrun_alpha3_13latent_v1_fullrun_alpha3_01latent_v1_fullrun_alpha2_04latent_v1_fullrun_alpha3_03latent_v1_fullrun_alpha2_01WebLINX-full
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead:
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models
💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.latent_v1_fullrun_alpha3_05latent_v1_fullrun_alpha1_08latent_v1_fullrun_alpha3_14latent_v1_fullrun_alpha3_10latent_v1_fullrun_alpha2_03yelp_review_full
Dataset Card for YelpReviewFull
Dataset Summary
The Yelp reviews dataset consists of reviews from Yelp.
It is extracted from the Yelp Dataset Challenge 2015 data.
Supported Tasks and Leaderboards
text-classification, sentiment-classification: The dataset is mainly used for text classification: given the text, predict the sentiment.
Languages
The reviews were mainly written in english.
Dataset Structure
Data Instances
A… See the full description on the dataset page: https://huggingface.co/datasets/Yelp/yelp_review_full.s2orc_full
S2ORC Full — Semantic Scholar Open Research Corpus
A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information.
Dataset Description
S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.latent_v1_fullrun_alpha3_08latent_v1_fullrun_alpha1_03latent_v1_fullrun_alpha1_09latent_v1_fullrun_alpha1_13latent_v1_fullrun_alpha1_06Shamela4_Full_DB
Shamela 4 — Full Islamic Library Corpus
A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text.
Dataset Structure
stage0_raw/
├── _meta/ # Cross-cutting metadata (Parquet + JSONL)
│ ├── extraction_manifest.json # Global extraction record
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.BioMysteryBench-full
BioMysteryBench (full set)
90 mystery-bioinformatics problems. Each problem provides anonymized
biological data files and asks a question that requires real analysis
(alignment, expression, variant calling, motif discovery, structure, etc.)
to answer — the source dataset cannot be looked up.
v11 (2026-07): 9 problems removed and 24 problems edited after an
answer-key audit — see CHANGELOG.md.
Contents
problems.csv / problems.parquet — one row per problem:
id —… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/BioMysteryBench-full.latent_v1_fullrun_alpha1_04latent_v1_fullrun_alpha1_11latent_v1_fullrun_alpha1_12latent_v1_fullrun_alpha1_07Total_Editing_Synthetic_Video_Albedo_Fullp2p-full-data
Open Pixel2Play (P2P) Full Dataset
Paper | GitHub | Project Page | Toy Dataset
The p2p-full-data dataset contains 8300+ hours of high-quality human annotated data, spanning across more than 40 popular 3D video games. All gameplay is recorded at 20 FPS by experienced players. Each frame is annotated with keyboard and mouse actions, and text instructionsare provided when available.
If you found the dataset helpful, please consider upvoting the paper so it can reach more people!… See the full description on the dataset page: https://huggingface.co/datasets/elefantai/p2p-full-data.konachan_full
konachan Full Dataset
This is the full dataset of konachan.com. And all the original images are maintained here.
Information
Images
There are 319040 images in total. The maximum ID of these images is 391069. Last updated at 2025-07-21 22:19:40 JST.
These are the information of recent 50 images:
id
filename
width
height
mimetype
tags
file_url
391069
391069.png
4774
2786
image/png
animal anthropomorphism azur_lane bird black_hair bondage building car… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/konachan_full.
