CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.tabulartext-classification1B<n<10B190 likes3.3m downloads2y agoHugging Face02m-a-p /PIN-200M PIN-200M A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents" Paper: https://arxiv.org/abs/2406.13923 This dataset contains around 200M samples in PIN format, with around 312 TB storage. 🚀 News [ 2025.09.22 ] !NEW! 🔥 We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it. [ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/PIN-200M.text10K<n<100K26 likes245k downloads5mo agoHugging Face03SKPark1 /ngii-map-full-light ngii-map-full-light Light point/line extract from NGII 1/1000 topographic data for Korea. Not for shipping into GitHub — use this Hugging Face dataset instead. CRS Korea_2000_Central_Belt_2010 projected meters [x, y] Layers (per region under by_region/<region>/) Layer Description C023 poles (전주/통신주) C022 lights (가로등·보안등) A002 roads (도로 중심선) B001_tiny building footprints <25 m² as centroids B002 lines (구분/재질 라인) Also:… See the full description on the dataset page: https://huggingface.co/datasets/SKPark1/ngii-map-full-light.geospatialother1M<n<10M0 likes230k downloads14d agoHugging Face04m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes57k downloads2y agoHugging Face05m-a-p /COIG-CQIA COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning Dataset Details Dataset Description 欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。 Welcome to the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/COIG-CQIA.textquestion-answering10K<n<100K775 likes36k downloads2y agoHugging Face06m-a-p /CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement [🏠Homepage] | [🛠️Code] OpenCodeInterpreter OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities. For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.textquestion-answering100K<n<1M208 likes28k downloads3y agoHugging Face07m-a-p /Matrix Matrix An open-source pretraining dataset containing 4690 billion tokens, this bilingual dataset with both English and Chinese texts is used for training neo models. Dataset Composition The dataset consists of several components, each originating from different sources and serving various purposes in language modeling and processing. Below is a brief overview of each component: Common Crawl Extracts from the Common Crawl project, featuring a rich diversity of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Matrix.texttext-generation1B<n<10B176 likes18k downloads2y agoHugging Face08sraimund /MapPool MapPool - Bubbling up an extremely large corpus of maps for AI MapPool is a dataset of 75 million potential maps and textual captions. It has been derived from CommonPool, a dataset consisting of 12 billion text-image pairs from the Internet. The images have been encoded by a vision transformer and classified into maps and non-maps by a support vector machine. This approach outperforms previous models and yields a validation accuracy of 98.5%. The MapPool dataset may help to train… See the full description on the dataset page: https://huggingface.co/datasets/sraimund/MapPool.image10M<n<100M5 likes18k downloads2y agoHugging Face09m-a-p /SuperGPQAThis repository contains the data presented in SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines. Tutorials for submitting to the official leadboard coming soon 📜 License SuperGPQA is a composite dataset that includes both original content and portions of data derived from other sources. The dataset is made available under the Open Data Commons Attribution License (ODC-BY), which asserts no copyright over the underlying content. This means that while the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/SuperGPQA.text10K<n<100K93 likes12k downloads1y agoHugging Face10bettergovph /project-noah-hazard-maps Project NOAH Hazard Maps Dataset Summary The Project NOAH (Nationwide Operational Assessment of Hazards) Hazard Maps is a comprehensive collection of geospatial datasets covering natural hazard assessments across the Philippines. The dataset includes three major hazard types: Flood Hazard Maps - Flood inundation maps for 5-year, 25-year, and 100-year rainfall return periods Landslide Hazard Maps - Shallow landslide susceptibility, structurally-controlled landslide… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/project-noah-hazard-maps.geospatialother2 likes11k downloads5mo agoHugging Face11incognitolm /USA-Map-Tiles Dataset Description The dataset is a set of map tiles for the various states of the USA. License: Open Data Commons Open Database License (ODbL) v1.0 Dataset Sources OpenStreetMap.org Geofabrik Download Server Bounding Boxes per State Dataset Structure Organized into folders by state Structure of graphml files: Top of files have a list of keys/ids that correspond to the properties of the segment: maxspeed: speed limit oneway: if it is a… See the full description on the dataset page: https://huggingface.co/datasets/incognitolm/USA-Map-Tiles.0 likes10k downloads8mo agoHugging Face12m-a-p /PIN-14M PIN-14M A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents" Paper: https://arxiv.org/abs/2406.13923 This dataset contains 14M samples in PIN format, with around 18.79 TB storage. 🚀 News [ 2025.09.04 ] !NEW! 🔥 We have completed the final version of the PIN-14M dataset and conducted some simple statistics on it. [ 2024.12.12 ] !NEW! 🔥 We have updated the quality signals for all subsets, with the dataset now containing 7.33B tokens… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/PIN-14M.text10K<n<100K38 likes9.8k downloads1y agoHugging Face13facebook /map-anything MapAnything Training Metadata Dataset Dataset Description This dataset contains pre-computed metadata and covisibility matrices for supporting the MapAnything codebase. This metadata enables easy reproducible training for feed-forward 3D reconstruction tasks. Please see our Data Processing README for more details. Citation If you use this dataset in your research, please cite our paper: @inproceedings{keetha2026mapanything, title={{MapAnything}: Universal… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything.image-to-3d100B<n<1T9 likes8.2k downloads8mo agoHugging Face14google /MapTrace MapTrace: A 2M-Sample Synthetic Dataset for Path Tracing on Maps Welcome to the MapTrace dataset! If you use this dataset in your work, please cite our paper below. For more details about our methodology and findings, please visit our project page or read the official white paper. This work was also recently featured on the Google Research Blog. Code & Scripts Official training and data loading scripts are available in our GitHub repository:… See the full description on the dataset page: https://huggingface.co/datasets/google/MapTrace.textimage-to-text10K<n<100K117 likes8k downloads7mo agoHugging Face15m-a-p /MTG2 likes7.2k downloads1y agoHugging Face16m-a-p /MAP-CC MAP-CC 🌐 Homepage | 🤗 MAP-CC | 🤗 CHC-Bench | 🤗 CT-LLM | 📖 arXiv | GitHub An open-source Chinese pretraining dataset with a scale of 800 billion tokens, offering the NLP community high-quality Chinese pretraining data. Disclaimer This model, developed for academic purposes, employs rigorously compliance-checked training data to uphold the highest standards of integrity and compliance. Despite our efforts, the inherent complexities of data and the broad spectrum of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MAP-CC.image1B<n<10B83 likes6.6k downloads2y agoHugging Face17PNW-GM /42_map_datasetimage10K<n<100K0 likes6.6k downloads3mo agoHugging Face18m-a-p /OProofs OProofs Formal Lean 4 theorem-proof pairs produced as part of the OProver project. Fields Field Type Description formal_statement string Lean 4 theorem statement formal_proof string Lean 4 proof body cot_proof string | null Chain-of-thought reasoning preceding the proof, if available prompt string | null Generation prompt, if available Stats Records: 6,804,694 Files: 73 parquet shards (zstd compressed) Loading from… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/OProofs.texttext-generation1M<n<10M4 likes6.3k downloads4mo agoHugging Face19GD-ML /MAPBench-V2For more details, please check our project page. Paper: https://arxiv.org/abs/2601.05432 Repository: https://github.com/AMAP-ML/Thinking-with-Map image1K<n<10K4 likes5.7k downloads8mo agoHugging Face20m-a-p /SciMMIR Dataset Card for "SciMMIR_dataset" SciMMIR This is the repo for the paper SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval. In this paper, we propose a novel SciMMIR benchmark and a corresponding dataset designed to address the gap in evaluating multi-modal information retrieval (MMIR) models in the scientific domain. It is worth mentioning that we define a data hierarchical architecture of "Two subsets, Five subcategories" and use human-created… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/SciMMIR.image100K<n<1M11 likes4.4k downloads3y agoHugging Face21Voxel51 /tomato-map Dataset Card for TomatoMAP This is a FiftyOne dataset with 68,069 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/tomato-map") # Launch the App session = fo.launch_app(dataset) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/tomato-map.imageobject-detection10K<n<100K4 likes4k downloads2mo agoHugging Face22shinonomelab /cleanvid-15m_map CleanVid Map (15M) 🎥 TempoFunk Video Generation Project CleanVid-15M is a large-scale dataset of videos with multiple metadata entries such as: Textual Descriptions 📃 Recording Equipment 📹 Categories 🔠 Framerate 🎞️ Aspect Ratio 📺 CleanVid aim is to improve the quality of WebVid-10M dataset by adding more data and cleaning the dataset by dewatermarking the videos in it. This dataset includes only the map with the urls and metadata, with 3,694,510 more entries than… See the full description on the dataset page: https://huggingface.co/datasets/shinonomelab/cleanvid-15m_map.tabulartext-to-video10M<n<100M23 likes4k downloads3y agoHugging Face23Farhadsakhodi /MapTrace MapTrace: A 2M-Sample Synthetic Dataset for Path Tracing on Maps Dataset Format The dataset contains 2M annotated paths designed to train models on route-tracing tasks. Splits: maptrace_parquet: Contains paths on more complex, stylized maps such as those found in brochures, park directories or shopping malls. floormap_parquet: Contains paths on simpler, structured floor maps, typical of office buildings appartment complexes, or campus maps. Each of these splits has… See the full description on the dataset page: https://huggingface.co/datasets/Farhadsakhodi/MapTrace.textimage-to-text10K<n<100K0 likes3.9k downloads7mo agoHugging Face24m-a-p /Code-Feedback OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement [🏠Homepage] | [🛠️Code] Introduction OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities. For further information and related… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Code-Feedback.textquestion-answering10K<n<100K240 likes3.7k downloads3y agoHugging Face25wdcqc /starcraft-remastered-melee-maps starcraft-remastered-melee-maps This is a dataset containing 1,815 Starcraft:Remastered melee maps, categorized into tilesets. The dataset is used to train this model: https://huggingface.co/wdcqc/starcraft-platform-terrain-32x32 The dataset is manually downloaded from Battle.net, bounding.net (scmscx.com) and broodwarmaps.com over a long period of time. To use this dataset, extract the staredit\\scenario.chk files from the map files using StormLib, then refer to Scenario.chk Format… See the full description on the dataset page: https://huggingface.co/datasets/wdcqc/starcraft-remastered-melee-maps.feature-extraction1K<n<10K2 likes3.5k downloads4y agoHugging Face26webninjasi /pk-map-statstabular1M<n<10M3 likes3.4k downloads21d agoHugging Face27m-a-p /GTZANaudio2 likes3.2k downloads1y agoHugging Face28astro-legacy-archive /wmap-single-year-maps WMAP DR5 Single-Year I/Q/U Maps The preview renders the canonical year-1 K1 TEMPERATURE field in its source NESTED order, using a Galactic Mollweide projection and a symmetric 99.5th percentile colour range. This dataset contains the complete full-resolution single-year I/Q/U release served by LAMBDA: ten WMAP differencing assemblies for each of nine observing years. These are year-specific, per-assembly measurements. They are distinct from wmap-band-maps-9yr, whose five maps… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/wmap-single-year-maps.tabular100M<n<1B0 likes2.6k downloads17d agoHugging Face29m-a-p /WildSongBench🤗 WildSongBench A benchmark for full-song music generation 192 prompts · 94 Chinese · 98 English 🎵&nbsp;YuE2&nbsp;project · 🚀&nbsp;Quick&nbsp;start · 📊&nbsp;Benchmarks · 🔁&nbsp;Reproduce · 📚&nbsp;Citation &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; WildSongBench (WSB) contains 192 song-generation prompts: 94 Chinese and 98 English, used in the YuE2 benchmarks. This repository provides prompts, exact inputs and seeds, reference scores… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/WildSongBench.documentn<1K11 likes2.4k downloads10d agoHugging Face30Maple222 /llmtcl ⚡ LitGPT 20+ high-performance LLMs with recipes to pretrain, finetune, and deploy at scale. ✅ From scratch implementations ✅ No abstractions ✅ Beginner friendly ✅ Flash attention ✅ FSDP ✅ LoRA, QLoRA, Adapter ✅ Reduce GPU memory (fp4/8/16/32) ✅ 1-1000+ GPUs/TPUs ✅ 20+ LLMs Quick start • Models • Finetune • Deploy • All workflows • Features • Recipes (YAML) • Lightning AI • Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/Maple222/llmtcl.text1 likes2.3k downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.