CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tahoebio /Tahoe-100M Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics' Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution. This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-100M.tabular1B<n<10B130 likes72k downloads1y agoHugging Face02KokosDev /tahoe-100m-zarr Tahoe-100M Zarr Collection Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub. Why Zarr Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.textn<1K1 likes18k downloads6mo agoHugging Face03alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.3k downloads3mo agoHugging Face04OpenGVLab /AS-100M AS-100M AS-100M is a subset of AS-1B. We release this dataset in both COCO format and JSONL format. NOTE: The bbox format in the COCO format is xywh, while in the JSONL format, it is x1y1x2y2. Introduction We present the All-Seeing Project with: All-Seeing 1B (AS-1B) dataset: we propose a new large-scale dataset (AS-1B) for open-world panoptic visual recognition and understanding, using an economical semi-automatic data engine that combines the power of off-the-shelf… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/AS-100M.textn<1K16 likes2.3k downloads3y agoHugging Face05DeepGlint-AI /DanQing100M 100M Chinese image-text pairs | 12TB dataset | 2024-2025 web data DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset Project Page | Paper | Code Hengyu Shen∗, Tiancheng Gu∗, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng‡, Kaicheng Yang† ∗ Equal Contribution | ‡ Team Leader | † Project Leader 📣 News [2026/01/16] ✨ We release the paper of DanQing. [2026/01/15] 🔥 We release the… See the full description on the dataset page: https://huggingface.co/datasets/DeepGlint-AI/DanQing100M.imagezero-shot-image-classification10M<n<100M52 likes1.9k downloads6mo agoHugging Face06slaf-project /Tahoe-100M Tahoe-100M Dataset (SLAF Format) Attribution This is a re-release of data originally generated by Tahoe Therapeutics. Original Dataset: tahoebio/Tahoe-100M Original Format: Parquet files This Release: Same data in SLAF (Sparse Lazy Array Format) License: CC0-1.0 (Creative Commons CC0 1.0 Universal - Public Domain) Original Citation: @article{zhang2025tahoe, title={Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/Tahoe-100M.tabular100B<n<1T3 likes1.8k downloads8mo agoHugging Face07thomas-schweich /pawn-stockfish-100m PAWN Stockfish 100M 100,000,000 self-play chess games generated with Stockfish 18, each annotated with per-position, per-legal-move evaluations — for chess policy-learning and NNUE-distillation research. Dataset Summary 100,000,000 machine-generated self-play chess games. Every position in every game is annotated with an evaluation of every legal move, not just the move played. The dataset was built as training data for PAWN — a testbed for finetuning and… See the full description on the dataset page: https://huggingface.co/datasets/thomas-schweich/pawn-stockfish-100m.tabularother100M<n<1B2 likes1.8k downloads4mo agoHugging Face08Mxode /Noah-Wukong-100M Noah-Wukong-100M The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset. The dataset contains 100 Million <image, text> pairs Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 ) Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration. The original website: wukong-dataset.github.io Terms of Use -… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Noah-Wukong-100M.textimage-to-text100M<n<1B0 likes1.4k downloads1y agoHugging Face09enryu43 /twitter100m_tweets Dataset Card for "twitter100m_tweets" Dataset with tweets for this post. DOI: 10.5281/zenodo.15086029 tabular10M<n<100M35 likes1.3k downloads2y agoHugging Face10facebook /action100m-preview Action100M: A Large-scale Video Action Dataset Paper | GitHub Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling. Load Action100M Annotations Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.textvideo-classification100K<n<1M151 likes1.1k downloads8mo agoHugging Face11do-me /foursquare_places_100M Foursquare OS Places 100M Full Foursquare OS Places dump from https://opensource.foursquare.com/os-places/. This is a single (geo-)parquet file based on the 81 individual parquet files from fused.io on https://source.coop/fused/fsq-os-places/2024-11-19/places. As it's just 10Gb, it's fairly easy to handle as a single file and can easily be queried over modern technologies like httpfs. Ways to query the file & visualize the results If you just want to poke around in… See the full description on the dataset page: https://huggingface.co/datasets/do-me/foursquare_places_100M.tabularfeature-extraction100M<n<1B18 likes904 downloads2y agoHugging Face12velaiola /Tahoe-100M Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics' Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution. This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/velaiola/Tahoe-100M.tabular1B<n<10B0 likes901 downloads10mo agoHugging Face13Kaichengalex /RealSyn100M [ACM MM25] RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm Tiancheng Gu, Kaicheng Yang, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Weidong Cai, Jiankang Deng 💡 Introduction Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of non-paired data, such as multimodal interleaved documents, remains… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/RealSyn100M.image10M<n<100M16 likes898 downloads1y agoHugging Face14wytro /dclm-refinedweb-100m-sampletext100M<n<1B0 likes805 downloads5mo agoHugging Face15witfoo /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data. Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B3 likes777 downloads3d agoHugging Face16totuta /youtube_subs_howto100M Dataset Card for youtube_subs_howto100M Dataset Summary The youtube_subs_howto100M dataset is an English-language dataset of instruction-response pairs extracted from 309136 YouTube videos. The dataset was orignally inspired by and sourced from the HowTo100M dataset, which was developed for natural language search for video clips. Supported Tasks and Leaderboards conversational: The dataset can be used to train a model for instruction(request) and a long form… See the full description on the dataset page: https://huggingface.co/datasets/totuta/youtube_subs_howto100M.text100K<n<1M4 likes573 downloads4y agoHugging Face17artham123 /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B0 likes480 downloads4mo agoHugging Face18sade-adrien /redpajama_v2_sample_100M Dataset Card for "redpajama_v2_sample_100M" More Information needed text100M<n<1B0 likes372 downloads3y agoHugging Face19Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes312 downloads2mo agoHugging Face20violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1 harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1 Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m-kl-0p1, revision 04390b5989d8d9adfe5597339f24a27b9b411c37, job 122416. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 17.06% mc_logprob 1,780 59.38% negative_abstain 1,799 27.24% reversed_qa 1,839 6.42%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1.textquestion-answering1K<n<10K0 likes312 downloads2d agoHugging Face21wanng /wukong100m wukong100m 简介 Brief Introduction 取自Noah-Wukong多语言多模态数据集中的中文部分,一共100M个图文对。 A subset from Noah-Wukong (a multimodal dataset), around 100M image-text pairs (only Chinese). 数据集信息 Dataset Information 大约一共100M个中文图文对。大约占用16GB空间(仅仅是url等文本信息,不包含图片)。下载成功率在80%左右。(虽然我没有统计下载之后会占用多少空间,但是,可以说非常非常大) Homepage: Noah-Wukong 下载 Download mkdir wukong100m && cd wukong100m for i in {00000..00031}; do wget… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wukong100m.textfeature-extraction10M<n<100M17 likes311 downloads4y agoHugging Face22lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes296 downloads28d agoHugging Face23lapp0 /100M_random_chess_boards 100M Random Chess Boards This dataset contains 100 million randomly sampled chess board positions. The process to generate each board state is as follows: Game Simulation 100 million games were played, with each move being randomly determined until the game ended (i.e., until checkmate, stalemate, or draw). Board Sampling From each game, a single random board state was selected at some point in the game. Columns other than fen represent statistics for the sample. Most boards… See the full description on the dataset page: https://huggingface.co/datasets/lapp0/100M_random_chess_boards.text100M<n<1B0 likes281 downloads2y agoHugging Face24violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-100m harvey-closed-book-qwen35-9b-notes-conditioned-100m Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m, revision 0c295885100d6eba4f514752aa081c5b0c73fdec, job 122452. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 31.33% mc_logprob 1,780 61.85% negative_abstain 1,799 71.10% reversed_qa 1,839 19.85%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m.textquestion-answering1K<n<10K0 likes281 downloads2d agoHugging Face25isthatshan /WestGenesis-Coder-SFT-100M Dataset Overview WestGenesis-Coder-Dataset is a meticulously curated coding dataset designed specifically for instruction-based model tuning and fine-tuning of existing models with enhanced code generation capabilities. This represents one of the largest and most comprehensively filtered corpora of publicly available coding data on the Hugging Face platform, with a non-thinking approach that emphasizes direct, concise code outputs for rapid model training. Key… See the full description on the dataset page: https://huggingface.co/datasets/isthatshan/WestGenesis-Coder-SFT-100M.texttext-generation1M<n<10M0 likes280 downloads3mo agoHugging Face26bpiyush /howto100m_captions_with_verb_nounstabular10M<n<100M1 likes247 downloads2y agoHugging Face27TornadoLabs /howto100m HowTo100M 105,128 videos (11721.4 GB) with 104,584 subtitle files, downloaded at source quality and re-hosted for direct use — no more dead YouTube links, no more flaky downloader scripts. Coverage: 105,128 of the 1,238,911 source video IDs (8.5%). 19,160 source videos were unavailable at fetch time (private, removed, members-only or geo-blocked) and are excluded. The dataset is refreshed as more videos are delivered. What's inside metadata.jsonl — one row per… See the full description on the dataset page: https://huggingface.co/datasets/TornadoLabs/howto100m.textvideo-classification100K<n<1M0 likes205 downloads5d agoHugging Face28TempoFunk /hdvila-100Mtexttext-to-video10M<n<100M17 likes184 downloads3y agoHugging Face29Gaolaboratory /MSConsensus-100M MSConsensus A hundred-million-scale, batch-effect-suppressed dataset and benchmark for proteomics machine learning. MSConsensus holds 110,209,043 consensus MS/MS spectra. They were built from 1.01 PB of public raw mass-spectrometry data taken from 1,500 PRIDE repositories and covering all major Orbitrap and timsTOF platforms. A new consensus algorithm, Weighted-Score Binning (WSBIN), produces several consensus spectra per precursor instead of one representative spectrum. This… See the full description on the dataset page: https://huggingface.co/datasets/Gaolaboratory/MSConsensus-100M.textfeature-extraction100M<n<1B0 likes178 downloads2d agoHugging Face30mohdumar /SPHERE_100Mtext100M<n<1B0 likes147 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.