CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tahoebio /Tahoe-100M Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics' Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution. This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-100M.tabular1B<n<10B130 likes72k downloads1y agoHugging Face02KokosDev /tahoe-100m-zarr Tahoe-100M Zarr Collection Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub. Why Zarr Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.textn<1K1 likes18k downloads6mo agoHugging Face03dalle-mini /YFCC100M_OpenAI_subsetThe YFCC100M is one of the largest publicly and freely useable multimedia collection, containing the metadata of around 99.2 million photos and 0.8 million videos from Flickr, all of which were shared under one of the various Creative Commons licenses. This version is a subset defined in openai/CLIP.28 likes11k downloads5y agoHugging Face04VDBBench /multimodal-embedding-100M Multimodal Embedding 100M This dataset contains a 100M-row multimodal embedding corpus generated from LAION-style image-text data exported with img2dataset as WebDataset shards. Images were resized to 256 during the WebDataset creation step before embedding generation. The dataset is intended for large-scale vector database ingestion, ANN index construction, nearest-neighbor search, and retrieval benchmark experiments. The dataset is stored as Parquet files and organized to keep… See the full description on the dataset page: https://huggingface.co/datasets/VDBBench/multimodal-embedding-100M.feature-extraction100M<n<1B1 likes3.7k downloads3mo agoHugging Face05alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.3k downloads3mo agoHugging Face06OpenGVLab /AS-100M AS-100M AS-100M is a subset of AS-1B. We release this dataset in both COCO format and JSONL format. NOTE: The bbox format in the COCO format is xywh, while in the JSONL format, it is x1y1x2y2. Introduction We present the All-Seeing Project with: All-Seeing 1B (AS-1B) dataset: we propose a new large-scale dataset (AS-1B) for open-world panoptic visual recognition and understanding, using an economical semi-automatic data engine that combines the power of off-the-shelf… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/AS-100M.textn<1K16 likes2.3k downloads3y agoHugging Face07DeepGlint-AI /DanQing100M 100M Chinese image-text pairs | 12TB dataset | 2024-2025 web data DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset Project Page | Paper | Code Hengyu Shen∗, Tiancheng Gu∗, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng‡, Kaicheng Yang† ∗ Equal Contribution | ‡ Team Leader | † Project Leader 📣 News [2026/01/16] ✨ We release the paper of DanQing. [2026/01/15] 🔥 We release the… See the full description on the dataset page: https://huggingface.co/datasets/DeepGlint-AI/DanQing100M.imagezero-shot-image-classification10M<n<100M52 likes1.9k downloads6mo agoHugging Face08slaf-project /Tahoe-100M Tahoe-100M Dataset (SLAF Format) Attribution This is a re-release of data originally generated by Tahoe Therapeutics. Original Dataset: tahoebio/Tahoe-100M Original Format: Parquet files This Release: Same data in SLAF (Sparse Lazy Array Format) License: CC0-1.0 (Creative Commons CC0 1.0 Universal - Public Domain) Original Citation: @article{zhang2025tahoe, title={Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/Tahoe-100M.tabular100B<n<1T3 likes1.8k downloads8mo agoHugging Face09thomas-schweich /pawn-stockfish-100m PAWN Stockfish 100M 100,000,000 self-play chess games generated with Stockfish 18, each annotated with per-position, per-legal-move evaluations — for chess policy-learning and NNUE-distillation research. Dataset Summary 100,000,000 machine-generated self-play chess games. Every position in every game is annotated with an evaluation of every legal move, not just the move played. The dataset was built as training data for PAWN — a testbed for finetuning and… See the full description on the dataset page: https://huggingface.co/datasets/thomas-schweich/pawn-stockfish-100m.tabularother100M<n<1B2 likes1.8k downloads4mo agoHugging Face10Voxel51 /action100m_tiny_subset Dataset Card for action100m This is a FiftyOne dataset with 1144 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/action100m_tiny_subset") # Launch the App session = fo.launch_app(dataset) Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/action100m_tiny_subset.video1K<n<10K3 likes1.4k downloads8mo agoHugging Face11Mxode /Noah-Wukong-100M Noah-Wukong-100M The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset. The dataset contains 100 Million <image, text> pairs Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 ) Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration. The original website: wukong-dataset.github.io Terms of Use -… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Noah-Wukong-100M.textimage-to-text100M<n<1B0 likes1.4k downloads1y agoHugging Face12enryu43 /twitter100m_tweets Dataset Card for "twitter100m_tweets" Dataset with tweets for this post. DOI: 10.5281/zenodo.15086029 tabular10M<n<100M35 likes1.3k downloads2y agoHugging Face13zkchen /papers100M0 likes1.3k downloads3y agoHugging Face14FedGraph /fedgraph_ogbn-papers100M_195trainer_0hop_iid_beta_10000.0_v10 likes1.2k downloads22d agoHugging Face15facebook /action100m-preview Action100M: A Large-scale Video Action Dataset Paper | GitHub Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling. Load Action100M Annotations Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.textvideo-classification100K<n<1M151 likes1.1k downloads8mo agoHugging Face16do-me /foursquare_places_100M Foursquare OS Places 100M Full Foursquare OS Places dump from https://opensource.foursquare.com/os-places/. This is a single (geo-)parquet file based on the 81 individual parquet files from fused.io on https://source.coop/fused/fsq-os-places/2024-11-19/places. As it's just 10Gb, it's fairly easy to handle as a single file and can easily be queried over modern technologies like httpfs. Ways to query the file & visualize the results If you just want to poke around in… See the full description on the dataset page: https://huggingface.co/datasets/do-me/foursquare_places_100M.tabularfeature-extraction100M<n<1B18 likes904 downloads2y agoHugging Face17velaiola /Tahoe-100M Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics' Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution. This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/velaiola/Tahoe-100M.tabular1B<n<10B0 likes901 downloads10mo agoHugging Face18Kaichengalex /RealSyn100M [ACM MM25] RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm Tiancheng Gu, Kaicheng Yang, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Weidong Cai, Jiankang Deng 💡 Introduction Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of non-paired data, such as multimodal interleaved documents, remains… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/RealSyn100M.image10M<n<100M16 likes898 downloads1y agoHugging Face19wytro /dclm-refinedweb-100m-sampletext100M<n<1B0 likes805 downloads5mo agoHugging Face20witfoo /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data. Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B3 likes777 downloads3d agoHugging Face21maknee /sift100m0 likes691 downloads8mo agoHugging Face22kaptaan45 /KapInstruct-100M KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.question-answering10K<n<100K0 likes631 downloads1mo agoHugging Face23mignonjia /acav100m ACAV100M Video+Caption Dataset Dataset Structure The dataset is split into sharded .tar.gz archives (~1000 video+caption pairs each). Each shard has the following structure: shard_XXXX.tar.gz └── shard_XXXX/ ├── videos/ │ ├── <video_id>_clip.mp4 │ └── ... └── captions/ ├── <video_id>_clip.txt └── ... videos/: 5-second 1080p MP4 clips with audio captions/: Corresponding text caption for each video clip Only videos with a matching… See the full description on the dataset page: https://huggingface.co/datasets/mignonjia/acav100m.0 likes577 downloads7mo agoHugging Face24totuta /youtube_subs_howto100M Dataset Card for youtube_subs_howto100M Dataset Summary The youtube_subs_howto100M dataset is an English-language dataset of instruction-response pairs extracted from 309136 YouTube videos. The dataset was orignally inspired by and sourced from the HowTo100M dataset, which was developed for natural language search for video clips. Supported Tasks and Leaderboards conversational: The dataset can be used to train a model for instruction(request) and a long form… See the full description on the dataset page: https://huggingface.co/datasets/totuta/youtube_subs_howto100M.text100K<n<1M4 likes573 downloads4y agoHugging Face25cerebros /NotGPT-mythos-base-en-1B-tokens-for-100M-model100K<n<1M1 likes525 downloads5mo agoHugging Face26artham123 /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B0 likes480 downloads4mo agoHugging Face27sade-adrien /redpajama_v2_sample_100M Dataset Card for "redpajama_v2_sample_100M" More Information needed text100M<n<1B0 likes372 downloads3y agoHugging Face28Odeinjul /bigann-100m-static-search-eval bigann-100m-static-search-eval Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts. Files Base: base.u8bin HNSW index: index_m_32_ef_500 Query: orig_query_10k.u8bin Ground truth: groundtruth.bin Checksums: checksums.sha256 Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/bigann-100m-static-search-eval.0 likes372 downloads3d agoHugging Face29Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes312 downloads2mo agoHugging Face30Odeinjul /deep-100m-static-search-eval deep-100m-static-search-eval Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts. Files Base: base.fbin HNSW index: index_m_32_ef_500 Query: orig_query_10k.fbin Ground truth: groundtruth.bin Additional held-out 500K query/GT set: queries/heldout_deep1b_random500k_seed20260922/ Checksums: checksums.sha256 Kaggle metadata is provided via dataset-metadata.json during… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-100m-static-search-eval.0 likes312 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.