datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tahoe-100M
Tahoe-100M
Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from
50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics'
Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution.
This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-100M.tahoe-100m-zarr
Tahoe-100M Zarr Collection
Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.YFCC100M_OpenAI_subsetThe YFCC100M is one of the largest publicly and freely useable multimedia collection, containing the metadata of around 99.2 million photos and 0.8 million videos from Flickr, all of which were shared under one of the various Creative Commons licenses.
This version is a subset defined in openai/CLIP.multimodal-embedding-100M
Multimodal Embedding 100M
This dataset contains a 100M-row multimodal embedding corpus generated from LAION-style image-text data exported with img2dataset as WebDataset shards. Images were resized to 256 during the WebDataset creation step before embedding generation. The dataset is intended for large-scale vector database ingestion, ANN index construction, nearest-neighbor search, and retrieval benchmark experiments.
The dataset is stored as Parquet files and organized to keep… See the full description on the dataset page: https://huggingface.co/datasets/VDBBench/multimodal-embedding-100M.fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.AS-100M
AS-100M
AS-100M is a subset of AS-1B. We release this dataset in both COCO format and JSONL format.
NOTE: The bbox format in the COCO format is xywh, while in the JSONL format, it is x1y1x2y2.
Introduction
We present the All-Seeing Project with:
All-Seeing 1B (AS-1B) dataset: we propose a new large-scale dataset (AS-1B) for open-world panoptic visual recognition and understanding, using an economical semi-automatic data engine that combines the power of off-the-shelf… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/AS-100M.DanQing100M
100M Chinese image-text pairs | 12TB dataset | 2024-2025 web data
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
Project Page | Paper | Code
Hengyu Shen∗, Tiancheng Gu∗, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng‡, Kaicheng Yang†
∗ Equal Contribution | ‡ Team Leader | † Project Leader
📣 News
[2026/01/16] ✨ We release the paper of DanQing.
[2026/01/15] 🔥 We release the… See the full description on the dataset page: https://huggingface.co/datasets/DeepGlint-AI/DanQing100M.Tahoe-100M
Tahoe-100M Dataset (SLAF Format)
Attribution
This is a re-release of data originally generated by Tahoe Therapeutics.
Original Dataset: tahoebio/Tahoe-100M
Original Format: Parquet files
This Release: Same data in SLAF (Sparse Lazy Array Format)
License: CC0-1.0 (Creative Commons CC0 1.0 Universal - Public Domain)
Original Citation:
@article{zhang2025tahoe,
title={Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/Tahoe-100M.pawn-stockfish-100m
PAWN Stockfish 100M
100,000,000 self-play chess games generated with Stockfish 18, each
annotated with per-position, per-legal-move evaluations — for chess
policy-learning and NNUE-distillation research.
Dataset Summary
100,000,000 machine-generated self-play chess games. Every position in
every game is annotated with an evaluation of every legal move, not just
the move played. The dataset was built as training data for
PAWN — a testbed for finetuning
and… See the full description on the dataset page: https://huggingface.co/datasets/thomas-schweich/pawn-stockfish-100m.action100m_tiny_subset
Dataset Card for action100m
This is a FiftyOne dataset with 1144 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/action100m_tiny_subset")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/action100m_tiny_subset.Noah-Wukong-100M
Noah-Wukong-100M
The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset.
The dataset contains 100 Million <image, text> pairs
Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 )
Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration.
The original website: wukong-dataset.github.io
Terms of Use -… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Noah-Wukong-100M.twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
papers100Mfedgraph_ogbn-papers100M_195trainer_0hop_iid_beta_10000.0_v1action100m-preview
Action100M: A Large-scale Video Action Dataset
Paper | GitHub
Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling.
Load Action100M Annotations
Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.foursquare_places_100M
Foursquare OS Places 100M
Full Foursquare OS Places dump from https://opensource.foursquare.com/os-places/.
This is a single (geo-)parquet file based on the 81 individual parquet files from fused.io on https://source.coop/fused/fsq-os-places/2024-11-19/places.
As it's just 10Gb, it's fairly easy to handle as a single file and can easily be queried over modern technologies like httpfs.
Ways to query the file & visualize the results
If you just want to poke around in… See the full description on the dataset page: https://huggingface.co/datasets/do-me/foursquare_places_100M.Tahoe-100M
Tahoe-100M
Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from
50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics'
Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution.
This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/velaiola/Tahoe-100M.RealSyn100M
[ACM MM25] RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
Tiancheng Gu,
Kaicheng Yang,
Chaoyi Zhang,
Yin Xie,
Xiang An,
Ziyong Feng,
Dongnan Liu,
Weidong Cai,
Jiankang Deng
💡 Introduction
Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of non-paired data, such as multimodal interleaved documents, remains… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/RealSyn100M.dclm-refinedweb-100m-sampleprecinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury
PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph,
mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data.
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.sift100mKapInstruct-100M
KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset
KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.acav100m
ACAV100M Video+Caption Dataset
Dataset Structure
The dataset is split into sharded .tar.gz archives (~1000 video+caption pairs each).
Each shard has the following structure:
shard_XXXX.tar.gz
└── shard_XXXX/
├── videos/
│ ├── <video_id>_clip.mp4
│ └── ...
└── captions/
├── <video_id>_clip.txt
└── ...
videos/: 5-second 1080p MP4 clips with audio
captions/: Corresponding text caption for each video clip
Only videos with a matching… See the full description on the dataset page: https://huggingface.co/datasets/mignonjia/acav100m.youtube_subs_howto100M
Dataset Card for youtube_subs_howto100M
Dataset Summary
The youtube_subs_howto100M dataset is an English-language dataset of instruction-response pairs extracted from 309136 YouTube videos. The dataset was orignally inspired by and sourced from the HowTo100M dataset, which was developed for natural language search for video clips.
Supported Tasks and Leaderboards
conversational: The dataset can be used to train a model for instruction(request) and a long form… See the full description on the dataset page: https://huggingface.co/datasets/totuta/youtube_subs_howto100M.NotGPT-mythos-base-en-1B-tokens-for-100M-modelprecinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.redpajama_v2_sample_100M
Dataset Card for "redpajama_v2_sample_100M"
More Information needed
bigann-100m-static-search-eval
bigann-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.u8bin
HNSW index: index_m_32_ef_500
Query: orig_query_10k.u8bin
Ground truth: groundtruth.bin
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/bigann-100m-static-search-eval.cyber-security-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.deep-100m-static-search-eval
deep-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.fbin
HNSW index: index_m_32_ef_500
Query: orig_query_10k.fbin
Ground truth: groundtruth.bin
Additional held-out 500K query/GT set: queries/heldout_deep1b_random500k_seed20260922/
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-100m-static-search-eval.
