datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tahoe-100M
Tahoe-100M
Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from
50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics'
Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution.
This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-100M.tahoe-100m-zarr
Tahoe-100M Zarr Collection
Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.AS-100M
AS-100M
AS-100M is a subset of AS-1B. We release this dataset in both COCO format and JSONL format.
NOTE: The bbox format in the COCO format is xywh, while in the JSONL format, it is x1y1x2y2.
Introduction
We present the All-Seeing Project with:
All-Seeing 1B (AS-1B) dataset: we propose a new large-scale dataset (AS-1B) for open-world panoptic visual recognition and understanding, using an economical semi-automatic data engine that combines the power of off-the-shelf… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/AS-100M.DanQing100M
100M Chinese image-text pairs | 12TB dataset | 2024-2025 web data
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
Project Page | Paper | Code
Hengyu Shen∗, Tiancheng Gu∗, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng‡, Kaicheng Yang†
∗ Equal Contribution | ‡ Team Leader | † Project Leader
📣 News
[2026/01/16] ✨ We release the paper of DanQing.
[2026/01/15] 🔥 We release the… See the full description on the dataset page: https://huggingface.co/datasets/DeepGlint-AI/DanQing100M.Tahoe-100M
Tahoe-100M Dataset (SLAF Format)
Attribution
This is a re-release of data originally generated by Tahoe Therapeutics.
Original Dataset: tahoebio/Tahoe-100M
Original Format: Parquet files
This Release: Same data in SLAF (Sparse Lazy Array Format)
License: CC0-1.0 (Creative Commons CC0 1.0 Universal - Public Domain)
Original Citation:
@article{zhang2025tahoe,
title={Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/Tahoe-100M.pawn-stockfish-100m
PAWN Stockfish 100M
100,000,000 self-play chess games generated with Stockfish 18, each
annotated with per-position, per-legal-move evaluations — for chess
policy-learning and NNUE-distillation research.
Dataset Summary
100,000,000 machine-generated self-play chess games. Every position in
every game is annotated with an evaluation of every legal move, not just
the move played. The dataset was built as training data for
PAWN — a testbed for finetuning
and… See the full description on the dataset page: https://huggingface.co/datasets/thomas-schweich/pawn-stockfish-100m.Noah-Wukong-100M
Noah-Wukong-100M
The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset.
The dataset contains 100 Million <image, text> pairs
Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 )
Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration.
The original website: wukong-dataset.github.io
Terms of Use -… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Noah-Wukong-100M.twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
action100m-preview
Action100M: A Large-scale Video Action Dataset
Paper | GitHub
Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling.
Load Action100M Annotations
Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.foursquare_places_100M
Foursquare OS Places 100M
Full Foursquare OS Places dump from https://opensource.foursquare.com/os-places/.
This is a single (geo-)parquet file based on the 81 individual parquet files from fused.io on https://source.coop/fused/fsq-os-places/2024-11-19/places.
As it's just 10Gb, it's fairly easy to handle as a single file and can easily be queried over modern technologies like httpfs.
Ways to query the file & visualize the results
If you just want to poke around in… See the full description on the dataset page: https://huggingface.co/datasets/do-me/foursquare_places_100M.Tahoe-100M
Tahoe-100M
Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from
50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics'
Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution.
This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/velaiola/Tahoe-100M.RealSyn100M
[ACM MM25] RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
Tiancheng Gu,
Kaicheng Yang,
Chaoyi Zhang,
Yin Xie,
Xiang An,
Ziyong Feng,
Dongnan Liu,
Weidong Cai,
Jiankang Deng
💡 Introduction
Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of non-paired data, such as multimodal interleaved documents, remains… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/RealSyn100M.dclm-refinedweb-100m-sampleprecinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury
PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph,
mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data.
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.youtube_subs_howto100M
Dataset Card for youtube_subs_howto100M
Dataset Summary
The youtube_subs_howto100M dataset is an English-language dataset of instruction-response pairs extracted from 309136 YouTube videos. The dataset was orignally inspired by and sourced from the HowTo100M dataset, which was developed for natural language search for video clips.
Supported Tasks and Leaderboards
conversational: The dataset can be used to train a model for instruction(request) and a long form… See the full description on the dataset page: https://huggingface.co/datasets/totuta/youtube_subs_howto100M.precinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.redpajama_v2_sample_100M
Dataset Card for "redpajama_v2_sample_100M"
More Information needed
cyber-security-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1
harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m-kl-0p1,
revision 04390b5989d8d9adfe5597339f24a27b9b411c37, job 122416. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
17.06%
mc_logprob
1,780
59.38%
negative_abstain
1,799
27.24%
reversed_qa
1,839
6.42%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1.wukong100m
wukong100m
简介 Brief Introduction
取自Noah-Wukong多语言多模态数据集中的中文部分,一共100M个图文对。
A subset from Noah-Wukong (a multimodal dataset), around 100M image-text pairs (only Chinese).
数据集信息 Dataset Information
大约一共100M个中文图文对。大约占用16GB空间(仅仅是url等文本信息,不包含图片)。下载成功率在80%左右。(虽然我没有统计下载之后会占用多少空间,但是,可以说非常非常大)
Homepage: Noah-Wukong
下载 Download
mkdir wukong100m && cd wukong100m
for i in {00000..00031}; do wget… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wukong100m.lucky-initialization-atlas-100m-v2
Lucky initialization atlas v2 evidence
Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs,
provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses
after those artifacts complete. It excludes
credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B
binary logs.
100M_random_chess_boards
100M Random Chess Boards
This dataset contains 100 million randomly sampled chess board positions. The process to generate each board state is as follows:
Game Simulation
100 million games were played, with each move being randomly determined until the game ended (i.e., until checkmate, stalemate, or draw).
Board Sampling
From each game, a single random board state was selected at some point in the game.
Columns other than fen represent statistics for the sample. Most boards… See the full description on the dataset page: https://huggingface.co/datasets/lapp0/100M_random_chess_boards.harvey-closed-book-qwen35-9b-notes-conditioned-100m
harvey-closed-book-qwen35-9b-notes-conditioned-100m
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m,
revision 0c295885100d6eba4f514752aa081c5b0c73fdec, job 122452. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
31.33%
mc_logprob
1,780
61.85%
negative_abstain
1,799
71.10%
reversed_qa
1,839
19.85%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m.WestGenesis-Coder-SFT-100M
Dataset Overview
WestGenesis-Coder-Dataset is a meticulously curated coding dataset designed specifically for instruction-based model tuning and fine-tuning of existing models with enhanced code generation capabilities. This represents one of the largest and most comprehensively filtered corpora of publicly available coding data on the Hugging Face platform, with a non-thinking approach that emphasizes direct, concise code outputs for rapid model training.
Key… See the full description on the dataset page: https://huggingface.co/datasets/isthatshan/WestGenesis-Coder-SFT-100M.howto100m_captions_with_verb_nounshowto100m
HowTo100M
105,128 videos (11721.4 GB) with 104,584 subtitle files, downloaded at source
quality and re-hosted for direct use — no more dead YouTube links, no more flaky
downloader scripts.
Coverage: 105,128 of the 1,238,911 source video IDs (8.5%). 19,160 source videos were unavailable at fetch time (private, removed, members-only or geo-blocked) and are excluded. The dataset is refreshed as more videos are delivered.
What's inside
metadata.jsonl — one row per… See the full description on the dataset page: https://huggingface.co/datasets/TornadoLabs/howto100m.hdvila-100MMSConsensus-100M
MSConsensus
A hundred-million-scale, batch-effect-suppressed dataset and benchmark for proteomics machine learning.
MSConsensus holds 110,209,043 consensus MS/MS spectra. They were built from 1.01 PB of public raw mass-spectrometry data taken from 1,500 PRIDE repositories and covering all major Orbitrap and timsTOF platforms. A new consensus algorithm, Weighted-Score Binning (WSBIN), produces several consensus spectra per precursor instead of one representative spectrum. This… See the full description on the dataset page: https://huggingface.co/datasets/Gaolaboratory/MSConsensus-100M.SPHERE_100M
