datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tahoe-100M
Tahoe-100M
Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from
50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics'
Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution.
This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-100M.fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.Tahoe-100M
Tahoe-100M Dataset (SLAF Format)
Attribution
This is a re-release of data originally generated by Tahoe Therapeutics.
Original Dataset: tahoebio/Tahoe-100M
Original Format: Parquet files
This Release: Same data in SLAF (Sparse Lazy Array Format)
License: CC0-1.0 (Creative Commons CC0 1.0 Universal - Public Domain)
Original Citation:
@article{zhang2025tahoe,
title={Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/Tahoe-100M.pawn-stockfish-100m
PAWN Stockfish 100M
100,000,000 self-play chess games generated with Stockfish 18, each
annotated with per-position, per-legal-move evaluations — for chess
policy-learning and NNUE-distillation research.
Dataset Summary
100,000,000 machine-generated self-play chess games. Every position in
every game is annotated with an evaluation of every legal move, not just
the move played. The dataset was built as training data for
PAWN — a testbed for finetuning
and… See the full description on the dataset page: https://huggingface.co/datasets/thomas-schweich/pawn-stockfish-100m.twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
foursquare_places_100M
Foursquare OS Places 100M
Full Foursquare OS Places dump from https://opensource.foursquare.com/os-places/.
This is a single (geo-)parquet file based on the 81 individual parquet files from fused.io on https://source.coop/fused/fsq-os-places/2024-11-19/places.
As it's just 10Gb, it's fairly easy to handle as a single file and can easily be queried over modern technologies like httpfs.
Ways to query the file & visualize the results
If you just want to poke around in… See the full description on the dataset page: https://huggingface.co/datasets/do-me/foursquare_places_100M.Tahoe-100M
Tahoe-100M
Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from
50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics'
Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution.
This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/velaiola/Tahoe-100M.precinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury
PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph,
mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data.
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.precinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.cyber-security-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.lucky-initialization-atlas-100m-v2
Lucky initialization atlas v2 evidence
Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs,
provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses
after those artifacts complete. It excludes
credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B
binary logs.
howto100m_captions_with_verb_nounsthermoshift-100m
ThermoShift
Hourly synthetic building-cooling trajectories with observed decisions and paired
counterfactual outcomes.
The dataset contains three cooling actions, their logging probabilities, factual
outcomes, and an oracle table with latent state and potential outcomes for every
action. Buildings are assigned to train, validation, and test splits, including
heatwave and sensor-degradation conditions.
Configurations
Name
Contents
logged
Decision-time… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/thermoshift-100m.Rainbow-Pony-100m-Flutter-steps-eval
Rainbow-Pony-100M Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In steps mode, the model is given an existing file and an edit instruction and
generates a sequence of localized search/replace edit actions, each mechanically
applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.chess-soft-100m-swa-mistakes
avewright/chess-soft-100m-swa-mistakes
Positions where overnight eval_swa.pt
(outputs/sf19_ft/overnight_20260908, SF 8000-node screen score 0.719)
disagrees with a teacher best move.
In-PV severity uses teacher MultiPV STM cps. Off-PV rows are not given
an invented drop; they are queued for Stockfish 19 analysis (needs_sf=1).
Holdout + flip hashes from the overnight union (20,961) are excluded.
1,264,747 rows in 64 shards.
Tags
tag
rows
blunder
59,889… See the full description on the dataset page: https://huggingface.co/datasets/avewright/chess-soft-100m-swa-mistakes.chess-soft-100m-disagreements
avewright/chess-soft-100m-disagreements
Positions where the greedy policy of
avewright/chess-transformer-100m-squares64
disagrees with a strong teacher best move (move_idx).
Current upload: 1,782,505 rows in 90 shards.
Mix
split
rows
shards
labels
data/shard_*.parquet
1,768,622
83
teacher MultiPV from avewright/chess-soft-multipv-lichess
data/sf19/*.parquet
13,883
7
Stockfish 19 max-Elo MultiPV from new games
Stream rows are an argmax filter of… See the full description on the dataset page: https://huggingface.co/datasets/avewright/chess-soft-100m-disagreements.Ornith-1.5-35B-A3B-Nemotron-v2-100M
Ornith 1.5 35B A3B Nemotron v2 100M
This dataset contains 108,729 English conversations with
108,729 regenerated assistant turns and 100,014,884 generated
assistant completion tokens. 100M refers to the completion-token target, not the
number of examples.
The prompt mix is a deterministic sample from
nvidia/Nemotron-Post-Training-Dataset-v2.
It covers the source dataset's chat, code, math, and STEM subsets. Every assistant turn
was regenerated with ornith-ai/Ornith-1.5-35B-A3B;… See the full description on the dataset page: https://huggingface.co/datasets/jzinno/Ornith-1.5-35B-A3B-Nemotron-v2-100M.cyber-security-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/cyber-security-100m.precinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (114M)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114 million sanitized security events (signal logs) and provenance graphs (10,442 incident graphs with 23,362 nodes and 32,732,650 edges) from real enterprise network monitoring across 5 organizations.
Available in two sizes:… See the full description on the dataset page: https://huggingface.co/datasets/bijaye/precinct6-cybersecurity-100m.laion_text_debiased_100M
100M Text Debiased Subset from LAION 2B
Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images.
Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data.
CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.synthlog-100mRainbow-Pony-100m-Flutter-direct-eval
Rainbow-Pony-100M Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In direct mode, the model is given an existing file and an edit instruction and
generates the complete modified file in a single forward pass (as opposed to the
steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.E-MM1-100M
Dataset Card for E-MM1-100M
Dataset Summary
E-MM1-100M is a large-scale multimodal dataset of 100M+ data groups, pairing data from five modalities: audio, image, video, point cloud, and text.
Each pair is a 5-tuple of a caption and an item from one of the four other modalities.
The data and captions are sourced from public data sources.
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/encord-team/E-MM1-100M.bulgarian-medical-cpt-100m
Bulgarian text for MOSS continued pretraining
Exactly 100 million training tokens: 10M medical and 90M general Bulgarian.
An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers.
No model training has been performed as part of this dataset build.
Medical data
Exactly 10,000,000 training tokens and 50,000 additional validation tokens,
including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.FreeRealEstate100M
Free Synthetic Real Estate Listings — 100M Rows
A free, fully synthetic dataset of 100,000,000 US real estate listings, generated for developers and builders working on price-prediction models, real-estate analytics, search/filter UIs, and BI pipelines — realistic property data without touching any real listing, address, or owner.
Every value in this dataset is artificially generated. No real properties, no scraped listings, no real addresses or PII. What makes it useful:… See the full description on the dataset page: https://huggingface.co/datasets/ziadatalabs/FreeRealEstate100M.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.HowTo100M-subtitles-small
HowTo100M-subtitles-small
The subtitles from a subset of the HowTo100M dataset.
