CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tahoebio /Tahoe-100M Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics' Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution. This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-100M.tabular1B<n<10B130 likes72k downloads1y agoHugging Face02alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.3k downloads3mo agoHugging Face03slaf-project /Tahoe-100M Tahoe-100M Dataset (SLAF Format) Attribution This is a re-release of data originally generated by Tahoe Therapeutics. Original Dataset: tahoebio/Tahoe-100M Original Format: Parquet files This Release: Same data in SLAF (Sparse Lazy Array Format) License: CC0-1.0 (Creative Commons CC0 1.0 Universal - Public Domain) Original Citation: @article{zhang2025tahoe, title={Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/Tahoe-100M.tabular100B<n<1T3 likes1.8k downloads8mo agoHugging Face04thomas-schweich /pawn-stockfish-100m PAWN Stockfish 100M 100,000,000 self-play chess games generated with Stockfish 18, each annotated with per-position, per-legal-move evaluations — for chess policy-learning and NNUE-distillation research. Dataset Summary 100,000,000 machine-generated self-play chess games. Every position in every game is annotated with an evaluation of every legal move, not just the move played. The dataset was built as training data for PAWN — a testbed for finetuning and… See the full description on the dataset page: https://huggingface.co/datasets/thomas-schweich/pawn-stockfish-100m.tabularother100M<n<1B2 likes1.8k downloads4mo agoHugging Face05enryu43 /twitter100m_tweets Dataset Card for "twitter100m_tweets" Dataset with tweets for this post. DOI: 10.5281/zenodo.15086029 tabular10M<n<100M35 likes1.3k downloads2y agoHugging Face06do-me /foursquare_places_100M Foursquare OS Places 100M Full Foursquare OS Places dump from https://opensource.foursquare.com/os-places/. This is a single (geo-)parquet file based on the 81 individual parquet files from fused.io on https://source.coop/fused/fsq-os-places/2024-11-19/places. As it's just 10Gb, it's fairly easy to handle as a single file and can easily be queried over modern technologies like httpfs. Ways to query the file & visualize the results If you just want to poke around in… See the full description on the dataset page: https://huggingface.co/datasets/do-me/foursquare_places_100M.tabularfeature-extraction100M<n<1B18 likes904 downloads2y agoHugging Face07velaiola /Tahoe-100M Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics' Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution. This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/velaiola/Tahoe-100M.tabular1B<n<10B0 likes901 downloads10mo agoHugging Face08witfoo /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data. Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B3 likes777 downloads3d agoHugging Face09artham123 /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B0 likes480 downloads4mo agoHugging Face10Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes312 downloads2mo agoHugging Face11lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes296 downloads28d agoHugging Face12bpiyush /howto100m_captions_with_verb_nounstabular10M<n<100M1 likes247 downloads2y agoHugging Face13neuralsorcerer /thermoshift-100m ThermoShift Hourly synthetic building-cooling trajectories with observed decisions and paired counterfactual outcomes. The dataset contains three cooling actions, their logging probabilities, factual outcomes, and an oracle table with latent state and potential outcomes for every action. Buildings are assigned to train, validation, and test splits, including heatwave and sensor-degradation conditions. Configurations Name Contents logged Decision-time… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/thermoshift-100m.tabulartabular-regression100M<n<1B0 likes154 downloads13d agoHugging Face14bbidpa /Rainbow-Pony-100m-Flutter-steps-eval Rainbow-Pony-100M Flutter — Steps Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In steps mode, the model is given an existing file and an edit instruction and generates a sequence of localized search/replace edit actions, each mechanically applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.tabulartext-generation1K<n<10K1 likes144 downloads17d agoHugging Face15avewright /chess-soft-100m-swa-mistakes avewright/chess-soft-100m-swa-mistakes Positions where overnight eval_swa.pt (outputs/sf19_ft/overnight_20260908, SF 8000-node screen score 0.719) disagrees with a teacher best move. In-PV severity uses teacher MultiPV STM cps. Off-PV rows are not given an invented drop; they are queued for Stockfish 19 analysis (needs_sf=1). Holdout + flip hashes from the overnight union (20,961) are excluded. 1,264,747 rows in 64 shards. Tags tag rows blunder 59,889… See the full description on the dataset page: https://huggingface.co/datasets/avewright/chess-soft-100m-swa-mistakes.tabularother1M<n<10M0 likes144 downloads17d agoHugging Face16avewright /chess-soft-100m-disagreements avewright/chess-soft-100m-disagreements Positions where the greedy policy of avewright/chess-transformer-100m-squares64 disagrees with a strong teacher best move (move_idx). Current upload: 1,782,505 rows in 90 shards. Mix split rows shards labels data/shard_*.parquet 1,768,622 83 teacher MultiPV from avewright/chess-soft-multipv-lichess data/sf19/*.parquet 13,883 7 Stockfish 19 max-Elo MultiPV from new games Stream rows are an argmax filter of… See the full description on the dataset page: https://huggingface.co/datasets/avewright/chess-soft-100m-disagreements.tabularother1M<n<10M0 likes140 downloads19d agoHugging Face17jzinno /Ornith-1.5-35B-A3B-Nemotron-v2-100M Ornith 1.5 35B A3B Nemotron v2 100M This dataset contains 108,729 English conversations with 108,729 regenerated assistant turns and 100,014,884 generated assistant completion tokens. 100M refers to the completion-token target, not the number of examples. The prompt mix is a deterministic sample from nvidia/Nemotron-Post-Training-Dataset-v2. It covers the source dataset's chat, code, math, and STEM subsets. Every assistant turn was regenerated with ornith-ai/Ornith-1.5-35B-A3B;… See the full description on the dataset page: https://huggingface.co/datasets/jzinno/Ornith-1.5-35B-A3B-Nemotron-v2-100M.tabular100K<n<1M0 likes134 downloads1mo agoHugging Face18Nobody05 /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/cyber-security-100m.tabulartext-classification100M<n<1B0 likes122 downloads2mo agoHugging Face19bijaye /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (114M) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114 million sanitized security events (signal logs) and provenance graphs (10,442 incident graphs with 23,362 nodes and 32,732,650 edges) from real enterprise network monitoring across 5 organizations. Available in two sizes:… See the full description on the dataset page: https://huggingface.co/datasets/bijaye/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B0 likes113 downloads5mo agoHugging Face20linyq /laion_text_debiased_100M 100M Text Debiased Subset from LAION 2B Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images. Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data. CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.image100M<n<1B0 likes112 downloads2y agoHugging Face21dkstr /synthlog-100mtabular100M<n<1B0 likes106 downloads14d agoHugging Face22bbidpa /Rainbow-Pony-100m-Flutter-direct-eval Rainbow-Pony-100M Flutter — Direct Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In direct mode, the model is given an existing file and an edit instruction and generates the complete modified file in a single forward pass (as opposed to the steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.tabulartext-generation1K<n<10K0 likes102 downloads17d agoHugging Face23encord-team /E-MM1-100M Dataset Card for E-MM1-100M Dataset Summary E-MM1-100M is a large-scale multimodal dataset of 100M+ data groups, pairing data from five modalities: audio, image, video, point cloud, and text. Each pair is a 5-tuple of a caption and an item from one of the four other modalities. The data and captions are sourced from public data sources. The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/encord-team/E-MM1-100M.tabular100M<n<1B8 likes94 downloads10mo agoHugging Face24DimitarV /bulgarian-medical-cpt-100m Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.tabulartext-generation100K<n<1M0 likes88 downloads14d agoHugging Face25violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.tabulartext-generation1K<n<10K0 likes86 downloads2d agoHugging Face26violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.tabulartext-generation1K<n<10K0 likes85 downloads2d agoHugging Face27violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.tabulartext-generation1K<n<10K0 likes81 downloads2d agoHugging Face28ziadatalabs /FreeRealEstate100M Free Synthetic Real Estate Listings — 100M Rows A free, fully synthetic dataset of 100,000,000 US real estate listings, generated for developers and builders working on price-prediction models, real-estate analytics, search/filter UIs, and BI pipelines — realistic property data without touching any real listing, address, or owner. Every value in this dataset is artificially generated. No real properties, no scraped listings, no real addresses or PII. What makes it useful:… See the full description on the dataset page: https://huggingface.co/datasets/ziadatalabs/FreeRealEstate100M.tabulartabular-regression100M<n<1B3 likes79 downloads1mo agoHugging Face29violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.tabulartext-generation1K<n<10K0 likes78 downloads2d agoHugging Face30diyarhamedi /HowTo100M-subtitles-small HowTo100M-subtitles-small The subtitles from a subset of the HowTo100M dataset. tabular10K<n<100K2 likes74 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.