100M
Datasets
All datasets matching “100M”Tahoe-100M
Tahoe-100M
Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from
50 cancer cell lines exposed to 1,100 small-molecule perturbations. Generated using Vevo Therapeutics'
Mosaic high-throughput platform, Tahoe-100M enables deep, context-aware exploration of gene function, cellular states, and drug responses at unprecedented scale and resolution.
This dataset is designed to power the development of next-generation AI… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-100M.tahoe-100m-zarr
Tahoe-100M Zarr Collection
Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.YFCC100M_OpenAI_subsetThe YFCC100M is one of the largest publicly and freely useable multimedia collection, containing the metadata of around 99.2 million photos and 0.8 million videos from Flickr, all of which were shared under one of the various Creative Commons licenses.
This version is a subset defined in openai/CLIP.multimodal-embedding-100M
Multimodal Embedding 100M
This dataset contains a 100M-row multimodal embedding corpus generated from LAION-style image-text data exported with img2dataset as WebDataset shards. Images were resized to 256 during the WebDataset creation step before embedding generation. The dataset is intended for large-scale vector database ingestion, ANN index construction, nearest-neighbor search, and retrieval benchmark experiments.
The dataset is stored as Parquet files and organized to keep… See the full description on the dataset page: https://huggingface.co/datasets/VDBBench/multimodal-embedding-100M.AS-100M
AS-100M
AS-100M is a subset of AS-1B. We release this dataset in both COCO format and JSONL format.
NOTE: The bbox format in the COCO format is xywh, while in the JSONL format, it is x1y1x2y2.
Introduction
We present the All-Seeing Project with:
All-Seeing 1B (AS-1B) dataset: we propose a new large-scale dataset (AS-1B) for open-world panoptic visual recognition and understanding, using an economical semi-automatic data engine that combines the power of off-the-shelf… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/AS-100M.fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.
