CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KokosDev /tahoe-100m-zarr Tahoe-100M Zarr Collection Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub. Why Zarr Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.textn<1K1 likes18k downloads6mo agoHugging Face02alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.3k downloads3mo agoHugging Face03OpenGVLab /AS-100M AS-100M AS-100M is a subset of AS-1B. We release this dataset in both COCO format and JSONL format. NOTE: The bbox format in the COCO format is xywh, while in the JSONL format, it is x1y1x2y2. Introduction We present the All-Seeing Project with: All-Seeing 1B (AS-1B) dataset: we propose a new large-scale dataset (AS-1B) for open-world panoptic visual recognition and understanding, using an economical semi-automatic data engine that combines the power of off-the-shelf… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/AS-100M.textn<1K16 likes2.3k downloads3y agoHugging Face04lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes296 downloads28d agoHugging Face05TempoFunk /hdvila-100Mtexttext-to-video10M<n<100M17 likes184 downloads3y agoHugging Face06codelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes61 downloads7mo agoHugging Face07ltg /babylm-2024-baby-cosmo-fine-100m@misc{charpentier2024gptbertboth, title={GPT or BERT: why not both?}, author={Lucas Georges Gabriel Charpentier and David Samuel}, year={2024}, eprint={2410.24159}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2410.24159}, } text100K<n<1M2 likes53 downloads2y agoHugging Face08codelion /sutra-improved-100M Sutra Improved 100M A self-improved pedagogical dataset for LLM pretraining, containing 413,899 entries totaling 110,038,011 tokens (~110 million). This dataset was created by applying an iterative self-improvement process to the Sutra-10B dataset, where each sample was rewritten using Gemma-3-4B-IT and only the better version (original or rewritten) was kept, followed by comprehensive deduplication and quality filtering. Dataset Description This dataset explores… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-improved-100M.texttext-generation100K<n<1M2 likes39 downloads6mo agoHugging Face09stanford-crfm /DSIR-filtered-pile-100M-short Dataset Card for DSIR-filtered-pile-100M-short Dataset Summary This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile. Languages English (EN) Dataset Structure A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.text10M<n<100M1 likes37 downloads4y agoHugging Face10Changyeli03 /llavasae_obliec100k_moststeps100Mtabularn<1K0 likes32 downloads2y agoHugging Face11chrisagrams /consensus-100M-mszx Consensus 100M (mszx) A 100M-scale mass spectrometry consensus dataset, distributed as mszx archive shards for use with the msdatasets library. Contents 90 shards: consensus_00.mszx ... consensus_89.mszx ~1.5 GB per shard, ~140 GB total manifest.json — shard list with sizes and sha256 checksums SHA256SUMS — sha256sum-format file for direct verification The shards are independent; row order across shards is not meaningful. Loading with msdatasets The Hugging… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/consensus-100M-mszx.tabularn<1K0 likes23 downloads5mo agoHugging Face12VertexResearch /Vertex-0.6-100M-self-identification Vertex 0.6 100M Self Identification A self-identification SFT dataset for Vertex-0.6-100M-8192-Instruct: 459 ChatML-style conversations that teach the model who it is: its name, creator, family, architecture, parameter count and knowledge cutoff. Made from SupraLabs/LLM-self-identification (Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6 100M identity: Marker Value MODEL_ID VertexResearch/Vertex-0.6-100M-8192-Instruct MODEL_NAME Vertex 0.6… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-100M-self-identification.texttext-generationn<1K0 likes20 downloads2d agoHugging Face13koiwave /100MMLUpro MMLU-Pro 100: A Balanced and Curated Evaluation Set Dataset Description This dataset is a curated, balanced subset of 100 questions derived from the TIGER-Lab/MMLU-Pro test set. It is designed to provide a small, fast, yet representative benchmark for evaluating the knowledge and reasoning capabilities of large language models across a wide range of academic and professional domains. The key feature of this dataset is its stratified sampling method, ensuring that the… See the full description on the dataset page: https://huggingface.co/datasets/koiwave/100MMLUpro.tabularmultiple-choicen<1K0 likes18 downloads1y agoHugging Face14Lambent /100-multiturn-task-sampling-elementary-phisynth-sharegpttextn<1K0 likes15 downloads2y agoHugging Face15Augmentoolkit /Openthoughts-100mil-DifferentFormattext1K<n<10K0 likes15 downloads1y agoHugging Face16fxmeng /transmla_pretrain_100m_tokenstext10K<n<100K0 likes12 downloads1y agoHugging Face17NarsAI /Mecan-ASI-Conversations-100Mtextn<1K0 likes8 downloads5mo agoHugging Face18TaiMingLu /finewebedu-test-100Mtext100K<n<1M0 likes5 downloads10mo agoHugging Face19Augmentoolkit /openthoughts-100mil-sharegptsubset of openthoughts 100k, 100mil tokens from across all shards in sharegpt format text1K<n<10K0 likes3 downloads2y agoHugging Face20BackpropBuff /simple_math_2_numbers_100mtext100M<n<1B1 likes2 downloads2y agoHugging Face21Shortheadband /100Mtextn<1K0 likes1 downloads2y agoHugging Face22marcosremar2 /iaratts-100M-ptbr-erinome-datasettext1K<n<10K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.