CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes21k downloads2y agoHugging Face02asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face03asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes10k downloads2y agoHugging Face04asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.8k downloads2y agoHugging Face05asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.5k downloads2y agoHugging Face06thu-pacman /Puro-2B Puro-2B Pretraining Data: The Recipe Behind a 2B Model This is the materialized pretraining data release for Puro-2B-Base, a 2B base model trained from scratch on consumer-grade RTX 5090 GPUs. The repository contains the component-level data pools used to construct the two Puro-2B pretraining phases, together with the tokenizer used for token accounting. It is organized for inspection, selective streaming, and recipe reconstruction rather than as a small train/test… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/Puro-2B.texttext-generation100M<n<1B8 likes7.3k downloads1d agoHugging Face07asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.9k downloads2y agoHugging Face08asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.9k downloads2y agoHugging Face09NeelNanda /c4-tokenized-2b Dataset Card for "c4-tokenized-2b" More Information needed 1M<n<10M0 likes5.8k downloads4y agoHugging Face10asahi417 /seamless-align-enA-koA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes3.5k downloads2y agoHugging Face11GeroldMeisinger /laion2b-en-a65_cogvlm2-4bit_captions Abstract This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8). The synthetic images are best viewed locally by cloning this repo with: git lfs install git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.imageimage-classification1K<n<10K6 likes3k downloads2y agoHugging Face12thu-pacman /PCMind-2.1-Kaiyuan-2B This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/PCMind-2.1-Kaiyuan-2B.texttext-generation1B<n<10B5 likes2.4k downloads10mo agoHugging Face13laion /relaion2B-en-research-safegatedimage1B<n<10B227 likes2.3k downloads2y agoHugging Face14kisate-team /gemma-2b-suite-explanationstext1M<n<10M0 likes2.2k downloads2y agoHugging Face15andropar /relaion2b-natural LAION-Natural: Naturalness Scores for ReLAION-2B (CCN 2025, Roth & Hebart) LAION-Natural is a large-scale naturalness scoring dataset covering 2.1 billion images from ReLAION-2B-en-research-safe. Each image receives a score predicting how "natural" or "photographic" it looks versus artificial/rendered content. At the recommended threshold of 0.7, the dataset identifies ~500 million natural photographs suitable for vision research, cognitive science, and model training. Also… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural.imageimage-classification1B<n<10B5 likes2.1k downloads6mo agoHugging Face16dsaddsaf /sn80-data-run2b0 likes1.8k downloads2mo agoHugging Face17andropar /relaion2b-natural-embeddings LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart) LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7). Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.tabularfeature-extraction100M<n<1B1 likes1.5k downloads6mo agoHugging Face18NeelNanda /c4-code-tokenized-2b Dataset Card for "c4-code-tokenized-2b" More Information needed 1M<n<10M1 likes1.5k downloads4y agoHugging Face19laion /relaion2B-multi-research-safegatedimage1B<n<10B48 likes1.4k downloads2y agoHugging Face20Beetle-Data /tl-2B-pretok Beetle-Data/tl-2B-pretok Pretokenized chunks (input_ids = 513-token packed sequences, no cross-document bleeding). Sharded incrementally; the marker data/_finalized.json is committed once all parts are uploaded. 1M<n<10M1 likes1.3k downloads4mo agoHugging Face212bidoubi /SeeClear-396k Dataset Structure SeeClear-396K is a large-scale synthetic dataset accompanying our paper on transparent object depth estimation. It contains paired transparent and opaque renderings generated from identical scene geometry, camera poses, lighting conditions, and object configurations. The only difference between each pair is the material assigned to the target object, enabling explicit supervision for learning transparency-aware representations. For additional details about the… See the full description on the dataset page: https://huggingface.co/datasets/2bidoubi/SeeClear-396k.1 likes1.2k downloads3mo agoHugging Face22apple /DFNDR-2B Dataset Card for DFNDR-2B This dataset contains synthetic captions, embeddings, and metadata for DFNDR-2B. The metadata has been generated using pretrained image-text models on DFN-2B, a 2B filtered subset of DataComp-12B. For details on how to use the metadata, please visit our ml-mobileclip repository. For code to generate multi-modal reinforced datasets at large scale see ml-mobileclip-dr repository. Note that this release does not contain original ground-truth captions. Please… See the full description on the dataset page: https://huggingface.co/datasets/apple/DFNDR-2B.text-to-image1B<n<10B8 likes1.2k downloads4mo agoHugging Face23Louisnguyen /sft-robo2-data-place_a2b_left SFT-Robo2 Expert Data: place_a2b_left Expert demonstration dataset for the place_a2b_left task from RoboTwin 2.0, for SFT training of OpenVLA-OFT following SimpleVLA-RL (arXiv:2509.09674). Structure aloha/ - ALOHA-format HDF5 (950 train / 50 val) rlds/ - RLDS/TFDS format (training-ready for OpenVLA-OFT) Details 1000 expert demonstrations via curobo motion planner Single-view (head camera) + proprioception 14D action space (bimanual ALOHA: 7 per arm including… See the full description on the dataset page: https://huggingface.co/datasets/Louisnguyen/sft-robo2-data-place_a2b_left.robotics0 likes1.2k downloads5mo agoHugging Face24laion /CLIP-ViT-H-14-laion2B-s32B-b79K-all-checkpointsThis repository contains the intermediate checkpoints for the model https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K. Each "epoch" corresponds to an additional (32B / 256) samples seen, consituting total of 256 "epochs" The purpose of releasing these checkpoints and optimizer states is to enable analysis. For the first 121 "epochs", training was done with float16 mixed precision before switching to bfloat16 after a loss blow up. 1 likes1k downloads8mo agoHugging Face25laion /relaion2B-multi-researchgatedimage1B<n<10B12 likes1k downloads2y agoHugging Face26NeelNanda /pile-small-tokenized-2b Dataset Card for "pile-small-tokenized-2b" More Information needed 10M<n<100M0 likes978 downloads4y agoHugging Face27Bingsu /laion2b_multi_korean_subset_with_image laion2b_multi_korean_subset_with_image img2dataset을 통해 다운로드에 성공한 Bingsu/laion2B-multi-korean-subset 이미지를 정리한 데이터셋입니다. 이미지는 9,800,137장입니다. 이미지는 짧은 쪽 길이가 256이 되도록 리사이즈 되었으며, 품질 100인 webp파일로 다운로드 되었습니다. Usage 1. datasets >>> from datasets import load_dataset >>> dataset = load_dataset("Bingsu/laion2b_multi_korean_subset_with_image", streaming=True, split="train") >>> dataset.features {'image': Image(decode=True, id=None), 'text': Value(dtype='string', id=None)… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2b_multi_korean_subset_with_image.imagefeature-extraction100K<n<1M6 likes940 downloads4y agoHugging Face28lyakaap /laion2B-japanese-subsetimage100M<n<1B4 likes848 downloads4y agoHugging Face29laion /relaion2B-en-researchgatedimage1B<n<10B45 likes823 downloads2y agoHugging Face30juiceb0xc0de /gemma-4-e2b-atlas image1M<n<10M4 likes800 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.