datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-align-enA-viA.speaker-embedding.xlsr-2bseamless-align-deA-enA.speaker-embedding.xlsr-2bseamless-align-enA-frA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.xlsr-2bPuro-2B
Puro-2B Pretraining Data: The Recipe Behind a 2B Model
This is the materialized pretraining data release for
Puro-2B-Base, a 2B base model
trained from scratch on consumer-grade RTX 5090 GPUs.
The repository contains the component-level data pools used to construct the
two Puro-2B pretraining phases, together with the tokenizer used for token
accounting. It is organized for inspection, selective streaming, and recipe
reconstruction rather than as a small train/test… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/Puro-2B.seamless-align-enA-jaA.speaker-embedding.xlsr-2bseamless-align-enA-hiA.speaker-embedding.xlsr-2bc4-tokenized-2b
Dataset Card for "c4-tokenized-2b"
More Information needed
seamless-align-enA-koA.speaker-embedding.xlsr-2blaion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.PCMind-2.1-Kaiyuan-2B
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/PCMind-2.1-Kaiyuan-2B.relaion2B-en-research-safegemma-2b-suite-explanationsrelaion2b-natural
LAION-Natural: Naturalness Scores for ReLAION-2B (CCN 2025, Roth & Hebart)
LAION-Natural is a large-scale naturalness scoring dataset covering 2.1 billion images from ReLAION-2B-en-research-safe. Each image receives a score predicting how "natural" or "photographic" it looks versus artificial/rendered content. At the recommended threshold of 0.7, the dataset identifies ~500 million natural photographs suitable for vision research, cognitive science, and model training.
Also… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural.sn80-data-run2brelaion2b-natural-embeddings
LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart)
LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7).
Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings
Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.c4-code-tokenized-2b
Dataset Card for "c4-code-tokenized-2b"
More Information needed
relaion2B-multi-research-safetl-2B-pretok
Beetle-Data/tl-2B-pretok
Pretokenized chunks (input_ids = 513-token packed sequences,
no cross-document bleeding). Sharded incrementally; the marker
data/_finalized.json is committed once all parts are uploaded.
SeeClear-396k
Dataset Structure
SeeClear-396K is a large-scale synthetic dataset accompanying our paper on transparent object depth estimation. It contains paired transparent and opaque renderings generated from identical scene geometry, camera poses, lighting conditions, and object configurations. The only difference between each pair is the material assigned to the target object, enabling explicit supervision for learning transparency-aware representations. For additional details about the… See the full description on the dataset page: https://huggingface.co/datasets/2bidoubi/SeeClear-396k.DFNDR-2B
Dataset Card for DFNDR-2B
This dataset contains synthetic captions, embeddings, and metadata for DFNDR-2B.
The metadata has been generated using pretrained image-text models on DFN-2B, a 2B filtered subset of DataComp-12B.
For details on how to use the metadata, please visit our ml-mobileclip repository.
For code to generate multi-modal reinforced datasets at large scale see ml-mobileclip-dr repository.
Note that this release does not contain original ground-truth captions. Please… See the full description on the dataset page: https://huggingface.co/datasets/apple/DFNDR-2B.sft-robo2-data-place_a2b_left
SFT-Robo2 Expert Data: place_a2b_left
Expert demonstration dataset for the place_a2b_left task from RoboTwin 2.0, for SFT training of OpenVLA-OFT following SimpleVLA-RL (arXiv:2509.09674).
Structure
aloha/ - ALOHA-format HDF5 (950 train / 50 val)
rlds/ - RLDS/TFDS format (training-ready for OpenVLA-OFT)
Details
1000 expert demonstrations via curobo motion planner
Single-view (head camera) + proprioception
14D action space (bimanual ALOHA: 7 per arm including… See the full description on the dataset page: https://huggingface.co/datasets/Louisnguyen/sft-robo2-data-place_a2b_left.CLIP-ViT-H-14-laion2B-s32B-b79K-all-checkpointsThis repository contains the intermediate checkpoints for the model https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K.
Each "epoch" corresponds to an additional (32B / 256) samples seen, consituting total of 256 "epochs"
The purpose of releasing these checkpoints and optimizer states is to enable analysis.
For the first 121 "epochs", training was done with float16 mixed precision before switching to bfloat16 after a loss blow up.
relaion2B-multi-researchpile-small-tokenized-2b
Dataset Card for "pile-small-tokenized-2b"
More Information needed
laion2b_multi_korean_subset_with_image
laion2b_multi_korean_subset_with_image
img2dataset을 통해 다운로드에 성공한 Bingsu/laion2B-multi-korean-subset 이미지를 정리한 데이터셋입니다.
이미지는 9,800,137장입니다.
이미지는 짧은 쪽 길이가 256이 되도록 리사이즈 되었으며, 품질 100인 webp파일로 다운로드 되었습니다.
Usage
1. datasets
>>> from datasets import load_dataset
>>> dataset = load_dataset("Bingsu/laion2b_multi_korean_subset_with_image", streaming=True, split="train")
>>> dataset.features
{'image': Image(decode=True, id=None),
'text': Value(dtype='string', id=None)… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2b_multi_korean_subset_with_image.laion2B-japanese-subsetrelaion2B-en-researchgemma-4-e2b-atlas
