datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.relaion2B-en-research-saferelaion2b-natural
LAION-Natural: Naturalness Scores for ReLAION-2B (CCN 2025, Roth & Hebart)
LAION-Natural is a large-scale naturalness scoring dataset covering 2.1 billion images from ReLAION-2B-en-research-safe. Each image receives a score predicting how "natural" or "photographic" it looks versus artificial/rendered content. At the recommended threshold of 0.7, the dataset identifies ~500 million natural photographs suitable for vision research, cognitive science, and model training.
Also… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural.relaion2B-multi-research-saferelaion2B-multi-researchlaion2b_multi_korean_subset_with_image
laion2b_multi_korean_subset_with_image
img2dataset을 통해 다운로드에 성공한 Bingsu/laion2B-multi-korean-subset 이미지를 정리한 데이터셋입니다.
이미지는 9,800,137장입니다.
이미지는 짧은 쪽 길이가 256이 되도록 리사이즈 되었으며, 품질 100인 webp파일로 다운로드 되었습니다.
Usage
1. datasets
>>> from datasets import load_dataset
>>> dataset = load_dataset("Bingsu/laion2b_multi_korean_subset_with_image", streaming=True, split="train")
>>> dataset.features
{'image': Image(decode=True, id=None),
'text': Value(dtype='string', id=None)… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2b_multi_korean_subset_with_image.laion2B-japanese-subsetrelaion2B-en-researchgemma-4-e2b-atlas
tmax-2b-atlas
juiceb0xc0de/tmax-2b-atlas
A brain atlas for allenai/tmax-2b, a hybrid SSM/Mamba/transformer language model. This is not a chat dataset or a benchmark — it is an internal-mechanics map of the model, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where the model stores compliance style, which late-layer directions you can edit without breaking reasoning, or whether the… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-2b-atlas.laion-2b-en-unsafe-quarter-one-downloadRoboTwin_place_a2b_left_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 71244,
"total_tasks": 500,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_place_a2b_left_randomized.RoboTwin_place_a2b_right_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 71049,
"total_tasks": 500,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_place_a2b_right_randomized.Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts
Snowball 67B-A2B RL artifact release
2026 mixed-domain RLVR campaign
This release also contains the complete releasable record of the September 2026 Snowball mixed-domain RLVR campaign.
It covers the September 11 synchronous and bounded-staleness asynchronous RLVR1→RLVR2 lineages and the 5.7T
Agentic-start RLVR1 lineage. All training arms are terminal. The final campaign figure,
trace audit, canonical configs, timing reports, retained
traces, and operational… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts.screen-highlighter-2b-highlight-only-v1-rolloutsQwen3.5-2B-Base
juiceb0xc0de/Qwen3.5-2B-Base
A brain atlas for Qwen/Qwen3.5-2B-Base, a 24-layer hybrid that runs linear attention on 18 layers and full attention on the other 6. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-2B-Base.laion2b_seed
Dataset Card for "laion2b_seed"
This dataset is a subset of laion2B-en-aesthetic, with SEED v1 tokens.
laion-2b-en-unsafe-quarter-four-downloadscreen-highlighter-2b-v1-rolloutslaion2B-en-aestheticlaion2B-multi-chinese-subset
laion2B-multi-chinese-subset
Github: Fengshenbang-LM
Docs: Fengshenbang-Docs
简介 Brief Introduction
取自Laion2B多语言多模态数据集中的中文部分,一共143M个图文对。
A subset from Laion2B (a multimodal dataset), around 143M image-text pairs (only Chinese).
数据集信息 Dataset Information
大约一共143M个中文图文对。大约占用19GB空间(仅仅是url等文本信息,不包含图片)。
Homepage: laion-5b
Huggingface: laion/laion2B-multi
下载 Download
mkdir laion2b_chinese_release && cd laion2b_chinese_release
for i in {00000..00012}; do… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-CCNL/laion2B-multi-chinese-subset.laion-2b-en-unsafe-quarter-four-download-filtered2b2t_1mil
The 1,024,000² 2b2t World Download Project (1M²). And More.
It's finally here. Totalling 13.7 TiB of highly compressed 2b2t world data, which includes the following:
a 1,024,000² (1M²) area of the Overworld (Dec 25 2025 - Apr 13 2026),
a 512,000² (512k²) area of the Overworld (Nov 11 2024 - Dec 12 2024),
a 256,000² (256k²) area of the End (Jan 23 2026 - Feb 15 2026),
a 100,000² (100k²) area of the Nether (Jun 9 2025 - Jun 14 2025)
This wasn't easy to accomplish in the… See the full description on the dataset page: https://huggingface.co/datasets/obvtiger/2b2t_1mil.laion-2b-en-unsafe-quarter-four-download-hdrelaion2B-en-research-safe-japanese-translation
relaion2B-en-research-safe-japanese-translation
This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it.
We used text2dataset for translating with open-weight LLMs.
By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese.
Prompt
The following is the prompt used for translation with Gemma.
You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.laion2b_en_sd2.1baseA dataset for SD2.1 base training, which contains metadata filtered from laion2b-en with the following conditions.
WIDTH>=512
HEIGHT>=512
punsafe<=0.98
AESTHETIC_SCORE>=4.5
laion-2b-en-unsafe-quarter-three-download-hdlaion2B-multi-joined-translated-to-en-smolmhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.
