6M
Models
All models matching “6M”Datasets
All datasets matching “6M”RefVideo6M
RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
If you use RefVideo-6M in your research, please cite our work as follows:
@article{zi2026refvideo6m
title={RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing},
author={Bojia Zi and Xiaoyan Yang and Yu Zhou and Ruijie Sun and Lihan Zhang and Bin Liang and Kam-Fai Wong and Haibin Huang and Chi Zhang and Xuelong Li},
journal={arXiv preprint arXiv:2608.26101}… See the full description on the dataset page: https://huggingface.co/datasets/RefVideo6M/RefVideo6M.FLUX-Reason-6M
FLUX-Reason-6M
FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems.
This dataset contains:
6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model.
20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.FaceID-6M
FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset
This repository contains the dataset described in FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset.
Links
FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset
Introduction
Comparison with Previous Works
FaceID Fidelity
Scaling Results
Released FaceID-6M dataset
Released FaceID Customization Models
Usage
Contact
Introduction
FaceID-6M, is the first… See the full description on the dataset page: https://huggingface.co/datasets/Super-shuhe/FaceID-6M.LAION-High-Qualtiy-Pro-6M-VLV
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
LAION-High-Qualtiy-Pro-6M Dataset
This repository hosts LAION-High-Quality-Pro-6M, the image-text dataset we used to train Vision-Language-Vision models.
Example Usage:
# pip install -U datasets pillow
from datasets import load_dataset
from PIL import Image
import base64
import io
# Robust decoder: works if the column is base64 *or* raw bytes
import io
import… See the full description on the dataset page: https://huggingface.co/datasets/ccvl/LAION-High-Qualtiy-Pro-6M-VLV.BM-6M
Dataset Card for ByteMorph-6M
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present ByteMorph… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M.Demeter-LongCoT-6M
Demeter-LongCoT-6M
Demeter-LongCoT-6M is a high-quality, compact chain-of-thought reasoning dataset curated for tasks in mathematics, science, and coding. While the dataset spans diverse domains, it is primarily driven by mathematical reasoning, reflecting a major share of math-focused prompts and long-form logical solutions.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Demeter-LongCoT-6M.
