datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FLUX-Reason-6M
FLUX-Reason-6M
FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems.
This dataset contains:
6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model.
20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.LAION-High-Qualtiy-Pro-6M-VLV
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
LAION-High-Qualtiy-Pro-6M Dataset
This repository hosts LAION-High-Quality-Pro-6M, the image-text dataset we used to train Vision-Language-Vision models.
Example Usage:
# pip install -U datasets pillow
from datasets import load_dataset
from PIL import Image
import base64
import io
# Robust decoder: works if the column is base64 *or* raw bytes
import io
import… See the full description on the dataset page: https://huggingface.co/datasets/ccvl/LAION-High-Qualtiy-Pro-6M-VLV.Demeter-LongCoT-6M
Demeter-LongCoT-6M
Demeter-LongCoT-6M is a high-quality, compact chain-of-thought reasoning dataset curated for tasks in mathematics, science, and coding. While the dataset spans diverse domains, it is primarily driven by mathematical reasoning, reflecting a major share of math-focused prompts and long-form logical solutions.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Demeter-LongCoT-6M.ua-case-outcome-6m
Ukrainian Court Decisions: Case Outcome Prediction (6.7M)
The largest publicly available dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (EDRSR). Contains 6,690,284 substantive decisions from civil and commercial courts spanning 2008--2026, with temporal splits across three wartime epochs.
Overview
Ukraine's EDRSR is one of the world's largest open judicial databases, containing 100M+ judicial… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-case-outcome-6m.Helios-R-6M
Helios-R-6M
Helios-R-6M is a high-quality, compact reasoning dataset designed to strengthen multi-step problem solving across mathematics, computer science, and scientific inquiry. While the dataset covers a range of disciplines, math constitutes the largest share of examples and drives the reasoning complexity.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset = load_dataset("prithivMLmods/Helios-R-6M"… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Helios-R-6M.ByteMorph-6M-Demo
Dataset Card for ByteMorph-6M-Demo
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/Boese0601/ByteMorph-6M-Demo.qa_verify_cot_new_6M_unfiltered_v7dataset_names = [
"HayatoHongoEveryonesAI/qa_verify_1m_cot_1",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_2",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_3",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_4",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_5",
"HayatoHongoEveryonesAI/qa_verify_2m_cot_2",
"HayatoHongoEveryonesAI/qa_verify_2m_cot_3",
]
https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing
fineweb-edu-dedup6mBM-6M-Demo
Dataset Card for ByteMorph-6M-Demo
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M-Demo.filesystem_huggingface_5053_cl6lee6m
Support Ticket Triage Corpus
Dataset ID: ZorakTriage94b837
Customer support ticket records with priority, status, and satisfaction annotations.
cornstack-6-langs-v1-tevatron-6Mlaion6m_recapAE_english_data_stage_3-6M-6.5Mspell_6m_mixMARIO-6MThis datasets is curated by TextDiffuser team in their work:
TextDiffuser: Diffusion Models as Text Painters (NeurIPS 2023)
MARIO-6M contains 6M images with text rendered on, filtered from LAION-400M.
qa_verify_cot_new_6M_v6dataset_names = [
"HayatoHongoEveryonesAI/qa_verify_cot_new_5.1M_v7",
"HayatoHongoEveryonesAI/qa_verify_new_v6",
]
BM-6M-Demo
Dataset Card for ByteMorph-6M-Demo
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteMorph/BM-6M-Demo.arbiter_6mvast-rtx3090-market-6mo
Vast.ai RTX 3090 Spot Market, February-August 2026
Panel data from the vast.ai GPU rental marketplace, restricted to NVIDIA RTX 3090 offers. The public offer listing was polled every 10 minutes between 2026-02-13 and 2026-08-15. Each observation records price, hardware specifications, host reliability, and location. A derived lifecycle table gives the listing duration of every offer. Vast.ai does not publish historical listing data; this dataset was collected independently.… See the full description on the dataset page: https://huggingface.co/datasets/MarcusLammers/vast-rtx3090-market-6mo.cosmopedia_6Mnllb-200-6M-sample-embeddingOriginal dataset
SONAR's author message
What happens to the original dataset?
Filter by blaser_sim >= 3.5
Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline
What are the use cases?
Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly
scaleedit-filtered-6m
ScaleEdit Filtered 6M
This repository contains a portable selection manifest for high-quality samples
from ScaleEdit-12M. It does not redistribute the source images. Download the
original ScaleEdit-12M Parquet files separately, then join each manifest row to
the source file named by source_relative_path at zero-based row_index.
Selection
For each evaluated source row, the latest successful stage-2 review was used.
A row is included when result.final_decision ==… See the full description on the dataset page: https://huggingface.co/datasets/QingyuShi/scaleedit-filtered-6m.concept_coverage_laion_6m
📦 Freeze-Align Dataset
The Freeze-Align Dataset (concept_coverage_laion_6m) is a curated collection of high-quality image-text pairs designed to facilitate efficient multimodal alignment using frozen unimodal encoders. This dataset supports the research presented in our CVPR 2025 paper, "Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment", enabling models to achieve CLIP-level performance with significantly reduced computational resources.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/mayug/concept_coverage_laion_6m.SyntheticTexts6MThis synthetic dataset is generated with russian context-free grammar. It contains about ~88M tokens.
fineweb-edu-2024-10-from-5M-to-6MBM-6Mdl_alchemy_seq9p6m_context1024fineweb-edu-2024-10-from-6M-to-7MAllTheBacteria-FCGR-6mer
Dataset Card for AllTheBacteria-FCGR-6mer
This dataset contains the
Frequency matrix of the Chaos Game Representation of DNA (FCGR) using 6-mers for the AllTheBacteria dataset paper
The set of assemblies used to create the FCGRs can be found here
https://ftp.ebi.ac.uk/pub/databases/AllTheBacteria/Releases/0.2/
Dataset Details
The dataset is ordered the same way the assemblies are provided, each .tar.gz file contains a collection of FCGR, one for each assembly.
FCGRs… See the full description on the dataset page: https://huggingface.co/datasets/koke/AllTheBacteria-FCGR-6mer.btc-15min-6months
