datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RefVideo6M
RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
If you use RefVideo-6M in your research, please cite our work as follows:
@article{zi2026refvideo6m
title={RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing},
author={Bojia Zi and Xiaoyan Yang and Yu Zhou and Ruijie Sun and Lihan Zhang and Bin Liang and Kam-Fai Wong and Haibin Huang and Chi Zhang and Xuelong Li},
journal={arXiv preprint arXiv:2608.26101}… See the full description on the dataset page: https://huggingface.co/datasets/RefVideo6M/RefVideo6M.FLUX-Reason-6M
FLUX-Reason-6M
FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems.
This dataset contains:
6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model.
20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.FaceID-6M
FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset
This repository contains the dataset described in FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset.
Links
FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset
Introduction
Comparison with Previous Works
FaceID Fidelity
Scaling Results
Released FaceID-6M dataset
Released FaceID Customization Models
Usage
Contact
Introduction
FaceID-6M, is the first… See the full description on the dataset page: https://huggingface.co/datasets/Super-shuhe/FaceID-6M.LAION-High-Qualtiy-Pro-6M-VLV
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
LAION-High-Qualtiy-Pro-6M Dataset
This repository hosts LAION-High-Quality-Pro-6M, the image-text dataset we used to train Vision-Language-Vision models.
Example Usage:
# pip install -U datasets pillow
from datasets import load_dataset
from PIL import Image
import base64
import io
# Robust decoder: works if the column is base64 *or* raw bytes
import io
import… See the full description on the dataset page: https://huggingface.co/datasets/ccvl/LAION-High-Qualtiy-Pro-6M-VLV.BM-6M
Dataset Card for ByteMorph-6M
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present ByteMorph… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M.Demeter-LongCoT-6M
Demeter-LongCoT-6M
Demeter-LongCoT-6M is a high-quality, compact chain-of-thought reasoning dataset curated for tasks in mathematics, science, and coding. While the dataset spans diverse domains, it is primarily driven by mathematical reasoning, reflecting a major share of math-focused prompts and long-form logical solutions.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Demeter-LongCoT-6M.ua-case-outcome-6m
Ukrainian Court Decisions: Case Outcome Prediction (6.7M)
The largest publicly available dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (EDRSR). Contains 6,690,284 substantive decisions from civil and commercial courts spanning 2008--2026, with temporal splits across three wartime epochs.
Overview
Ukraine's EDRSR is one of the world's largest open judicial databases, containing 100M+ judicial… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-case-outcome-6m.BM-6M
Dataset Card for ByteMorph-6M
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present ByteMorph… See the full description on the dataset page: https://huggingface.co/datasets/ByteMorph/BM-6M.Helios-R-6M
Helios-R-6M
Helios-R-6M is a high-quality, compact reasoning dataset designed to strengthen multi-step problem solving across mathematics, computer science, and scientific inquiry. While the dataset covers a range of disciplines, math constitutes the largest share of examples and drives the reasoning complexity.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset = load_dataset("prithivMLmods/Helios-R-6M"… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Helios-R-6M.fineweb-edu-dedup6m
Stage 1 (S1): General Knowledge Anchor — 6M FineWeb-Edu-Dedup
1. Project Overview
This dataset represents the General Knowledge Acquisition Phase (S1) for a research project focused on developing a Domain-Adaptive LLM for ISO 27001 Information Security Auditing.
S1 serves as the cognitive foundation. This corpus is designed to establish high-level linguistic proficiency and general reasoning before the introduction of specialized regulatory standards in Stage 2.… See the full description on the dataset page: https://huggingface.co/datasets/JoTeqtheFirstAI/fineweb-edu-dedup6m.ByteMorph-6M-Demo
Dataset Card for ByteMorph-6M-Demo
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/Boese0601/ByteMorph-6M-Demo.fineweb-edu-dedup6mqa_verify_cot_new_6M_unfiltered_v7dataset_names = [
"HayatoHongoEveryonesAI/qa_verify_1m_cot_1",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_2",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_3",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_4",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_5",
"HayatoHongoEveryonesAI/qa_verify_2m_cot_2",
"HayatoHongoEveryonesAI/qa_verify_2m_cot_3",
]
https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing
openai-vs-anthropic-news-coverage-6mo-2025-2026
OpenAI vs Anthropic News Coverage (6 Months)
Dataset Summary
This dataset contains English-language news articles that mention OpenAI or Anthropic in the article title or description. It is designed for analyzing media coverage volume and trends over time, not sentiment or opinion.
The dataset covers approximately six months of news and includes both article-level data and a derived weekly aggregation.
Data Collection
Articles were collected using keyword-based… See the full description on the dataset page: https://huggingface.co/datasets/NewsDataHub/openai-vs-anthropic-news-coverage-6mo-2025-2026.BM-6M-Demo
Dataset Card for ByteMorph-6M-Demo
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M-Demo.filesystem_huggingface_5053_cl6lee6m
Support Ticket Triage Corpus
Dataset ID: ZorakTriage94b837
Customer support ticket records with priority, status, and satisfaction annotations.
AE_english_data_stage_3-6M-6.5Mcornstack-6-langs-v1-tevatron-6MAuto-6MLlaion6m_recapspell_6m_mixqa_verify_cot_new_6M_v6dataset_names = [
"HayatoHongoEveryonesAI/qa_verify_cot_new_5.1M_v7",
"HayatoHongoEveryonesAI/qa_verify_new_v6",
]
MARIO-6MThis datasets is curated by TextDiffuser team in their work:
TextDiffuser: Diffusion Models as Text Painters (NeurIPS 2023)
MARIO-6M contains 6M images with text rendered on, filtered from LAION-400M.
BM-6M-Demo
Dataset Card for ByteMorph-6M-Demo
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteMorph/BM-6M-Demo.vast-rtx3090-market-6mo
Vast.ai RTX 3090 Spot Market, February-August 2026
Panel data from the vast.ai GPU rental marketplace, restricted to NVIDIA RTX 3090 offers. The public offer listing was polled every 10 minutes between 2026-02-13 and 2026-08-15. Each observation records price, hardware specifications, host reliability, and location. A derived lifecycle table gives the listing duration of every offer. Vast.ai does not publish historical listing data; this dataset was collected independently.… See the full description on the dataset page: https://huggingface.co/datasets/MarcusLammers/vast-rtx3090-market-6mo.cosmopedia_6Mnllb-200-6M-sample-embeddingOriginal dataset
SONAR's author message
What happens to the original dataset?
Filter by blaser_sim >= 3.5
Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline
What are the use cases?
Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly
scaleedit-filtered-6m
ScaleEdit Filtered 6M
This repository contains a portable selection manifest for high-quality samples
from ScaleEdit-12M. It does not redistribute the source images. Download the
original ScaleEdit-12M Parquet files separately, then join each manifest row to
the source file named by source_relative_path at zero-based row_index.
Selection
For each evaluated source row, the latest successful stage-2 review was used.
A row is included when result.final_decision ==… See the full description on the dataset page: https://huggingface.co/datasets/QingyuShi/scaleedit-filtered-6m.Collection_1_6Maywiki6m-selfdoc-final_v2
wiki6m-selfdoc-final_v2
GJ's cleaned-context wiki6M blocks (V4 arm search_1p7B_selfdoc_v2): tiered queries, = LLM-written answer from the source doc, no BM25. 640 parquet shards at repo root, 1,562,273 blocks x 4096, 26% masked. meta/: holdout + stats.
RCP source: /mloscratch/homes/ponkshe/searchllm_dt_runs/hf_stage/selfdoc_parts_search. Pushed 2026-09-15 by ops/push_datasets_to_hf.py. Provenance: Search-LLM EXPERIMENT_PLAN.md / RESULTS.md / INVENTORY.md.
