4k
Models
All models matching “4k”Datasets
All datasets matching “4k”Aesthetic-4K
Aesthetic-4K Dataset
We introduce Aesthetic-4K, a high-quality dataset for ultra-high-resolution image generation, featuring carefully selected images and captions generated by GPT-4o.
Additionally, we have meticulously filtered out low-quality images through manual inspection, excluding those with motion blur, focus issues, or mismatched text prompts.
For more details, please refer to our paper:
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models (CVPR… See the full description on the dataset page: https://huggingface.co/datasets/zhang0jhon/Aesthetic-4K.4KLSDB
4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation
DataCV @ CVPR 2026 · Accepted 🎉
4KLSDB is a native-4K image dataset with 129,484 train / 2,000 val / 1,984 test images, spanning nature, urban scenes, people, food, artwork, CGI, animals, and architecture. It supports both image restoration (super-resolution) and 4K text-to-image generation.
Quick links · 🌐 Project page · 💻 Code (GitHub) · 📄 Paper (arXiv) · 🤗 Dataset · 🧱 Checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/SingleBicycle/4KLSDB.DL3DV-ALL-4K
DL3DV-Dataset
This repo has all the 4K frames with camera poses of DL3DV-10K Dataset. We are working hard to review all the dataset to avoid sensitive information. Thank you for your patience.
Download
If you have enough space, you can use git to download a dataset from huggingface. See this link. 480P/960P versions should satisfies most needs.
If you do not have enough space, we further provide a download script here to download a subset. The usage:
usage: download.py… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-ALL-4K.reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs.
The dataset has these columns for users to filter out:
repo_id
tok_len
user
thought_trace
assistant
ChatML
Repositories… See the full description on the dataset page: https://huggingface.co/datasets/Qyrou/reasoning-corpus-4K-5M-v1.DRIFT-TL-Distill-4K
DRIFT-TL-Distill-4K Dataset
This dataset contains multimodal reasoning examples with images and step-by-step thinking processes.
Paper: Directional Reasoning Injection for Fine-Tuning MLLMs
Code/Project Page: https://github.com/WikiChao/DRIFT
Dataset Structure
Each example contains:
messages: Conversation between user and assistant with image references
images: Paths to associated images
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/ChaoHuangCS/DRIFT-TL-Distill-4K.qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4
qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3890625
Action score: 0.4359375
Valid samples: 320/320
