datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.Krea-2-Turbo-Checkpoint-Format-Benchmark
Krea 2 Turbo ComfyUI Format Fidelity Benchmark
This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code.
Main result
BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.SEA-Dataset
SEA-Dataset by Kreasof AI
The SEA-Dataset is a large-scale, multilingual, and instruction-based dataset curated by Kreasof AI. It combines over 34 high-quality, publicly available datasets, with a significant focus on enhancing the representation of Southeast Asian (SEA) languages. This dataset is designed for training and fine-tuning large language models (LLMs) to be more capable in a variety of domains including reasoning, mathematics, coding, and multilingual tasks, while also… See the full description on the dataset page: https://huggingface.co/datasets/kreasof-ai/SEA-Dataset.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.krea2-wildcards
krea2-wildcards — Review Evidence
Human-review contact sheets and their source renders for the
krea2-wildcards prompt library: a project that
compiles researched visual vocabulary — art styles, linework and coloring, poses, lighting, camera
composition, character design — into ComfyUI Dynamic Prompts wildcards for the Krea2 model family.
This dataset isn't the wildcard library itself. It's the photographic evidence a reviewer looked at
when deciding, for each catalog entry… See the full description on the dataset page: https://huggingface.co/datasets/innofree/krea2-wildcards.Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
The images were created in ComfyUI with the
bf16 version
of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.kream-product-blip-captions
KREAM Product Blip Captions Dataset Information
KREAM Product Blip Captions Dataset is a dataset card for finetuning a text-to-image generative model collected from KREAM, one of the best online-resell market in Korea.
This dataset consists of 'image' and 'text' key pairs.
The format of 'text' is 'category (e.g. outer), product original name (e.g. The North Face 1996 Eco Nuptse Jacket Black), blip captions (e.g. a photography of the north face black down jacket)'.
You can easily… See the full description on the dataset page: https://huggingface.co/datasets/hahminlew/kream-product-blip-captions.SPL-Combinedkrea2anna-lush-krea2-datasetkrea2-skin-lora
Inline Skin LoRA - training dataset
The 26 image + caption pairs used to train inlineresearch/skin-lora-krea-2-raw, a photorealistic skin-texture LoRA for Krea 2. Close-up and half-face portraits chosen for real, unretouched skin: visible pores, fine hair, natural texture and imperfection, across a range of skin tones, ages, and lighting.
Contents
data/
metadata.jsonl captions for the dataset viewer (one row per image: file_name + text)
0000.jpg… See the full description on the dataset page: https://huggingface.co/datasets/inlineresearch/krea2-skin-lora.kream-fashion-anchor-positive-imageskrea2-benchKAC-QuantFin-1M
Kreasof AI Capital (KAC) QuantFin-1M
We employ GARCH-based synthetic price generation, combined with our proprietary trading algorithm. The synthetic price path originally length in 28800 steps with 1-minute interval, but subsampled to 1028 steps with 28-minutes interval to save storage.
License: cc-by-nc-sa-4.0 (non-commercial)
bigc-bem-eng
Dataset Details
This is dataset of speech translation task for Bemba-to-English Language. This dataset is acquired from (Big-C)[https://github.com/csikasote/bigc] github repository.
Big-C is a large conversations dataset between Bemba Speakers based on Image [1]. This dataset provide data for speech translation.
Preprocessing Steps
Some preprocessing was done in this dataset.
Drop some unused columns other than audio_id, sentence, translation, and speaker_id.
Investigate… See the full description on the dataset page: https://huggingface.co/datasets/kreasof-ai/bigc-bem-eng.Akihito_Tsukushi_Krea_2_Captionskrea2-scene-linear-hdr-dataset
Krea 2 scene-linear HDR — LogC4 training dataset
320 scene-linear HDR training images for the Krea 2 HDR LoRA
(see model: LAXMAYDAY/krea2-scene-linear-hdr-lora
· tools: krea2-hdr-toolkit).
Built by re-projecting CC0 Poly Haven HDRIs into normal perspective scene images,
exposure-normalizing (mid-gray → 0.18), and LogC4-encoding to [0,1] so a wide
scene-linear range is carried through an ordinary 8-bit pipeline.
What each sample is
images/<slug>__v<k>.png — an… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/krea2-scene-linear-hdr-dataset.GLM-Kimi-OpenThoughts-HunterAlpha-Filtered
domain
Mean_In
P95_In
Mean_Out
P95_Out
Total_Tokens
General-Distillation
92.81
378
2032.89
3777
784362123
General-Math
51.65
72
3422
4007
5578684
Math
58.61
85
3323.74
3991
899705
Multilingual-STEM
64.01
89
3548.07
3974
159495080
MultilingualSTEM
76.5
119
3097.35
3917
147082677
PHD-Science
44.39
55
3082.65
3857
532525113
code
404.92
1209
2554.44
3721
3536439
general
62.62
2941427.94
3496
301834703
main
96.01
390
2170.94
3776
832633275
math
76.88
147
3313.32
3979… See the full description on the dataset page: https://huggingface.co/datasets/kreasof-ai/GLM-Kimi-OpenThoughts-HunterAlpha-Filtered.SEA-Dataset-LiteThis is lite version of kreasof-ai/SEA-Dataset with max_seq_length = 1024
danbooru-krea2-2m-metadata
Danbooru Krea2 2M Metadata
This gated dataset is the image-free, portable definition of a frozen Danbooru
training selection. It contains no images, latent tensors, text-encoder tensors,
or visual embeddings.
Contents
records: 2,133,904 unique selected posts in deterministic
manifest_order, with original categorized tags, supplemental general tags,
final training captions, alternate captions, repaired VLM caption fields, OCR,
quality/duplicate signals, exact… See the full description on the dataset page: https://huggingface.co/datasets/SumomoLee/danbooru-krea2-2m-metadata.krea2workspaceKrea2-workflowskrea2_hires_inpaintkiss-shot_krea2_lora_captionsaistudio-realistic-snapshot-krea2krea_outputskrea-2-models-backupSPL-400K-math-shepherdvinne_krea2_captionspose-data-1
