CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LucasFang /FLUX-Reason-6M FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems. This dataset contains: 6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model. 20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.image1M<n<10M112 likes9.4k downloads8mo agoHugging Face02ccvl /LAION-High-Qualtiy-Pro-6M-VLV Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models LAION-High-Qualtiy-Pro-6M Dataset This repository hosts LAION-High-Quality-Pro-6M, the image-text dataset we used to train Vision-Language-Vision models. Example Usage: # pip install -U datasets pillow from datasets import load_dataset from PIL import Image import base64 import io # Robust decoder: works if the column is base64 *or* raw bytes import io import… See the full description on the dataset page: https://huggingface.co/datasets/ccvl/LAION-High-Qualtiy-Pro-6M-VLV.textimage-to-text1M<n<10M4 likes1.9k downloads1y agoHugging Face03prithivMLmods /Demeter-LongCoT-6M Demeter-LongCoT-6M Demeter-LongCoT-6M is a high-quality, compact chain-of-thought reasoning dataset curated for tasks in mathematics, science, and coding. While the dataset spans diverse domains, it is primarily driven by mathematical reasoning, reflecting a major share of math-focused prompts and long-form logical solutions. Quick Start with Hugging Face Datasets🤗 pip install -U datasets from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Demeter-LongCoT-6M.texttext-generation1M<n<10M5 likes1.3k downloads4mo agoHugging Face04overthelex /ua-case-outcome-6m Ukrainian Court Decisions: Case Outcome Prediction (6.7M) The largest publicly available dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (EDRSR). Contains 6,690,284 substantive decisions from civil and commercial courts spanning 2008--2026, with temporal splits across three wartime epochs. Overview Ukraine's EDRSR is one of the world's largest open judicial databases, containing 100M+ judicial… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-case-outcome-6m.texttext-classification1M<n<10M0 likes654 downloads4mo agoHugging Face05prithivMLmods /Helios-R-6M Helios-R-6M Helios-R-6M is a high-quality, compact reasoning dataset designed to strengthen multi-step problem solving across mathematics, computer science, and scientific inquiry. While the dataset covers a range of disciplines, math constitutes the largest share of examples and drives the reasoning complexity. Quick Start with Hugging Face Datasets🤗 pip install -U datasets from datasets import load_dataset dataset = load_dataset("prithivMLmods/Helios-R-6M"… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Helios-R-6M.texttext-generation1M<n<10M3 likes470 downloads4mo agoHugging Face06Boese0601 /ByteMorph-6M-Demo Dataset Card for ByteMorph-6M-Demo The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/Boese0601/ByteMorph-6M-Demo.imageimage-to-image100K<n<1M1 likes405 downloads1y agoHugging Face07HayatoHongoEveryonesAI /qa_verify_cot_new_6M_unfiltered_v7dataset_names = [ "HayatoHongoEveryonesAI/qa_verify_1m_cot_1", "HayatoHongoEveryonesAI/qa_verify_1m_cot_2", "HayatoHongoEveryonesAI/qa_verify_1m_cot_3", "HayatoHongoEveryonesAI/qa_verify_1m_cot_4", "HayatoHongoEveryonesAI/qa_verify_1m_cot_5", "HayatoHongoEveryonesAI/qa_verify_2m_cot_2", "HayatoHongoEveryonesAI/qa_verify_2m_cot_3", ] https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing tabular1M<n<10M0 likes305 downloads8mo agoHugging Face08NanoMatriX /fineweb-edu-dedup6mtext1M<n<10M0 likes266 downloads8mo agoHugging Face09ByteDance-Seed /BM-6M-Demo Dataset Card for ByteMorph-6M-Demo The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M-Demo.imageimage-to-image100K<n<1M3 likes245 downloads1y agoHugging Face10TianfuXinqu /filesystem_huggingface_5053_cl6lee6m Support Ticket Triage Corpus Dataset ID: ZorakTriage94b837 Customer support ticket records with priority, status, and satisfaction annotations. textn<1K0 likes223 downloads1mo agoHugging Face11nthakur /cornstack-6-langs-v1-tevatron-6Mtext1M<n<10M0 likes150 downloads1y agoHugging Face12wusize /laion6m_recapimage1M<n<10M0 likes131 downloads1y agoHugging Face13AdoCleanCode /AE_english_data_stage_3-6M-6.5Mtext100K<n<1M0 likes131 downloads8mo agoHugging Face14minhbui /spell_6m_mixtext1M<n<10M0 likes112 downloads2y agoHugging Face15stzhao /MARIO-6MThis datasets is curated by TextDiffuser team in their work: TextDiffuser: Diffusion Models as Text Painters (NeurIPS 2023) MARIO-6M contains 6M images with text rendered on, filtered from LAION-400M. imagetext-to-image1M<n<10M4 likes99 downloads2y agoHugging Face16HayatoHongoEveryonesAI /qa_verify_cot_new_6M_v6dataset_names = [ "HayatoHongoEveryonesAI/qa_verify_cot_new_5.1M_v7", "HayatoHongoEveryonesAI/qa_verify_new_v6", ] tabular1M<n<10M0 likes95 downloads8mo agoHugging Face17ByteMorph /BM-6M-Demo Dataset Card for ByteMorph-6M-Demo The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteMorph/BM-6M-Demo.imageimage-to-image100K<n<1M0 likes78 downloads1y agoHugging Face18jrahn /arbiter_6mtext1M<n<10M0 likes64 downloads2y agoHugging Face19MarcusLammers /vast-rtx3090-market-6mo Vast.ai RTX 3090 Spot Market, February-August 2026 Panel data from the vast.ai GPU rental marketplace, restricted to NVIDIA RTX 3090 offers. The public offer listing was polled every 10 minutes between 2026-02-13 and 2026-08-15. Each observation records price, hardware specifications, host reliability, and location. A derived lifecycle table gives the listing duration of every offer. Vast.ai does not publish historical listing data; this dataset was collected independently.… See the full description on the dataset page: https://huggingface.co/datasets/MarcusLammers/vast-rtx3090-market-6mo.tabular1M<n<10M0 likes64 downloads1mo agoHugging Face20HuggingFaceTB /cosmopedia_6Mtext1M<n<10M6 likes58 downloads2y agoHugging Face21jasonrichdarmawan /nllb-200-6M-sample-embeddingOriginal dataset SONAR's author message What happens to the original dataset? Filter by blaser_sim >= 3.5 Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline What are the use cases? Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly tabular1M<n<10M0 likes55 downloads1y agoHugging Face22QingyuShi /scaleedit-filtered-6m ScaleEdit Filtered 6M This repository contains a portable selection manifest for high-quality samples from ScaleEdit-12M. It does not redistribute the source images. Download the original ScaleEdit-12M Parquet files separately, then join each manifest row to the source file named by source_relative_path at zero-based row_index. Selection For each evaluated source row, the latest successful stage-2 review was used. A row is included when result.final_decision ==… See the full description on the dataset page: https://huggingface.co/datasets/QingyuShi/scaleedit-filtered-6m.tabular1M<n<10M0 likes51 downloads26d agoHugging Face23mayug /concept_coverage_laion_6m 📦 Freeze-Align Dataset The Freeze-Align Dataset (concept_coverage_laion_6m) is a curated collection of high-quality image-text pairs designed to facilitate efficient multimodal alignment using frozen unimodal encoders. This dataset supports the research presented in our CVPR 2025 paper, "Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment", enabling models to achieve CLIP-level performance with significantly reduced computational resources. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/mayug/concept_coverage_laion_6m.imagezero-shot-classification1M<n<10M1 likes46 downloads1y agoHugging Face24v-urushkin /SyntheticTexts6MThis synthetic dataset is generated with russian context-free grammar. It contains about ~88M tokens. texttext-generation1M<n<10M0 likes42 downloads1y agoHugging Face25ohsuz /fineweb-edu-2024-10-from-5M-to-6Mtext1M<n<10M0 likes41 downloads2y agoHugging Face26JackyZhuo /BM-6Mtext1M<n<10M0 likes39 downloads1y agoHugging Face27kothasuhas /dl_alchemy_seq9p6m_context1024tabularn<1K0 likes37 downloads15d agoHugging Face28ohsuz /fineweb-edu-2024-10-from-6M-to-7Mtext1M<n<10M0 likes34 downloads2y agoHugging Face29koke /AllTheBacteria-FCGR-6mer Dataset Card for AllTheBacteria-FCGR-6mer This dataset contains the Frequency matrix of the Chaos Game Representation of DNA (FCGR) using 6-mers for the AllTheBacteria dataset paper The set of assemblies used to create the FCGRs can be found here https://ftp.ebi.ac.uk/pub/databases/AllTheBacteria/Releases/0.2/ Dataset Details The dataset is ordered the same way the assemblies are provided, each .tar.gz file contains a collection of FCGR, one for each assembly. FCGRs… See the full description on the dataset page: https://huggingface.co/datasets/koke/AllTheBacteria-FCGR-6mer.text1M<n<10M0 likes29 downloads1y agoHugging Face30Mazino0 /btc-15min-6monthstabular10K<n<100K0 likes28 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.