CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RefVideo6M /RefVideo6Mgated RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing If you use RefVideo-6M in your research, please cite our work as follows: @article{zi2026refvideo6m title={RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing}, author={Bojia Zi and Xiaoyan Yang and Yu Zhou and Ruijie Sun and Lihan Zhang and Bin Liang and Kam-Fai Wong and Haibin Huang and Chi Zhang and Xuelong Li}, journal={arXiv preprint arXiv:2608.26101}… See the full description on the dataset page: https://huggingface.co/datasets/RefVideo6M/RefVideo6M.video17 likes55k downloads22d agoHugging Face02LucasFang /FLUX-Reason-6M FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems. This dataset contains: 6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model. 20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.image1M<n<10M109 likes9.4k downloads8mo agoHugging Face03Super-shuhe /FaceID-6M FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset This repository contains the dataset described in FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset. Links FaceID-6M: A Large-Scale, Open-Source FaceID Customization Dataset Introduction Comparison with Previous Works FaceID Fidelity Scaling Results Released FaceID-6M dataset Released FaceID Customization Models Usage Contact Introduction FaceID-6M, is the first… See the full description on the dataset page: https://huggingface.co/datasets/Super-shuhe/FaceID-6M.text-to-image18 likes2.6k downloads1y agoHugging Face04ccvl /LAION-High-Qualtiy-Pro-6M-VLV Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models LAION-High-Qualtiy-Pro-6M Dataset This repository hosts LAION-High-Quality-Pro-6M, the image-text dataset we used to train Vision-Language-Vision models. Example Usage: # pip install -U datasets pillow from datasets import load_dataset from PIL import Image import base64 import io # Robust decoder: works if the column is base64 *or* raw bytes import io import… See the full description on the dataset page: https://huggingface.co/datasets/ccvl/LAION-High-Qualtiy-Pro-6M-VLV.textimage-to-text1M<n<10M4 likes1.8k downloads1y agoHugging Face05ByteDance-Seed /BM-6M Dataset Card for ByteMorph-6M The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present ByteMorph… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M.image-to-image1M<n<10M14 likes1.6k downloads1y agoHugging Face06prithivMLmods /Demeter-LongCoT-6M Demeter-LongCoT-6M Demeter-LongCoT-6M is a high-quality, compact chain-of-thought reasoning dataset curated for tasks in mathematics, science, and coding. While the dataset spans diverse domains, it is primarily driven by mathematical reasoning, reflecting a major share of math-focused prompts and long-form logical solutions. Quick Start with Hugging Face Datasets🤗 pip install -U datasets from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Demeter-LongCoT-6M.texttext-generation1M<n<10M5 likes1.3k downloads4mo agoHugging Face07overthelex /ua-case-outcome-6m Ukrainian Court Decisions: Case Outcome Prediction (6.7M) The largest publicly available dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (EDRSR). Contains 6,690,284 substantive decisions from civil and commercial courts spanning 2008--2026, with temporal splits across three wartime epochs. Overview Ukraine's EDRSR is one of the world's largest open judicial databases, containing 100M+ judicial… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-case-outcome-6m.texttext-classification1M<n<10M0 likes640 downloads4mo agoHugging Face08ByteMorph /BM-6M Dataset Card for ByteMorph-6M The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present ByteMorph… See the full description on the dataset page: https://huggingface.co/datasets/ByteMorph/BM-6M.image-to-image1M<n<10M0 likes487 downloads1y agoHugging Face09prithivMLmods /Helios-R-6M Helios-R-6M Helios-R-6M is a high-quality, compact reasoning dataset designed to strengthen multi-step problem solving across mathematics, computer science, and scientific inquiry. While the dataset covers a range of disciplines, math constitutes the largest share of examples and drives the reasoning complexity. Quick Start with Hugging Face Datasets🤗 pip install -U datasets from datasets import load_dataset dataset = load_dataset("prithivMLmods/Helios-R-6M"… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Helios-R-6M.texttext-generation1M<n<10M3 likes477 downloads4mo agoHugging Face10JoTeqtheFirstAI /fineweb-edu-dedup6m Stage 1 (S1): General Knowledge Anchor — 6M FineWeb-Edu-Dedup 1. Project Overview This dataset represents the General Knowledge Acquisition Phase (S1) for a research project focused on developing a Domain-Adaptive LLM for ISO 27001 Information Security Auditing. S1 serves as the cognitive foundation. This corpus is designed to establish high-level linguistic proficiency and general reasoning before the introduction of specialized regulatory standards in Stage 2.… See the full description on the dataset page: https://huggingface.co/datasets/JoTeqtheFirstAI/fineweb-edu-dedup6m.text-generation1M<n<10M0 likes412 downloads8mo agoHugging Face11Boese0601 /ByteMorph-6M-Demo Dataset Card for ByteMorph-6M-Demo The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/Boese0601/ByteMorph-6M-Demo.imageimage-to-image100K<n<1M1 likes403 downloads1y agoHugging Face12NanoMatriX /fineweb-edu-dedup6mtext1M<n<10M0 likes313 downloads8mo agoHugging Face13HayatoHongoEveryonesAI /qa_verify_cot_new_6M_unfiltered_v7dataset_names = [ "HayatoHongoEveryonesAI/qa_verify_1m_cot_1", "HayatoHongoEveryonesAI/qa_verify_1m_cot_2", "HayatoHongoEveryonesAI/qa_verify_1m_cot_3", "HayatoHongoEveryonesAI/qa_verify_1m_cot_4", "HayatoHongoEveryonesAI/qa_verify_1m_cot_5", "HayatoHongoEveryonesAI/qa_verify_2m_cot_2", "HayatoHongoEveryonesAI/qa_verify_2m_cot_3", ] https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing tabular1M<n<10M0 likes305 downloads8mo agoHugging Face14NewsDataHub /openai-vs-anthropic-news-coverage-6mo-2025-2026 OpenAI vs Anthropic News Coverage (6 Months) Dataset Summary This dataset contains English-language news articles that mention OpenAI or Anthropic in the article title or description. It is designed for analyzing media coverage volume and trends over time, not sentiment or opinion. The dataset covers approximately six months of news and includes both article-level data and a derived weekly aggregation. Data Collection Articles were collected using keyword-based… See the full description on the dataset page: https://huggingface.co/datasets/NewsDataHub/openai-vs-anthropic-news-coverage-6mo-2025-2026.text-classification1K<n<10K1 likes265 downloads7mo agoHugging Face15ByteDance-Seed /BM-6M-Demo Dataset Card for ByteMorph-6M-Demo The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M-Demo.imageimage-to-image100K<n<1M3 likes244 downloads1y agoHugging Face16TianfuXinqu /filesystem_huggingface_5053_cl6lee6m Support Ticket Triage Corpus Dataset ID: ZorakTriage94b837 Customer support ticket records with priority, status, and satisfaction annotations. textn<1K0 likes228 downloads29d agoHugging Face17AdoCleanCode /AE_english_data_stage_3-6M-6.5Mtext100K<n<1M0 likes183 downloads8mo agoHugging Face18nthakur /cornstack-6-langs-v1-tevatron-6Mtext1M<n<10M0 likes150 downloads1y agoHugging Face19Junc1i /Auto-6ML0 likes135 downloads2mo agoHugging Face20wusize /laion6m_recapimage1M<n<10M0 likes131 downloads1y agoHugging Face21minhbui /spell_6m_mixtext1M<n<10M0 likes112 downloads2y agoHugging Face22HayatoHongoEveryonesAI /qa_verify_cot_new_6M_v6dataset_names = [ "HayatoHongoEveryonesAI/qa_verify_cot_new_5.1M_v7", "HayatoHongoEveryonesAI/qa_verify_new_v6", ] tabular1M<n<10M0 likes97 downloads8mo agoHugging Face23stzhao /MARIO-6MThis datasets is curated by TextDiffuser team in their work: TextDiffuser: Diffusion Models as Text Painters (NeurIPS 2023) MARIO-6M contains 6M images with text rendered on, filtered from LAION-400M. imagetext-to-image1M<n<10M4 likes96 downloads2y agoHugging Face24ByteMorph /BM-6M-Demo Dataset Card for ByteMorph-6M-Demo The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteMorph/BM-6M-Demo.imageimage-to-image100K<n<1M0 likes77 downloads1y agoHugging Face25MarcusLammers /vast-rtx3090-market-6mo Vast.ai RTX 3090 Spot Market, February-August 2026 Panel data from the vast.ai GPU rental marketplace, restricted to NVIDIA RTX 3090 offers. The public offer listing was polled every 10 minutes between 2026-02-13 and 2026-08-15. Each observation records price, hardware specifications, host reliability, and location. A derived lifecycle table gives the listing duration of every offer. Vast.ai does not publish historical listing data; this dataset was collected independently.… See the full description on the dataset page: https://huggingface.co/datasets/MarcusLammers/vast-rtx3090-market-6mo.tabular1M<n<10M0 likes64 downloads1mo agoHugging Face26HuggingFaceTB /cosmopedia_6Mtext1M<n<10M6 likes58 downloads2y agoHugging Face27jasonrichdarmawan /nllb-200-6M-sample-embeddingOriginal dataset SONAR's author message What happens to the original dataset? Filter by blaser_sim >= 3.5 Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline What are the use cases? Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly tabular1M<n<10M0 likes55 downloads1y agoHugging Face28QingyuShi /scaleedit-filtered-6m ScaleEdit Filtered 6M This repository contains a portable selection manifest for high-quality samples from ScaleEdit-12M. It does not redistribute the source images. Download the original ScaleEdit-12M Parquet files separately, then join each manifest row to the source file named by source_relative_path at zero-based row_index. Selection For each evaluated source row, the latest successful stage-2 review was used. A row is included when result.final_decision ==… See the full description on the dataset page: https://huggingface.co/datasets/QingyuShi/scaleedit-filtered-6m.tabular1M<n<10M0 likes51 downloads25d agoHugging Face29JeshmaSoph /Collection_1_6Mayvideon<1K0 likes50 downloads5mo agoHugging Face30Kausp11 /wiki6m-selfdoc-final_v2 wiki6m-selfdoc-final_v2 GJ's cleaned-context wiki6M blocks (V4 arm search_1p7B_selfdoc_v2): tiered queries, = LLM-written answer from the source doc, no BM25. 640 parquet shards at repo root, 1,562,273 blocks x 4096, 26% masked. meta/: holdout + stats. RCP source: /mloscratch/homes/ponkshe/searchllm_dt_runs/hf_stage/selfdoc_parts_search. Pushed 2026-09-15 by ops/push_datasets_to_hf.py. Provenance: Search-LLM EXPERIMENT_PLAN.md / RESULTS.md / INVENTORY.md. 1M<n<10M0 likes50 downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.