CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01happyhackingspace /dit12 likes418k downloads8mo agoHugging Face02rugds /ditec-wdn-- Dataset Card for DiTEC-WDN Dataset Summary DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics. Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series. A node can represent a reservoir, junction, or tank, while an edge… See the full description on the dataset page: https://huggingface.co/datasets/rugds/ditec-wdn.tabulargraph-ml1B<n<10B4 likes24k downloads10mo agoHugging Face03Dihaw00 /ditec-wdn-- Dataset Card for DiTEC-WDN Dataset Summary DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics. Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series. A node can represent a reservoir, junction, or tank, while an edge… See the full description on the dataset page: https://huggingface.co/datasets/Dihaw00/ditec-wdn.tabulargraph-ml1B<n<10B0 likes11k downloads8mo agoHugging Face04hulaba /ditec-wdn-- Dataset Card for DiTEC-WDN Dataset Summary DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics. Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series. A node can represent a reservoir, junction, or tank, while an edge… See the full description on the dataset page: https://huggingface.co/datasets/hulaba/ditec-wdn.tabulargraph-ml1B<n<10B0 likes6.4k downloads9mo agoHugging Face05QingyanBai /Ditto-1M Ditto-1M: A High-Quality Synthetic Dataset for Instruction-Based Video Editing Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset Qingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu, Hanlin Wang, Wen Wang, Ka Leong Cheng, Shuailei Ma, Yanhong Zeng, Zichen Liu, Yinghao Xu, Yujun Shen, Qifeng Chen Figure: Our proposed synthetic data generation pipeline can automatically produce high-quality and highly diverse video editing data, encompassing both… See the full description on the dataset page: https://huggingface.co/datasets/QingyanBai/Ditto-1M.video-to-videon>1T52 likes5.7k downloads4mo agoHugging Face06jattcoder00 /ditec-wdn-- Dataset Card for DiTEC-WDN Dataset Summary DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics. Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series. A node can represent a reservoir, junction, or tank, while an edge… See the full description on the dataset page: https://huggingface.co/datasets/jattcoder00/ditec-wdn.tabulargraph-ml100M<n<1B0 likes5.6k downloads10mo agoHugging Face07lioooox /DiTFakeHere is the released dataset (DiTFake) for Synthetic Image Detection (SID) proposed in our paper. Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective This dataset contains 30,000 images in total, including synthetic images generated by three recent DiT-based models (Flux, PixArt, and SD3) and equal numbers of real images from COCO. More implementation details can be found in our GitHub repository. imageimage-classification1K<n<10K0 likes3.6k downloads1y agoHugging Face08razaali10 /ditec-wdn-- Dataset Card for DiTEC-WDN Dataset Summary DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics. Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series. A node can represent a reservoir, junction, or tank, while an edge… See the full description on the dataset page: https://huggingface.co/datasets/razaali10/ditec-wdn.tabulargraph-ml1B<n<10B0 likes1.6k downloads8mo agoHugging Face09Jnaranjo /video-dit-latents-hq Video DiT Latents - Animals (HQ) Pre-computed VAE latents for training video generation models. Dataset Info Property Value Resolution 256×256 pixels Latent Shape (4, 16, 32, 32) Frames 16 @ 8fps (2 seconds) VAE stabilityai/sd-vae-ft-mse Classes dog, cat, bird, horse, fish, lion, elephant, monkey, butterfly, deer Usage import torch from pathlib import Path # Load a single latent latent = torch.load("dog/12345.pt")… See the full description on the dataset page: https://huggingface.co/datasets/Jnaranjo/video-dit-latents-hq.text10K<n<100K0 likes1.5k downloads9mo agoHugging Face10andyx10 /dit-loras-interpreting andyx10/dit-loras-interpreting Experimenting with interpreting write vectors over 100 hidden-topic model organisms fromdiff-interpretation-tuning/loras implementation We use 'self_attn.o_proj andmlp.down_proj` for write vectors: two per block across 36 blocks, with a total of 72 write vectors per organism. Jacobian Lens from Neuronpedia neuronpedia/jacobian-lens (qwen3-4b/jlens/Salesforce-wikitext/Qwen3-4B_jacobian_lens.pt) layout test100/… See the full description on the dataset page: https://huggingface.co/datasets/andyx10/dit-loras-interpreting.tabular1K<n<10K0 likes1.3k downloads2d agoHugging Face11Darknsu /Ditto_emo_checkpoint_loss_2000 likes504 downloads6mo agoHugging Face12oyyggbond /ditto_local0 likes469 downloads7mo agoHugging Face13kingsidharth /zangei-dit-stage-1-250k-256px-dinov3 Zangei 256px DINOv3 Features — Repacked Repacked from the known Colab extraction layout. Source dataset: kingsidharth/zangei-dit-stage-1-250k-256px-img Feature model: DINOv3 ViT-S/16 Image variants: resized square Feature kinds: cls reg patch Files: shards/dinov3_vits16_256_.safetensors Matching row index: shards/dinov3_vits16_256_.parquet Tensor keys: cls reg patch Manifests: manifests/files.parquet manifests/files.csv config.json tabular1M<n<10M0 likes344 downloads4mo agoHugging Face14oyyggbond /ditto_local_ref0 likes329 downloads7mo agoHugging Face15oyyggbond /ditto_global_freeform30 likes301 downloads7mo agoHugging Face16Jouesmak /DiTFakeHere is the released dataset (DiTFake) for Synthetic Image Detection (SID) proposed in our paper. Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective This dataset contains 30,000 images in total, including synthetic images generated by three recent DiT-based models (Flux, PixArt, and SD3) and equal numbers of real images from COCO. More implementation details can be found in our GitHub repository. imageimage-classification10K<n<100K0 likes274 downloads8mo agoHugging Face17raynardj /ditto-source-videosFrom this video editing dataset all source video, but in parquet text100K<n<1M0 likes237 downloads4mo agoHugging Face18FreedomIntelligence /DitingBench Diting Benchmark Our paperGithub Our benchmark is designed to evaluate the speech comprehension capabilities of Speech LLMs. We tested both humans and Speech LLMs in terms of speech understanding and provided further analysis of the results, along with a comparative study between the two. This offers insights for the future development of Speech LLMs. For more details, please refer to our paper. Result Level Task Human Baseline GPT-4o MuLLaMA GAMA SALMONN… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/DitingBench.audio10K<n<100K1 likes217 downloads1y agoHugging Face19Tsu7am1 /Cinematic-DiT-Video-Dataset Cinematic DiT Video Dataset This dataset is publicly downloadable under a restricted research license. It is intended for non-commercial research on AI-generated video detection, media forensics, authenticity analysis, and content-safety evaluation. Dataset Summary This dataset contains 9,000 synthetic text-to-video samples generated from 3,000 Chinese cinematic prompts. Each source prompt has one video in each of three generation profiles. The prompt pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Tsu7am1/Cinematic-DiT-Video-Dataset.tabular1K<n<10K0 likes214 downloads2mo agoHugging Face20Ouzhang /low-high-ditto-codebook-analysis0 likes203 downloads3mo agoHugging Face21JOHNNY2026sadf /MOB-DiT-data MOB DiT multimodal data Private transfer repository for mouse olfactory-bulb multimodal spatial data. Layout raw/MOB_722/: complete copy of /home/songzq/SMulRe/data/MOB_722. processed/code_MOB_23/: the eight input artifacts used by DiT_224_Multimodal_MOB.py and DiT_MOB_CellInfo_vali_final.py. The processed inputs contain Visium and ISS/Xenium-style expression objects, aligned transcript and cell coordinates, morphology fields, gene embeddings, cluster-affinity… See the full description on the dataset page: https://huggingface.co/datasets/JOHNNY2026sadf/MOB-DiT-data.imagen<1K0 likes195 downloads8d agoHugging Face22aixk /dit-latents-cache30 likes125 downloads5d agoHugging Face23Darknsu /Ditto_emo_checkpoint_2000 likes119 downloads6mo agoHugging Face24Chocopy /PtychoFlow_DiT PtyRAD reconstruction eval on generated test split This folder contains PtyRAD reconstructions for a seeded random subset of test samples pooled from multiple HDF5 files. Dataset files simulation_data1.hdf5: /gpfs/scratch/ailab/ai4physic/gendata3/simulation_data1.hdf5 simulation_data2.hdf5: /gpfs/scratch/ailab/ai4physic/gendata3/simulation_data2.hdf5 simulation_data3.hdf5: /gpfs/scratch/ailab/ai4physic/gendata3/simulation_data3.hdf5 simulation_data4.hdf5:… See the full description on the dataset page: https://huggingface.co/datasets/Chocopy/PtychoFlow_DiT.image10K<n<100K0 likes98 downloads3mo agoHugging Face25BharathK333 /MMFace-DiT-Datasets MMFace-DiT Dataset: Multimodal Face Generation Benchmarks This repository contains the multimodal conditioning data and high-quality captions for MMFace-DiT, accepted to CVPR 2026. This dataset provides the necessary spatial (masks, sketches) and semantic (VLM-enriched captions) pairs to enable high-fidelity, controllable face synthesis. 📂 Dataset Components The dataset is organized to be plug-and-play with the MMFace-DiT repository: Celeb_Dataset/:… See the full description on the dataset page: https://huggingface.co/datasets/BharathK333/MMFace-DiT-Datasets.image-to-image1 likes94 downloads4mo agoHugging Face26danielsanjosepro /ditflow_drawer_vision_only_v1_evalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 26, "total_frames": 11562, "total_tasks": 1, "total_videos": 78, "total_chunks": 1, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:26"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/danielsanjosepro/ditflow_drawer_vision_only_v1_eval.tabularrobotics10K<n<100K0 likes87 downloads9mo agoHugging Face27KonstantinosKK /reflect-dit-train-images Reflect-DiT Training Images This is the official dataset repository for the training images used in Reflect-DiT, a Reflective Diffusion Transformer for image generation. 🔗 Paper: Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection Contents The dataset is stored in multiple .tar archives located in the data/ directory: data/ ├── gen_eval_sana_part_0.tar ├── gen_eval_sana_part_1.tar ├── ... └── gen_eval_sana_part_9.tar… See the full description on the dataset page: https://huggingface.co/datasets/KonstantinosKK/reflect-dit-train-images.text-to-image0 likes86 downloads1y agoHugging Face28zhiyang1 /scatter_dit0 likes82 downloads2y agoHugging Face29jacklishufan /reflect-DiT1 likes79 downloads2y agoHugging Face30seriintan /rollout_eval_multi_task_dit_frazier_v2_20260907_124913This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/seriintan/rollout_eval_multi_task_dit_frazier_v2_20260907_124913.tabularrobotics10K<n<100K0 likes76 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.