CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes1.9k downloads22d agoHugging Face02SBMM75 /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes868 downloads10d agoHugging Face03b-remy /jade-samples-10000x1 JADE amortized posterior samples — 10,000 observations x 1 draw Noisy weak-lensing convergence observations paired with joint posterior draws of (convergence field, cosmology) from the amortized conditional diffusion model of JADE. [!IMPORTANT] This dataset is not a product of arXiv:2606.31988. It was generated afterwards, with the same trained model, to support posterior calibration diagnostics that do not appear in the paper. No number in the paper was computed from it, and… See the full description on the dataset page: https://huggingface.co/datasets/b-remy/jade-samples-10000x1.tabular10K<n<100K0 likes681 downloads12d agoHugging Face04Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes480 downloads1y agoHugging Face05stablellama /Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source The images were created in ComfyUI with the bf16 version of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.tabulartext-to-image1K<n<10K0 likes351 downloads22d agoHugging Face06dboemer /koch_50-samplesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "koch", "total_episodes": 50, "total_frames": 16417, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dboemer/koch_50-samples.tabularrobotics10K<n<100K0 likes299 downloads2y agoHugging Face07jozhang97 /ambient-short-samplestabular1K<n<10K0 likes279 downloads1y agoHugging Face08NinaCalvi /ultra-50k-samples-dataset-instruction_followingtabular10K<n<100K0 likes268 downloads2y agoHugging Face09sunovivid /sit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199tabular100K<n<1M0 likes239 downloads10mo agoHugging Face10data-is-better-together /fineweb2-2k-samplestabular100K<n<1M0 likes238 downloads1y agoHugging Face11kdcyberdude /cosmopedia_web_samples_v2_shards_entabular1M<n<10M0 likes213 downloads2y agoHugging Face12enjalot /latent-taxonomy-samplestabular1M<n<10M0 likes166 downloads2mo agoHugging Face13PleIAs /data_samples Multimodal Pretraining This section covers our large-scale collections at the source and is distributed in its original form (PDF with layout intact, audio attached to its transcript) rather than as text extracted after the fact. The emphasis is on what large-scale web collection misses: academic global production (badly indexed in scientific repositories); patents outside the US; the technical and regulatory archives of telecom and finance. These are long, structured… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/data_samples.tabular1M<n<10M0 likes158 downloads1mo agoHugging Face14cfahlgren1 /hermes-agent-trace-samples-2026-06-05 Hermes Agent Raw Session Samples Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers. Each file in sessions/ is the exact single-session output from: hermes sessions export sessions/<session_id>.jsonl --session-id <session_id> No derived tables, flattened rows, SQLite database, or formatted JSON copies are included. tabularn<1K0 likes142 downloads4mo agoHugging Face15gmreincglm /usta-feeds-samples Dated samples of United States public-record change files 15 samples, one folder per family. Each folder holds sample.csv, sample.json and a README naming the source, the columns, the sealing date and the row count. Every file is a change file, not a snapshot. We seal dated copies of a public source, compare two copies, and keep what appeared, what stopped being listed, and what quietly changed in between. Most of these sources publish only the list as it stands today and… See the full description on the dataset page: https://huggingface.co/datasets/gmreincglm/usta-feeds-samples.tabularn<1K0 likes137 downloads14d agoHugging Face16sunovivid /sit-latents-ode-heun-1000-class-0_1000-samples-segment-400-499tabular100K<n<1M0 likes129 downloads10mo agoHugging Face17lerobot /SMPL_samples SMPL motion dataset Episodes Episode Motion 0 dance_hiphop 1 dance_jackson 2 eating_apple 3 high_jump 4 hiphop_ii 5 horse_riding 6 jump_360 7 mambo_kicks 8 reach_jump 9 turn_jump_270 10 turn_jump_360 11 victory_dance 12 walk_forward tabularrobotics1K<n<10K0 likes123 downloads2mo agoHugging Face18fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes114 downloads3d agoHugging Face19superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes112 downloads5mo agoHugging Face20alirezaaminzadeh /docflow-invoice-samples-fa DocFlow Invoice Samples — Persian & Bilingual Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines. Published by Aria AI Engineering Team. Dataset Summary Property Value Samples 50 (synthetic, OCR-friendly) Languages Persian (FA), English (EN) Formats PNG images + JSON annotations Use case Invoice OCR benchmarking, AP automation R&D Synthetic Yes — no real PII Fields Annotated vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.imageimage-to-textn<1K0 likes106 downloads2mo agoHugging Face21chair0 /50_samples_20260713_185031This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/chair0/50_samples_20260713_185031.tabularrobotics1K<n<10K0 likes101 downloads1mo agoHugging Face22fluid-concepts /multimodal-expert-instruction-samplesgated Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside. ▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.audion<1K1 likes98 downloads3d agoHugging Face23Invicto69 /Job-dataset-samples My Jobs and Companies Dataset This dataset contains sample rows of job postings and company profiles. tabular100K<n<1M5 likes95 downloads4mo agoHugging Face24OwnedByDanes /Usenet-Corpus-1980-2013-Threaded-Samples Usenet Corpus 1980–2013 — Threaded (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded dataset: Usenet posts reconstructed into conversations via thread_id, thread_position, and thread_depth. This repo is a free preview; the full, commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at: Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.tabulartext-generation10K<n<100K0 likes91 downloads11d agoHugging Face25dboemer /eval_koch_50-samplesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "koch", "total_episodes": 10, "total_frames": 4979, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dboemer/eval_koch_50-samples.tabularrobotics1K<n<10K0 likes77 downloads2y agoHugging Face26mlfoundations-dev /multiple_samples_ground_truth_openr1_llm_verifiertabular100K<n<1M0 likes73 downloads2y agoHugging Face27mlfoundations-dev /multiple_samples_ground_truth_openr1_llm_verifier_cleantabular100K<n<1M0 likes69 downloads2y agoHugging Face28Inferencebench /pass-at-k-samplestabular1K<n<10K0 likes69 downloads6mo agoHugging Face29NinaCalvi /ultra-50k-samples-dataset-honestytabular10K<n<100K0 likes67 downloads2y agoHugging Face30HCAI-Lab-GT /archive-dolma3-pool-150b-samples archive-dolma3-pool-150b-samples ARCHIVE (pre-6T era): stratified working samples drawn from the 150B pool. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_pool_150B_samples Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory including the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-samples.tabular10M<n<100M0 likes67 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.