CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sponbebob4258 /recam-lerobotgatedvideo10K<n<100K0 likes8.9k downloads10d agoHugging Face02Ar4ikov /celebA_spoof Dataset Card for "celebA_spoof" More Information needed image100K<n<1M1 likes6.3k downloads3y agoHugging Face03mypThu /SportsSlomo-CVS 🎥 SportsSloMo-CVS Dataset This repository contains the dataset presented in the paper Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor. The Complementary Vision Sensor (CVS), known as Tianmouc, captures synchronized RGB frames together with high-frame-rate, multi-bit spatial difference (SD, encoding structural edges) and temporal difference (TD, encoding motion cues) data within a single RGB exposure. This dataset facilitates research in RGB… See the full description on the dataset page: https://huggingface.co/datasets/mypThu/SportsSlomo-CVS.textimage-to-imagen<1K4 likes5.7k downloads4mo agoHugging Face04slayone /uva_spoj_rawtranslation1K<n<10K1 likes5.2k downloads3y agoHugging Face05DuplexGen /duplexgen-spoken DuplexGen Spoken Rendered spoken audio for the DuplexGen turn-taking dialogues — the exact set of clips used to fine-tune the full-duplex model (PP-DG) in DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking Dialogues. Each clip is a full render of one generated dialogue variation: the mixed two-speaker dialogue audio, the isolated per-turn utterances, and the inserted backchannel clips, plus per-clip metadata. Audio is synthesized with Chatterbox TTS; the dialogue text it… See the full description on the dataset page: https://huggingface.co/datasets/DuplexGen/duplexgen-spoken.text-to-speech10K<n<100K4 likes4.7k downloads2mo agoHugging Face06maharshipandya /spotify-tracks-dataset Content This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly. Usage The dataset can be used for: Building a Recommendation System based on some user input or preference Classification purposes based on audio features and available genres Any other application that you can think of. Feel free to discuss! Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.tabularfeature-extraction100K<n<1M136 likes4.6k downloads3y agoHugging Face07karunapu /sports-calendar Sports Calendar Feeds Public calendar and scoreboard artifacts for MLB, NHL, NFL, College Football (NCAA Division I), the English Premier League, and selected men's cricket competitions. The feeds update on a deterministic four-hour GitHub Actions schedule. Calendar events are transparent (Free), use stable first-party UIDs, update completed games with final scores, and mark called-off fixtures as cancelled. For normal calendar use, choose the Daily Scoreboard feed. It condenses… See the full description on the dataset page: https://huggingface.co/datasets/karunapu/sports-calendar.0 likes4.5k downloads20m agoHugging Face08BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes4.5k downloads9mo agoHugging Face09ruslanmv /sports-trends-dataset ⚽🏀🎾🏏 Sports-Trends Dataset A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits. The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day. TL;DR — A continuously-updated, medallion-architecture data lake for football, basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.tabulartabular-classificationn<1K5 likes4.4k downloads5h agoHugging Face10Lekim89 /sportsmot SportsMOT SportsMOT is a large-scale dataset for single-camera multi-object tracking (MOT) in sports videos. It focuses on tracking players in professional sports scenes, where targets exhibit fast and variable-speed motion, frequent occlusion, motion blur, camera motion, and similar team uniforms. This Hugging Face repository provides SportsMOT in a MOTChallenge-style structure for research, benchmarking, training, and evaluation of multi-object tracking systems in sports… See the full description on the dataset page: https://huggingface.co/datasets/Lekim89/sportsmot.imageobject-detection100K<n<1M2 likes2.9k downloads4mo agoHugging Face11Torch-Trade /btcusdt_spot_1m_03_2023_to_12_2025tabular1M<n<10M0 likes2.5k downloads8mo agoHugging Face12MLCommons /ml_spoken_wordsMultilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.0. The dataset contains more than 340,000 keywords, totaling 23.4 million 1-second spoken examples (over 6,000 hours). The dataset has many use cases, ranging from voice-enabled consumer devices to call center automation. This dataset is generated by applying forced alignment on crowd-sourced sentence-level audio to produce per-word timing estimates for extraction. All alignments are included in the dataset.audio-classification10M<n<100M37 likes2.2k downloads4y agoHugging Face13spongy /physical-ai-bench-step-30000-videos PhysicalAIBench Step 30000 Model Comparison This repository contains two filename-aligned sets of 5,220 MP4 outputs from step_30000 evaluation runs. The files are presented through the companion PhysicalAI Video Gallery. Model sets Directory Model Files Bytes videos/ DC-AE v0.2, Cosmos encoder + causal decoder 5,220 4,244,965,541 videos_wan22_vae/ Wan 2.2 VAE, phase 3 PDX 5,220 4,818,900,749 The two directories have an exact 1:1 basename match.… See the full description on the dataset page: https://huggingface.co/datasets/spongy/physical-ai-bench-step-30000-videos.video10K<n<100K0 likes1.9k downloads2mo agoHugging Face14spongy /tempimagen<1K0 likes1.9k downloads4mo agoHugging Face15lerobot-raw /spoc_raw0 likes1.9k downloads2y agoHugging Face16YOLO2431 /kitchen_rack_combo_v2_spoon_onlyThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "yam_bimanual", "total_episodes": 183, "total_frames": 75265, "total_tasks": 1, "total_videos": 549, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:183" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/kitchen_rack_combo_v2_spoon_only.tabularrobotics10K<n<100K0 likes1.6k downloads2mo agoHugging Face17BEE-spoke-data /wikipedia-20230901.en-deduped wikipedia - 20230901.en - deduped purpose: train with less data while maintaining (most) of the quality This is really more of a "high quality diverse sample" rather than "we are trying to remove literal duplicate documents". Source dataset: graelo/wikipedia. configs default command: python -m text_dedup.minhash \ --path $ds_name \ --name $dataset_config \ --split $data_split \ --cache_dir "./cache" \ --output $out_dir \ --column $text_column \… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/wikipedia-20230901.en-deduped.texttext-generation10M<n<100M6 likes1.4k downloads9mo agoHugging Face18SpongeCakeAllDay /cad-bench-freecad-agent-initial-runs0 likes1.4k downloads4mo agoHugging Face19RoboCOIN /AI2_Alphabot_2_stir_spoon AI2_Alphabot_2_stir_spoon Dataset Description This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Task Preview View Video Directly Overview Total Episodes: 405 Total Frames: 334844 FPS: 30 Dataset Size: 14.30 GB Robot Name: AI2_Alphabot_2 End-Effector Type: two_finger_end_effector Teleoperation Type: vr_controller Sensors: cam_front_chest_rgb, cam_front_head_rgb, cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_stir_spoon.robotics0 likes1.4k downloads3mo agoHugging Face20BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.3k downloads9mo agoHugging Face21hkadxqq /spooky-author-identificationtext10K<n<100K0 likes1.2k downloads4y agoHugging Face22BEE-spoke-data /Long-Data-Col-rp_pile_pretrain Dataset Card for "Long-Data-Col-rp_pile_pretrain" This dataset is a subset of togethercomputer/Long-Data-Collections, namely the rp_sub.jsonl.zst and pile_sub.jsonl.zst files from the pretrain split. Like the source dataset, we do not attempt to modify/change licenses of underlying data. Refer to the source dataset (and its source datasets) for details. changes as this is supposed to be a "long text dataset", we drop all rows where text contains <= 250 characters.… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/Long-Data-Col-rp_pile_pretrain.texttext-generation10M<n<100M3 likes1.2k downloads9mo agoHugging Face23AxonData /face-anti-spoofing-dataset Face Antispoofing dataset for liveness detection Anti-Spoofing dataset: live, replay, cut, print, 3D masks - large-scale face anti spoofing This dataset delivers a single, end-to-end resource for training and benchmarking facial liveness-detection systems. By aggregating live sessions and eleven realistic presentation-attack classes into one collection, it accelerates development toward iBeta Level 1/2 compliance and strengthens model robustness against the full spectrum of spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/face-anti-spoofing-dataset.imagevideo-classificationn<1K2 likes1.1k downloads7mo agoHugging Face24ssz1111 /SpokenWOZ-Test-Audioaudio1K<n<10K1 likes1k downloads9mo agoHugging Face25UniverseTBD /mmu_tess_spoc mmu_tess_spoc HATS Catalog Collection This is the collection of HATS catalogs representing mmu_tess_spoc. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_tess_spoc.tabular1M<n<10M0 likes1k downloads4mo agoHugging Face26Jack-ppkdczgx /SEA-Spoofgated SEA-Spoof Access And License Access requires author approval. Please email the authors before requesting or using the dataset: wu_jinyang@a-star.edu.sg imcc.sg@gmail.com This dataset is released for non-commercial academic research only. Use is restricted to academic institutions and approved research users. Commercial use is not permitted, and this dataset may not be used by commercial companies or for commercial products, services, model training, evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Jack-ppkdczgx/SEA-Spoof.audioaudio-classification100K<n<1M6 likes989 downloads2mo agoHugging Face27ConquestAce /spotify-songs5 likes941 downloads2y agoHugging Face28BEE-spoke-data /consumer-finance-complaints BEE-spoke-data/consumer-finance-complaints consumer-finance-complaints but in a format that actually works. Pulled Feb 2024 texttext-classification1M<n<10M5 likes896 downloads9mo agoHugging Face29TommyKwok /binance-top50-spot-v1 Binance Top 50 Backtesting Dataset Built at: 2026-05-21T11:53:45.468482+00:00 Parameters Lookback: 1 days Top N: 3 Trade Types: spot, um Data Types: klines, aggTrades Build Status SPOT: 3 symbols klines: 3/3 healthy aggTrades: 3/3 healthy UM: 3 symbols klines: 3/3 healthy aggTrades: 3/3 healthy fundingRate: 3/3 healthy textrobotics100K<n<1M0 likes879 downloads4mo agoHugging Face30Ustiniansy /SportsTimegated SportsTime SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026. It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball. Dataset This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.textvisual-question-answering10K<n<100K1 likes876 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.