CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aharley /pointodysseyFeb 21: updated to v1.2. Please see https://github.com/y-zheng18/point_odyssey/tree/main for release notes. 4 likes7.1k downloads3y agoHugging Face02BangumiBase /aharensanwahakarenaiseason2 Bangumi Image Base of Aharen-san Wa Hakarenai Season 2 This is the image base of bangumi Aharen-san wa Hakarenai Season 2, we detected 65 characters, 6341 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/aharensanwahakarenaiseason2.image1K<n<10K0 likes3.8k downloads1y agoHugging Face03ahashanahmed /csv0 likes2.7k downloads4m agoHugging Face04Ahad690 /fiverr-gigs Fiverr Gigs — community dataset Public Fiverr listing metadata collected via the fiverr-gig-optimizer Claude Code skill. Grown via opt-in contributions from users who run contribute.py. Licensed CC-BY-4.0. What's in it Each record follows the canonical schema (title, category/subcategory, tier prices, delivery days, tags, rating, review_count, gig_count_in_search, currency). Prices are normalized to USD. Privacy Contributions are anonymized before… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/fiverr-gigs.tabular10K<n<100K1 likes1.5k downloads19h agoHugging Face05Ahad690 /growthkit-trends GrowthKit Trends A community, opt-in, federated dataset of public, anonymized short-form short-form-video trend and benchmark observations, contributed by users of the open-source GrowthKit Claude Code skill. It improves GrowthKit's default benchmarks over time so every founder starts from better, source-tagged ranges instead of fabricated numbers. Honesty first. GrowthKit never lets a model invent a market metric. Numbers come from deterministic scripts run on a founder's own… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/growthkit-trends.tabular10K<n<100K0 likes1.4k downloads18h agoHugging Face06aharley /rvl_cdipThe RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class. There are 320,000 training images, 40,000 validation images, and 40,000 test images.image-classification100K<n<1M92 likes1.4k downloads2y agoHugging Face07aharoon /fpp-ml-bench FPP-ML-Bench: Fringe Projection Profilometry Benchmarking Dataset The first open-source, photorealistic synthetic dataset for single-shot fringe projection profilometry (FPP), generated using VIRTUS-FPP in NVIDIA Isaac Sim. This dataset enables standardized benchmarking and systematic comparison of deep learning approaches for single-shot 3D depth reconstruction from fringe patterns. Dataset Summary Property Value Total fringe images 15,600 (52 per viewpoint… See the full description on the dataset page: https://huggingface.co/datasets/aharoon/fpp-ml-bench.imagedepth-estimation1K<n<10K3 likes1.1k downloads8mo agoHugging Face08ahanjaya /crossed_arm_point_clouds Crossed Arm Point Clouds Dataset This dataset contains 3D point cloud data captured from a LiDAR scanner for crossed arm classification in the context of robot magic trick performance. Overview This dataset was collected for training and evaluating the Crossed Arm Voxel Network (CAVN) architecture, a deep learning model designed for 3D point cloud classification in human-robot interaction magic performances. The data supports classification of human arm positions during… See the full description on the dataset page: https://huggingface.co/datasets/ahanjaya/crossed_arm_point_clouds.image-classification1K<n<10K0 likes1k downloads8mo agoHugging Face09Ahad690 /app-rank-anchors App Rank Anchors Community-federated public app-store calibration anchors for the AppScope open app-intelligence stack. Each row is a public fact — a segment + rank + observed download flow — derived from the public Google Play realInstalls delta over a time window paired with an app's chart rank in that window. Pooling these anchors across self-hosting contributors lets the Garg–Telang download estimator calibrate absolute scale (scale_b) per (platform, category, country)… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/app-rank-anchors.n<1K0 likes960 downloads19h agoHugging Face10Ahaskar04 /aloha-handover-no-mocap ALOHA 2-Arm Handover (no mocap markers) 163 successful scripted demonstrations of a bimanual box handover in MuJoCo, recorded on the ALOHA / ViperX 300 dual-arm platform. Arm A picks a box off the table and hands it to arm B, which retracts with it. Why "no mocap" This is a re-collection of Ahaskar04/aloha-handover-data with a visual confound removed. The scripted expert drives the arms through MuJoCo mocap bodies (left/target, right/target) — 2 cm spheres that… See the full description on the dataset page: https://huggingface.co/datasets/Ahaskar04/aloha-handover-no-mocap.robotics10K<n<100K0 likes617 downloads1mo agoHugging Face11aharley /alltracker_dataThis repo contains the data we produced/postprocessed as part of AllTracker: Efficient Dense Point Tracking at High Resolution. This data is used by the training scripts in our github repo, and leads to the models in our model page, which you can test in our Gradio demo. 0 likes539 downloads1y agoHugging Face12Ahaskar04 /viperx-3arm-handover-demo COLA 3-Arm Sequential Handover — Scripted Demonstrations 1,500 simulated demonstrations of a sequential 3-arm handover task in MuJoCo, generated with a hand-tuned waypoint policy. Three ViperX 300s robot arms (A → B → C) cooperate to lift an object off a table, hand it between two arms, and place it into a box rigidly attached to the third arm. This dataset is intended for multi-agent imitation learning, coordination research, and as the 3-arm extension of the 2-arm… See the full description on the dataset page: https://huggingface.co/datasets/Ahaskar04/viperx-3arm-handover-demo.videorobotics1K<n<10K0 likes400 downloads5mo agoHugging Face13BangumiBase /aharensanwahakarenai Bangumi Image Base of Aharen-san Wa Hakarenai This is the image base of bangumi Aharen-san wa Hakarenai, we detected 43 characters, 5875 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/aharensanwahakarenai.image1K<n<10K0 likes256 downloads2y agoHugging Face14ahancock516 /oai-5g-srs-ranging-dataset OAI 5G NR SRS Ranging Captures Uplink Sounding Reference Signal (SRS) channel-estimate captures from a monolithic 5G NR software-defined-radio testbed, collected for SRS-based ranging experiments. The gNB (OpenAirInterface on a USRP X410, Band n78, 40 MHz / 106 PRB) configures each UE to transmit SRS; the gNB's per-SRS frequency-domain channel estimate, oversampled IDFT CIR, and ToA estimate are streamed off the PHY via OAI's T_tracer and recorded at a series of known… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-srs-ranging-dataset.audion<1K0 likes239 downloads2mo agoHugging Face15Jill111 /AHA-WAM-SO101-HIL-training-assets AHA-WAM SO101 Plug HIL Training Assets Reproducibility bundle for the SO101 power-adapter insertion experiments. It contains a compact, ZIP-based representation of the paths expected by the AHA-WAM training configuration: the released AHA-WAM-pretrained.pt initialization checkpoint; so101_ahawam_plug.zip (the downloader extracts only task 02 and 04); ahawam_hil_raw.zip (raw rich-v1/v2/v3 HG-DAgger recordings); the fixed 02+04 action/state normalization statistics; cached T5… See the full description on the dataset page: https://huggingface.co/datasets/Jill111/AHA-WAM-SO101-HIL-training-assets.imagen<1K0 likes237 downloads22d agoHugging Face16Ahaskar04 /cola-handover-demosvideo1K<n<10K0 likes229 downloads5mo agoHugging Face17aHapBean /Wan21-1.3B-2000prompt0 likes227 downloads5mo agoHugging Face18Ahaskar04 /aloha-handover-data0 likes207 downloads2mo agoHugging Face19dsa1dsa12 /deepscaler-verl-aha-momenttext10K<n<100K0 likes200 downloads27d agoHugging Face20QCRI /AHA-MEMES AHA-Memes A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes Hateful memes carry their meaning in the interaction between an image and the text laid over it, and often through cultural references that neither modality states outright. Arabic has been badly served here: the meme resources that exist annotate propaganda or coarse "harmful content", not who is being attacked or how. AHA-Memes is a benchmark of 5,000 Arabic memes, each annotated by trained… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AHA-MEMES.imageimage-classification10K<n<100K0 likes188 downloads24d agoHugging Face21Aharneish /spirittext100K<n<1M0 likes156 downloads3y agoHugging Face22ahancock516 /oai-5g-plkg-dataset OAI 5G PLKG Paired Dataset PUSCH (uplink) and PDSCH (downlink) channel-estimate captures from an OAI 5G NR SA testbed (USRP X410 gNB + USRP B210 UE, band n78, 106 PRB, PCI=1), collected for a channel-informed neural PHY secret-key generation (PLKG) demo paper. Campaigns 1-3: captures are handshake/RRC-setup-phase only (strict DMRS-comb filtering rejects connected-mode single-layer data on both plugins). Campaign 4 relaxes this filtering on both plugins and includes ordinary… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-plkg-dataset.videon<1K0 likes154 downloads2mo agoHugging Face23safafaf311 /aha-moment-dataset verl_deepscaler — DeepScaleR for Verl GRPO Training Full DeepScaleR training dataset (from agentica-org/DeepScaleR-Preview-Dataset, the most downloaded DeepScaleR dataset with 25.4K downloads), converted to Verl framework format, with the DeepSeek-R1 "aha moment" problem included. Source Original: agentica-org/DeepScaleR-Preview-Dataset Size: 40,315 math problems from AIME, AMC, Omni-MATH, STILL Included: +1 "aha moment" problem (the exact problem from… See the full description on the dataset page: https://huggingface.co/datasets/safafaf311/aha-moment-dataset.textn<1K0 likes149 downloads28d agoHugging Face24ahazimeh /slide2svg Slide2SVG Slide2SVG was released as part of the paper “Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling” (AAAI 2026). Slide2SVG is a real-world dataset for semantic document derendering: transforming a rasterized presentation slide into a structured, editable SVG representation. Curated from publicly available academic conference presentations, it captures diverse design styles, font choices, image placements, and layout configurations found in real… See the full description on the dataset page: https://huggingface.co/datasets/ahazimeh/slide2svg.imageimage-to-text10K<n<100K11 likes145 downloads4mo agoHugging Face25ahad-j /robocasa_mt4_N216_zarr0 likes133 downloads5mo agoHugging Face26sdfffafag3 /countdown-verl-aha-moment Countdown Dataset for Verl - DeepSeek-R1 "Aha Moment" Reproduction This dataset is prepared for reproducing DeepSeek-R1's "aha moment" using the Verl framework. It uses the Countdown task from Jiayi-Pan/Countdown-Tasks-3to4. Overview The "aha moment" was first observed in DeepSeek-R1-Zero when training on the Countdown task — the model learns to allocate more thinking time and re-evaluate its initial approach rather than rushing to an answer. Files… See the full description on the dataset page: https://huggingface.co/datasets/sdfffafag3/countdown-verl-aha-moment.text100K<n<1M0 likes127 downloads28d agoHugging Face27Ahaskar04 /cola-pushblock-1033 COLA Push-Block — Two-Agent Coordination (1033 demos) Teleoperated demonstrations of a two-arm cooperative push-block task in MuJoCo, collected for the COLA (Coordination via Latent Adapters) paper. Two 4-DoF arms on opposite ends of a near-frictionless ("icy") table cooperate to (1) push a small block across the midline and (2) stop it inside a target zone before it slides off the far edge. This is the cached, flattened version of the dataset used directly by the COLA training… See the full description on the dataset page: https://huggingface.co/datasets/Ahaskar04/cola-pushblock-1033.robotics100K<n<1M0 likes126 downloads4mo agoHugging Face28asfafaaf3434 /deepscaler-aha-momenttext10K<n<100K0 likes123 downloads27d agoHugging Face29ahamedddd /svarah_processedThe input_features are nothing but the values generated after passing the dataset's audio array through a whisper processor's feature extraction and the field 'labels' consists of the tokenized(using whisper tokenizer) ground truths. The following is the link for what I did with the sarvah dataset and how I trained it on whisper-large-v3-turbo. The training steps for whisper-large-v3 are same. https://colab.research.google.com/drive/1oD0v7MWZ9WJqk7tZYThwgTUM85PTEhMN?usp=sharing 1K<n<10K0 likes119 downloads1y agoHugging Face30ACIDE /AHA-Calvin-1pimage10K<n<100K0 likes119 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.