CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dalle-mini /witimage1M<n<10M7 likes140k downloads5y agoHugging Face02adams-story /datacomp200m Datacomp200m This is a smaller version of the datacomp_1b dataset. Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows. The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling. Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.image100M<n<1B3 likes110k downloads3y agoHugging Face03tencent /Hy-Embodied-0.5-VLA-Data Hy-Embodied-0.5-VLA From Vision-Language-Action Models to a Real-World Robot Learning Stack Tencent Robotics X × Tencent Hy Team 📖 Abstract We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.tabularroboticsn<1K23 likes84k downloads3mo agoHugging Face04lmarena-ai /leaderboard-dataset Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load_dataset # Load all historical text style control data ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full") # Load the current text style control leaderboard ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest") # Filter to overall category ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.tabular1M<n<10M26 likes46k downloads14h agoHugging Face05autogluon /fev_datasets Forecast evaluation datasets This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models. The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities. The datasets follow a format that is compatible with the fev package. Data format and usage Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.tabulartime-series-forecasting100K<n<1M13 likes45k downloads8mo agoHugging Face06allenai /molmobot-data MolmoBot-data Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms: DoorOpeningDataGenConfig RBY1OpenDataGenConfig RBY1PickDataGenConfig FrankaPickOmniCamConfig RBY1PickAndPlaceDataGenConfig FrankaPickAndPlaceOmniCamConfig FrankaPickAndPlaceColorOmniCamConfig FrankaPickAndPlaceNextToOmniCamConfig Please note that every package indexed by the parquet files can contain several instances of episode data. We also provide an… See the full description on the dataset page: https://huggingface.co/datasets/allenai/molmobot-data.tabular100K<n<1M8 likes43k downloads2mo agoHugging Face07mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes39k downloads3y agoHugging Face08zekaiwang /trex_dataset T-Rex Dataset A large-scale, tactile-reactive bimanual manipulation dataset, collected via teleoperation on a Dexmate Vega-1 robot with two Sharpa Wave dexterous hands. Stored as a LeRobotDataset v3.0. 🌐 Project Page · ✍️ Paper (arXiv) · 💻 Code (T-Rex) · 🚀 Dataset Quickstart · 📓 Colab notebook One episode from each of 20 motor primitives (head-camera view, cropped to the workspace), each with a different object. Teleoperation setup: Manus gloves + VIVE… See the full description on the dataset page: https://huggingface.co/datasets/zekaiwang/trex_dataset.tabularrobotics1M<n<10M30 likes34k downloads3mo agoHugging Face09autogluon /chronos_datasets Chronos datasets Time series datasets used for training and evaluation of the Chronos forecasting models. Note that some Chronos datasets (ETTh, ETTm, brazilian_cities_temperature and spanish_energy_and_weather) that rely on a custom builder script are available in the companion repo autogluon/chronos_datasets_extra. See the paper for more information. Data format and usage The recommended way to use these datasets is via https://github.com/autogluon/fev. All datasets… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets.tabulartime-series-forecasting10M<n<100M76 likes33k downloads2y agoHugging Face10huggingface-projects /drlc-leaderboard-datatabular10K<n<100K2 likes32k downloads2d agoHugging Face11Forceless /PPTAgent-parsed_dataimage1K<n<10K8 likes31k downloads2y agoHugging Face12wayslab /llm-network-study-data LLM-Network-Study-Data Per-request network captures (.pcapng) collected by the LLM-Network-Study benchmark harness (benchmark.py and the per-workload test scripts). Each directory holds one capture file per request, named request_<id>_run<n>_<timestamp>.pcapng. A directory name encodes four dimensions: <capture-env>_<provider/model>_<workload>[_<dataset/variant>]_results Dimension legend Dimension Values Meaning Capture env ethernet Wired connection to… See the full description on the dataset page: https://huggingface.co/datasets/wayslab/llm-network-study-data.tabularn<1K0 likes30k downloads29d agoHugging Face13open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads58m agoHugging Face14allenai /MolmoAct2-BimanualYAM-DatasetThis dataset was created using LeRobot. MolmoAct2-BimanualYAM Dataset This repository is the merged ckpt / merged LeRobot dataset artifact for the MolmoAct2-BimanualYAM Dataset, a large-scale collection of bimanual robot manipulation demonstrations collected for MolmoAct2. Across the full collection, MolmoAct2-BimanualYAM contains more than 720 hours of training demonstrations spanning diverse tabletop manipulation tasks. Language Annotations This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-BimanualYAM-Dataset.tabularrobotics10M<n<100M11 likes27k downloads4mo agoHugging Face15openfoodfacts /product-database Open Food Facts Database What is 🍊 Open Food Facts? A food products database Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels. Made by everyone Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/product-database.tabular1M<n<10M147 likes23k downloads59m agoHugging Face16facebook /show3d-dataset SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen, Alex Wong, Tomas Hodan, and Kun He CVPR 2026; https://arxiv.org/abs/2603.28760 News September 18, 2026: Released synchronized exocentric views and camera calibrations. SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild. It is intended to advance research on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/show3d-dataset.tabularother1K<n<10K6 likes21k downloads6d agoHugging Face17chewwt /po_qwen14b_tabular_data BoLT Prompt Optimization — Tabular Dataset For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks. Dataset Description The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores. Evaluation details: Model: Qwen/Qwen3-14B Task: minerva_math500 (4-shot) (from lm-eval library) System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabulartext-generation1K<n<10K1 likes21k downloads5mo agoHugging Face18SII-WANGZJ /Polymarket_data Polymarket Data Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze A comprehensive dataset of 1.9 billion trading records from Polymarket, processed into multiple analysis-ready formats. Features cleaned data, unified token perspectives, and user-level transformations — ready for market research, behavioral studies, and quantitative analysis. Zhengjie Wang1,2, Leiyu Chao1,3, Yu Bao1,4, Lian Cheng1,3, Jianhan Liao1,5, Yikang Li1,† 1Shanghai Innovation Institute… See the full description on the dataset page: https://huggingface.co/datasets/SII-WANGZJ/Polymarket_data.tabular1B<n<10B84 likes20k downloads2mo agoHugging Face19huggingface /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/CADS-dataset.tabularimage-segmentation10K<n<100K4 likes18k downloads9mo agoHugging Face20datamatastudios /ai-model-popularity Datamata AI Model Popularity Index Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot. Latest snapshot: 2026-09-20 Models in this release: 50 Updated: weekly Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution. Source & methodology: https://www.datamatastudios.com/datasets Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.tabularn<1K0 likes17k downloads5d agoHugging Face21mlfoundations /datacomp_1b DataComp-1B This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.image1B<n<10B53 likes15k downloads3y agoHugging Face22google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M267 likes14k downloads3y agoHugging Face23weizhiwang /Open-Qwen2VL-Data Introduction This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources. Project page: https://victorwz.github.io/Open-Qwen2VL Code: https://github.com/Victorwz/Open-Qwen2VL Dataset ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1 datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.imageimage-text-to-text10M<n<100M25 likes12k downloads1y agoHugging Face24paperswithbacktest /Stocks-Daily-Price Stocks Daily Price This dataset includes daily price data for various stocks. 25,986,919 rows over 7,764 symbols, 8 columns, covering 1962-01-02 to 2026-08-05. Refreshed monthly. Strategies Built on This Data 2,401 papers in the Papers With Backtest catalogue declare this dataset as an input. 2,238 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.35, and 45% clear a t-statistic of 1.96 on their own sample… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Daily-Price.tabulartime-series-forecasting10M<n<100M66 likes12k downloads24d agoHugging Face25InternRobotics /InternData-fractal20220817_datatabular1K<n<10K1 likes11k downloads1y agoHugging Face26toxigen /toxigen-data Dataset Card for ToxiGen Sign up for Data Access To access ToxiGen, first fill out this form. Dataset Summary This dataset is for implicit hate speech detection. All instances were generated using GPT-3 and the methods described in our paper. Languages All text is written in English. Dataset Structure Data Fields We release TOXIGEN as a dataframe with the following fields: prompt is the prompt used for generation. generation is… See the full description on the dataset page: https://huggingface.co/datasets/toxigen/toxigen-data.tabulartext-classification100K<n<1M78 likes11k downloads2y agoHugging Face27piebro /deutsche-bahn-data Deutsche Bahn Train Data This dataset contains public historical data from Deutsche Bahn, the largest German train company. It includes train schedules, delays, and cancellations from stations across Germany. For more info visit the project page at GitHub: https://github.com/piebro/deutsche-bahn-data Dataset Structure Monthly Processed Data The monthly processed data is located in monthly_processed_data/ and contains files named… See the full description on the dataset page: https://huggingface.co/datasets/piebro/deutsche-bahn-data.tabulartime-series-forecasting100M<n<1B15 likes11k downloads2h agoHugging Face28leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes11k downloads5mo agoHugging Face29dazhiyang /bsrn-merra2 BSRN MERRA-2 Atmospheric Inputs Point-extracted MERRA-2 reanalysis data for Baseline Surface Radiation Network (BSRN) stations. These parquet files provide atmospheric and aerosol inputs for the REST2 clear-sky radiation model [2]. Dataset Description Each file contains hourly MERRA-2 variables [1] at a single BSRN station location. Extraction is performed via Google Earth Engine (GEE) from NASA's 0.5° × 0.625° global grid. Data are aligned to the MERRA-2 grid… See the full description on the dataset page: https://huggingface.co/datasets/dazhiyang/bsrn-merra2.tabular10M<n<100M2 likes9.8k downloads24d agoHugging Face30UCSC-VLAA /Recap-DataComp-1B Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.imagezero-shot-classification1B<n<10B205 likes9.5k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.