CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01riotu-lab /Synthetic-UAV-Flight-Trajectories UAV Trajectory Dataset Summary This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training. Data Description The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Synthetic-UAV-Flight-Trajectories.tabular100K<n<1M17 likes5.1k downloads2y agoHugging Face02Longitude-Labs /spreadsheet-arena-release Spreadsheet Arena A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models. This is the public release accompanying the Spreadsheet Arena paper. Contents battles.csv models.csv outputs/<id>/ sheet.json sheet.xlsx <id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.tabulartabular-classificationn<1K5 likes4.8k downloads4mo agoHugging Face03MALTA-Lab /MALTA_LIBRAS malta_libras_minds_subset: Dataset tensors corresponding to all 20 LIBRAS signs from MINDS dataset. malta_libras_complete: Complete dataset tensors of all MALTA-LIBRAS collection. tabular10K<n<100K4 likes2.1k downloads1y agoHugging Face04leibnitz-lab /mdsaimagen<1K0 likes1.4k downloads1y agoHugging Face05PLAN-Lab /mTSBench mTSBench mTSBench is a collection of 344 multivariate time series from 19 datasets commonly used in anomaly detection research. Each folder corresponds to one dataset and contains *_train.csv, *_test.csv, and *_val.csv files. See data_summary.csv for per-file statistics. How to download This repository uses Git LFS for the CSV files. git lfs install git clone https://huggingface.co/datasets/PLAN-Lab/mTSBench Load with Hugging Face Select one of the… See the full description on the dataset page: https://huggingface.co/datasets/PLAN-Lab/mTSBench.tabular10M<n<100M3 likes1k downloads6d agoHugging Face06shahadalkhalifa /Crypto_Whitepaper_Labeledtabularn<1K3 likes876 downloads4y agoHugging Face07labos1 /LSV LSV: LabSuperVision Benchmark Dataset Description LSV is a multi-view video dataset of wet-lab biology experiments, captured from a mix of first-person (XMglass smart glasses), third-person (DJI action camera), and multiview (multiple synchronized phones) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors. The dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/labos1/LSV.tabularvideo-classificationn<1K0 likes787 downloads5mo agoHugging Face08stair-lab /fantastic_bugs_resulttabular100K<n<1M0 likes467 downloads1y agoHugging Face09av9ash /CSSR-S_labelled_suicidewatch_posts_reddit Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code. License and Citation This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following: @article{patil2025evaluating, title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.tabulartext-classification1K<n<10K0 likes449 downloads8mo agoHugging Face10jngb-labs /InvoiceBenchmark InvoiceBenchmark 200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number. The Pitch Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.documentquestion-answeringn<1K0 likes404 downloads5mo agoHugging Face11Mireu-Lab /NSL-KDD NSL-KDD The data set is a data set that converts the arff File provided by the link into CSV and results. The data set is personally stored by converting data to float64. If you want to obtain additional original files, they are organized in the Original Directory in the repo. Labels The label of the data set is as follows. # Column Non-Null Count Dtype 0 duration 151165 non-null int64 1 protocol_type 151165 non-null object 2 service 151165 non-null… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/NSL-KDD.tabular100K<n<1M6 likes397 downloads2y agoHugging Face12McAuley-Lab /Amazon-C4 Amazon-C4 A complex product search dataset built based on Amazon Reviews 2023 dataset. C4 is short for Complex Contexts Created by ChatGPT. Quick Start Loading Queries from datasets import load_dataset dataset = load_dataset('McAuley-Lab/Amazon-C4')['test'] >>> dataset Dataset({ features: ['qid', 'query', 'item_id', 'user_id', 'ori_rating', 'ori_review'], num_rows: 21223 }) >>> dataset[288] {'qid': 288, 'query': 'I need something that can entertain my… See the full description on the dataset page: https://huggingface.co/datasets/McAuley-Lab/Amazon-C4.tabular10K<n<100K8 likes330 downloads2y agoHugging Face13AI-Growth-Lab /patents_claims_1.5m_traim_testtabular1M<n<10M10 likes291 downloads4y agoHugging Face14Mireu-Lab /UNSW-NB15 UNSW-NB15 This data is provided through the Train, Test CSV file provided by UNSW-NB15. link Labels The label of the data set is as follows. # Column Non-Null Count Dtype 0 id 82332 non-null int64 1 dur 82332 non-null float64 2 proto 82332 non-null object 3 service 82332 non-null object 4 state 82332 non-null object 5 spkts 82332 non-null int64 6 dpkts 82332 non-null int64 7 sbytes 82332 non-null int64 8 dbytes 82332 non-null int64 9 rate… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/UNSW-NB15.tabular100K<n<1M1 likes287 downloads2y agoHugging Face15Glide-py /r_judge_labelled R-Judge with LLM-Judge Labels This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains. Files File Description r_judge_data.csv Base dataset extracted from R-Judge (568 rows, deduplicated) r_judge_labelled_anthropic_claude-sonnet-4-6.csv Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.tabulartext-classificationn<1K0 likes247 downloads3mo agoHugging Face16labhamlet /TUT2018-ov2audio10K<n<100K0 likes222 downloads8mo agoHugging Face17gvic-unb /beecrowd-beginner-labeled-topics Beecrowd Beginner Labeled Topics Dataset Summary This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.tabulartext-classificationn<1K0 likes208 downloads2mo agoHugging Face18RalphLabsAI /ralph-device-lab The same model. Small enough to sit on your phone. One configured 2.94 GB Round 7 sub2 crown. Four physical iPhones. Four open receipts. Each phone loaded the model and completed the same normal PocketPal chat prompt through the local llama.cpp Metal runtime. Watch the 30-second desktop cut · Watch the 30-second vertical cut · Watch the 15-second vertical teaser · Download the model · Inspect the machine-readable matrix · Verify every file · Read licensing and attribution… See the full description on the dataset page: https://huggingface.co/datasets/RalphLabsAI/ralph-device-lab.imagen<1K1 likes201 downloads8d agoHugging Face19murai-lab /WorcesterMA_Housing_Facades WorcesterMA_Housing_Facades: 🌐 GitHub | 🤗 Dataset Street-level photographs of housing facades from Worcester, MA, organized into four facade classes. Each image filename is the property PID (integer). The dataset links housing registry metadata (e.g., year_built) with facade images collected for research in visual housing classification. Dataset Card Dataset name: WorcesterMA_Housing_Facades Short description: Photographs of housing facades from Worcester, MA.… See the full description on the dataset page: https://huggingface.co/datasets/murai-lab/WorcesterMA_Housing_Facades.imageimage-classification10K<n<100K0 likes182 downloads9mo agoHugging Face20labhamlet /TUT2018-ov3Copyright (c) 2018 Tampere University of Technology and its licensors All rights reserved. Permission is hereby granted, without written agreement and without license or royalty fees, to use and copy the TUT Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset (“Work”) described in this document and composed of audio and metadata. This grant is only for experimental and non-commercial purposes, provided that the copyright notice in its entirety appear in all… See the full description on the dataset page: https://huggingface.co/datasets/labhamlet/TUT2018-ov3.audio10K<n<100K0 likes155 downloads8mo agoHugging Face21labhamlet /TUT2018-ov1audio1K<n<10K0 likes154 downloads8mo agoHugging Face22Orion-The-Lab /wooden_window_factory_01_enriched_v2 Real industrial data, AI-ready for Physical AI ORION WWF1 – Certified Sample Pack v2.0 (Enriched) Version Status Sector Pipeline v2.0-Enriched 🟢 Level 3 Certified Industrial-Manufacturing Orion Unified V5.2 🌟 The Evolution: Beyond Anonymization The ORION WWF1 v2.0 Enriched pack represents the professional evolution of our baseline industrial dataset. While previous versions focused on privacy-first anonymization, v2.0 transforms raw video… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v2.imagevideo-classificationn<1K0 likes140 downloads5mo agoHugging Face23idg101 /Aperture_Lab_Synthetic_Aperture_Sonar_v1 ApertureLab Synthetic SAS Dataset Version 1.0 (September 2026). Author: Isaac Gerg. Made with ApertureLab; samples, statistics and the generation pipeline are described on the dataset page. 1000 simulated synthetic aperture sonar (SAS) images, each an 80 m along-track by 200 m range swath from a HISAS 1030-class 100 kHz sonar on a straight track, beamformed by time-domain back-projection at 2.5 cm pixels and delivered as dynamic-range-compressed (DRC) TIFF LZW images with COCO… See the full description on the dataset page: https://huggingface.co/datasets/idg101/Aperture_Lab_Synthetic_Aperture_Sonar_v1.imageobject-detection1K<n<10K0 likes135 downloads1d agoHugging Face24SecureFinAI-Lab /FinRL_BTC_news_signals Overview This news dataset is created for FinAI Contest 2025 Task 1 FinRL-DeepSeek for Crypto Trading. We collected BTC news for the training and testing period from different sources [1] [2]. For each news, we use the DeepSeek chat model to extract the sentiment score, risk level, and their correpsonding confidence level and one-sentence reasoning. Column Description date_time Timestamp of when the news article was published (in UTC). title Title of the news article.… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/FinRL_BTC_news_signals.tabularn<1K0 likes121 downloads1y agoHugging Face25JacobiusMakes /lab-grown-diamond-import-monitor US Lab-Grown Diamond Import Monitor A reproducible monthly dataset on United States imports of loose cut laboratory-grown diamonds. The primary series uses US Census Bureau HTS 7104.91.10.00 data. UN Comtrade HS 710491 data provides a partner-country cross-check. The primary series covers stones cut but not set, suitable for jewelry manufacture. It excludes diamonds imported already set in finished jewelry and rough diamonds. It therefore does not measure total lab-grown diamond… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/lab-grown-diamond-import-monitor.tabular1K<n<10K0 likes117 downloads14d agoHugging Face26cleanorlabs /cleanor-storage-lab pretty_name: "Cleanor Storage Lab" license: cc-by-4.0 language: - en tags: - image-compression - avif - webp - jpeg-xl - heic - cloud-storage - benchmark - open-data size_categories: - n<1K configs: - config_name: compression-benchmark data_files: compression-benchmark.csv - config_name: heic-tax data_files: heic-tax-benchmark.csv - config_name: nextgen-formats data_files: nextgen-formats-benchmark.csv - config_name: cloud-price-index… See the full description on the dataset page: https://huggingface.co/datasets/cleanorlabs/cleanor-storage-lab.tabularn<1K1 likes115 downloads2mo agoHugging Face27kishan51 /llm-affect-lab LLM Affect Lab This dataset contains the API-level results for LLM Affect Lab, a study of functional affect signatures in language model behavior. Functional Affect Score (FAS) is a 0-1 behavioral proxy. It combines generated-token confidence, enthusiastic language, consistency across repeated samples, forced self-report computed from digit top-logprob probabilities, and length control. The goal is not to claim that models feel emotions; the goal is to measure whether different… See the full description on the dataset page: https://huggingface.co/datasets/kishan51/llm-affect-lab.tabulartext-generation1K<n<10K0 likes110 downloads5mo agoHugging Face28letrinhan /vn-provinces-labor-force-age-15-plus Vietnam provinces labor force aged 15 and over Provincial and regional labor force aged 15 years and over (thousand persons). Coverage 2005 and 2007-2024. Year 2024 is preliminary. Includes historical Ha Tay where present in the source. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Hero (continued) Comparison… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-labor-force-age-15-plus.tabular1K<n<10K0 likes103 downloads2d agoHugging Face29letrinhan /vn-provinces-labor-productivity Vietnam provinces labor productivity Provincial and regional labor productivity (million VND per worker). Coverage 2018-2024. Year 2024 is preliminary. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (441 rows) data/provinces.csv data/provinces.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-labor-productivity.tabularn<1K0 likes102 downloads2d agoHugging Face30abullard1 /steam-reviews-constructiveness-binary-label-annotations-1.5k 1.5K Steam Reviews Binary Labeled for Constructiveness Dataset Summary This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain. Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.tabulartext-classification1K<n<10K2 likes93 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.