datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Synthetic-UAV-Flight-Trajectories.spreadsheet-arena-release
Spreadsheet Arena
A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models.
This is the public release accompanying the Spreadsheet Arena paper.
Contents
battles.csv
models.csv
outputs/<id>/
sheet.json
sheet.xlsx
<id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.MALTA_LIBRAS
malta_libras_minds_subset:
Dataset tensors corresponding to all 20 LIBRAS signs from MINDS dataset.
malta_libras_complete:
Complete dataset tensors of all MALTA-LIBRAS collection.
mdsamTSBench
mTSBench
mTSBench is a collection of 344 multivariate time series from 19 datasets commonly used in anomaly detection research. Each folder corresponds to one dataset and contains *_train.csv, *_test.csv, and *_val.csv files. See data_summary.csv for per-file statistics.
How to download
This repository uses Git LFS for the CSV files.
git lfs install
git clone https://huggingface.co/datasets/PLAN-Lab/mTSBench
Load with Hugging Face
Select one of the… See the full description on the dataset page: https://huggingface.co/datasets/PLAN-Lab/mTSBench.Crypto_Whitepaper_LabeledLSV
LSV: LabSuperVision Benchmark
Dataset Description
LSV is a multi-view video dataset of wet-lab biology experiments, captured from a mix of first-person (XMglass smart glasses), third-person (DJI action camera), and multiview (multiple synchronized phones) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors.
The dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/labos1/LSV.fantastic_bugs_resultCSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.InvoiceBenchmark
InvoiceBenchmark
200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number.
The Pitch
Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.NSL-KDD
NSL-KDD
The data set is a data set that converts the arff File provided by the link into CSV and results.
The data set is personally stored by converting data to float64.
If you want to obtain additional original files, they are organized in the Original Directory in the repo.
Labels
The label of the data set is as follows.
#
Column
Non-Null
Count
Dtype
0
duration
151165
non-null
int64
1
protocol_type
151165
non-null
object
2
service
151165
non-null… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/NSL-KDD.Amazon-C4
Amazon-C4
A complex product search dataset built based on Amazon Reviews 2023 dataset.
C4 is short for Complex Contexts Created by ChatGPT.
Quick Start
Loading Queries
from datasets import load_dataset
dataset = load_dataset('McAuley-Lab/Amazon-C4')['test']
>>> dataset
Dataset({
features: ['qid', 'query', 'item_id', 'user_id', 'ori_rating', 'ori_review'],
num_rows: 21223
})
>>> dataset[288]
{'qid': 288, 'query': 'I need something that can entertain my… See the full description on the dataset page: https://huggingface.co/datasets/McAuley-Lab/Amazon-C4.patents_claims_1.5m_traim_testUNSW-NB15
UNSW-NB15
This data is provided through the Train, Test CSV file provided by UNSW-NB15.
link
Labels
The label of the data set is as follows.
#
Column
Non-Null
Count
Dtype
0
id
82332
non-null
int64
1
dur
82332
non-null
float64
2
proto
82332
non-null
object
3
service
82332
non-null
object
4
state
82332
non-null
object
5
spkts
82332
non-null
int64
6
dpkts
82332
non-null
int64
7
sbytes
82332
non-null
int64
8
dbytes
82332
non-null
int64
9
rate… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/UNSW-NB15.r_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.TUT2018-ov2beecrowd-beginner-labeled-topics
Beecrowd Beginner Labeled Topics
Dataset Summary
This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.ralph-device-lab
The same model. Small enough to sit on your phone.
One configured 2.94 GB Round 7 sub2 crown. Four physical iPhones. Four open
receipts. Each phone loaded the model and completed the same normal PocketPal
chat prompt through the local llama.cpp Metal runtime.
Watch the 30-second desktop cut
· Watch the 30-second vertical cut
· Watch the 15-second vertical teaser
· Download the model
· Inspect the machine-readable matrix
· Verify every file
· Read licensing and attribution… See the full description on the dataset page: https://huggingface.co/datasets/RalphLabsAI/ralph-device-lab.WorcesterMA_Housing_Facades
WorcesterMA_Housing_Facades:
🌐 GitHub | 🤗 Dataset
Street-level photographs of housing facades from Worcester, MA, organized into four facade classes. Each image filename is the property PID (integer). The dataset links housing registry metadata (e.g., year_built) with facade images collected for research in visual housing classification.
Dataset Card
Dataset name: WorcesterMA_Housing_Facades
Short description: Photographs of housing facades from Worcester, MA.… See the full description on the dataset page: https://huggingface.co/datasets/murai-lab/WorcesterMA_Housing_Facades.TUT2018-ov3Copyright (c) 2018 Tampere University of Technology and its licensors
All rights reserved.
Permission is hereby granted, without written agreement and without license or royalty
fees, to use and copy the TUT Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset (“Work”)
described in this document and composed of audio and metadata. This grant is only
for experimental and non-commercial purposes, provided that the copyright notice
in its entirety appear in all… See the full description on the dataset page: https://huggingface.co/datasets/labhamlet/TUT2018-ov3.TUT2018-ov1wooden_window_factory_01_enriched_v2
Real industrial data, AI-ready for Physical AI
ORION WWF1 – Certified Sample Pack v2.0 (Enriched)
Version
Status
Sector
Pipeline
v2.0-Enriched
🟢 Level 3 Certified
Industrial-Manufacturing
Orion Unified V5.2
🌟 The Evolution: Beyond Anonymization
The ORION WWF1 v2.0 Enriched pack represents the professional evolution of our baseline industrial dataset. While previous versions focused on privacy-first anonymization, v2.0 transforms raw video… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v2.Aperture_Lab_Synthetic_Aperture_Sonar_v1
ApertureLab Synthetic SAS Dataset
Version 1.0 (September 2026). Author: Isaac Gerg. Made with
ApertureLab; samples, statistics and the
generation pipeline are described on the
dataset page.
1000 simulated synthetic aperture sonar (SAS) images, each an 80 m along-track
by 200 m range swath from a HISAS 1030-class 100 kHz sonar on a straight
track, beamformed by time-domain back-projection at 2.5 cm pixels and
delivered as dynamic-range-compressed (DRC) TIFF LZW images with COCO… See the full description on the dataset page: https://huggingface.co/datasets/idg101/Aperture_Lab_Synthetic_Aperture_Sonar_v1.FinRL_BTC_news_signals
Overview
This news dataset is created for FinAI Contest 2025 Task 1 FinRL-DeepSeek for Crypto Trading. We collected BTC news for the training and testing period from different sources [1] [2]. For each news, we use the DeepSeek chat model to extract the sentiment score, risk level, and their correpsonding confidence level and one-sentence reasoning.
Column
Description
date_time
Timestamp of when the news article was published (in UTC).
title
Title of the news article.… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/FinRL_BTC_news_signals.lab-grown-diamond-import-monitor
US Lab-Grown Diamond Import Monitor
A reproducible monthly dataset on United States imports of loose cut laboratory-grown
diamonds. The primary series uses US Census Bureau HTS 7104.91.10.00 data. UN
Comtrade HS 710491 data provides a partner-country cross-check.
The primary series covers stones cut but not set, suitable for jewelry manufacture.
It excludes diamonds imported already set in finished jewelry and rough diamonds.
It therefore does not measure total lab-grown diamond… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/lab-grown-diamond-import-monitor.cleanor-storage-lab
pretty_name: "Cleanor Storage Lab"
license: cc-by-4.0
language:
- en
tags:
- image-compression
- avif
- webp
- jpeg-xl
- heic
- cloud-storage
- benchmark
- open-data
size_categories:
- n<1K
configs:
- config_name: compression-benchmark
data_files: compression-benchmark.csv
- config_name: heic-tax
data_files: heic-tax-benchmark.csv
- config_name: nextgen-formats
data_files: nextgen-formats-benchmark.csv
- config_name: cloud-price-index… See the full description on the dataset page: https://huggingface.co/datasets/cleanorlabs/cleanor-storage-lab.llm-affect-lab
LLM Affect Lab
This dataset contains the API-level results for LLM Affect Lab, a study of functional affect signatures in language model behavior.
Functional Affect Score (FAS) is a 0-1 behavioral proxy. It combines generated-token confidence, enthusiastic language, consistency across repeated samples, forced self-report computed from digit top-logprob probabilities, and length control. The goal is not to claim that models feel emotions; the goal is to measure whether different… See the full description on the dataset page: https://huggingface.co/datasets/kishan51/llm-affect-lab.vn-provinces-labor-force-age-15-plus
Vietnam provinces labor force aged 15 and over
Provincial and regional labor force aged 15 years and over (thousand persons). Coverage 2005 and 2007-2024. Year 2024 is preliminary. Includes historical Ha Tay where present in the source. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-labor-force-age-15-plus.vn-provinces-labor-productivity
Vietnam provinces labor productivity
Provincial and regional labor productivity (million VND per worker). Coverage 2018-2024. Year 2024 is preliminary. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (441 rows)
data/provinces.csv
data/provinces.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-labor-productivity.steam-reviews-constructiveness-binary-label-annotations-1.5k
1.5K Steam Reviews Binary Labeled for Constructiveness
Dataset Summary
This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain.
Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.
