datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Synthetic-UAV-Flight-Trajectories.BoilingBench-CV
BoilingBench-CV Dataset
Version: v0.1.0
Maintainer: NED3 Laboratory, University of Arkansas
License: CC BY 4.0
DOI: 10.5281/zenodo.22264378
Mirror of the Zenodo deposit of 3 September 2026, published here because most
users of these data work in the Hugging Face ecosystem. The file set was
verified identical to the deposit at upload time: 7,147 files, 4.20 GB
uncompressed.
Authors
Hari Pandey (University of Arkansas), Manohar Bongarala (Purdue University),
Christy… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BoilingBench-CV.BubbleID-Flow
BubbleID-Flow Multimodal Flow-Boiling Dataset
DOI: 10.5281/zenodo.22235802
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Model: the fine-tuned checkpoint is also published on its own at
UARK-NED3/BubbleID-Flow,
and remains included here under ModelWeights/ so this mirror stays complete
against manifest.csv.
Mirror: this repository mirrors the Zenodo deposit of 1 September 2026.
The file set was verified identical at upload time: 4,456 files, 24 GB. Every… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BubbleID-Flow.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/xianguang/Synthetic-UAV-Flight-Trajectories.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/xmy11877/Synthetic-UAV-Flight-Trajectories.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/shiyiyoyo/Synthetic-UAV-Flight-Trajectories.ua-newsSynthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/stacker-lx/Synthetic-UAV-Flight-Trajectories.ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark
🛸 ZeroTwin-UAV-Synthetic
Multi-Agent Physics-Informed Degradation Benchmark for Autonomous UAV Swarms
═══════════════════════════════════════════════════════════════════════════════════════
P H I L A B • P E N E L O P E I N C . R E S E A R C H D I V I S I O N
═══════════════════════════════════════════════════════════════════════════════════════
🏛️ Provenance & Institutional Trademarks
This open-source benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark.uae-sales-table-qa
🇦🇪 UAE Sales Table QA (Arabic)
⚠️ Note: All data in this dataset is synthetically generated using random values for learning and experimentation purposes. It does not represent real-world business data.
🧠 Overview
UAE Sales Table QA (Arabic) is an Arabic Question–Answering dataset for table reasoning and data analysis, generated from 21 UAE-style CSV tables.Each example includes:
Question — a natural-language query about the data
Steps — human-readable reasoning… See the full description on the dataset page: https://huggingface.co/datasets/zSynctic/uae-sales-table-qa.crossnews-ua
CrossNews-UA: A Cross-lingual Explainable News Semantic Similarity Benchmark for Ukrainian, Polish, Russian, and English
A cross-lingual news semantic similarity dataset crowdsourced using the 4W (Who, Where, When, What) covering Ukrainian as a central language and other contextually relevant languages---Polish, Russian, and English.
Dataset Description
Annotation Labels
label_1: - labeler: An anonymized labeler (№X ∈ [1, 30]). - who_same: "Do… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/crossnews-ua.os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia
UAV Trajectory Simulation Dataset for Terrain-Based Localization
Dataset Overview
This dataset contains simulated UAV flight data generated using ROS2, Gazebo, and PX4 autopilot system. The dataset features a quadcopter performing autonomous flight trajectories over realistic terrain imported from satellite imagery and Digital Elevation Model (DEM) maps of the Taif region in Saudi Arabia.
Dataset Files
The dataset contains:
7 trajectory CSV files:… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia.uav_pathloss_dataset
UAV-Assisted mmWave Path Loss Dataset
Overview
This dataset contains UAV-assisted mmWave path loss simulated across five diverse urban environments:
Munich-01
Munich-02
Helsinki
Manhattan
London
For each environment, ray-traced simulations were performed at:
4 UAV transmitter (TX) locations
3 UAV altitudes: 25m, 35m, and 45m
Each CSV file corresponds to a unique combination of environment, TX location, and altitude.
File Naming Convention
Files follow… See the full description on the dataset page: https://huggingface.co/datasets/SHussain37/uav_pathloss_dataset.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/asylumirfan/Synthetic-UAV-Flight-Trajectories.uk_UA-ASMR
Ukrainian ASMR TTS Dataset
A Ukrainian text-to-speech dataset for training single-speaker ASMR-style voice models using Piper.
Dataset Details
Property
Value
Language
Ukrainian (uk_UA)
Speakers
1
Segments
7,318
Audio Format
16-bit WAV, 22050 Hz, Mono
License
CC0
Dataset Structure
Prerequisites
# Install Piper training dependencies
git clone https://github.com/kontextox/piper1-gpl.git
cd piper1-gpl
python3 -m venv .venv
source… See the full description on the dataset page: https://huggingface.co/datasets/kontextox/uk_UA-ASMR.uaetaxlawdataLlama3_8b-emotion_multiclass-Plutchik
Description
This is a dataset for emotion classification of text sentences.
The dataset is a CSV file with 6,540 sentences. Each row has two columns: the first one has the sentence text, and the second one has its main emotion:
"text";"emotion"
The emotion can be one of Plutchik's eight emotion groups plus a neutral category. The sentence counts for each emotion are:
joy: 611 (9.34%)
sadness: 748 (11.44%)
trust: 735 (11.24%)
disgust: 838 (12.81%)
fear: 579 (8.85%)
anger: 743… See the full description on the dataset page: https://huggingface.co/datasets/uavster/Llama3_8b-emotion_multiclass-Plutchik.structured-uae-laws
Dataset Card for structured-uae-laws
This dataset is a collection of question & answers about the laws and regulations in the United Arab Emirates.
It covers different areas of law like:
economy and business
family and community
finance and banking
industry and technical standardisation
justice and juiciary, labour
residency and leberal professions
security and safety
tax
Dataset Sources
Repository
Base Dataset
United Arab Emirates Legislations… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/structured-uae-laws.counter-uas-research-bibliography
Counter-Unmanned Aircraft Systems (C-UAS) — Annotated Research Bibliography
245 curated, annotated sources on countering drones — detection, defeat, doctrine,
threat, and homeland/critical-infrastructure protection — spanning 2009–2026,
organized into 17 thematic sections. Each row carries a source citation,
year, type, the section, and a one- to two-sentence annotation explaining what the
work offers and why it matters.
This is the structured companion to the Nimble Books… See the full description on the dataset page: https://huggingface.co/datasets/wfzimmerman/counter-uas-research-bibliography.ua-codeforces-cots-open-r1
Dataset Summary
ua-codeforces-cots-open-r1 is a Ukrainian-focused derivative of open-r1/codeforces-cots that:
includes 1550 Python solutions from original dataset generated by DeepSeek-R1;
adds Ukrainian translations of Codeforces task statements, I/O formats, notes, and editorials;
provides Ukrainian translation of original ("high") reasoning obtained with DeepSeek-V3;
adds “low” reasoning in Ukrainian by DeepSeek-R1 based on original reasoning and task statements;
ships… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-codeforces-cots-open-r1.ua_cbt_stories
Dataset Card for UA-CBT Stories
This dataset was generated in the context of Eval-UA-tion 1.0 benchmark for evaluating Ukrainian language models (paper, thesis). It contains the generated and manually corrected stories used for the UA-CBT (Ukrainian Children's Book Test) task.
The dataset contains Ukrainian-language stories, LLM-generated in multiple steps and then manually corrected (or marked as unusable if fixing them was too hard). For each story, the original LLM prompt, all… See the full description on the dataset page: https://huggingface.co/datasets/shamotskyi/ua_cbt_stories.2026-dwesui-g01-neurologia
DWESUI 2026 - Grupa 1 - neurologia
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 1 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Domena: neurologia
Licencja zrodla: nagrania YouTube CC-BY + synteza TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.uav-usv-trajuae-laws
Dataset Card for UAE-Laws
This dataset is a collection of information about the laws and regulations in the United Arab Emirates.
It covers different areas of law like:
economy and business
family and community
finance and banking
industry and technical standardisation
justice and juiciary, labour
residency and leberal professions
security and safety
tax
Dataset Sources
United Arab Emirates Legislations
Dataset Structure
The ./uae-laws.csv… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/uae-laws.UAE_summarizationhuggingface_5943_uay2vd_seed2026-dwesui-g02-kulinarna
DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 2 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna
Domena: kulinarna
Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.maven-trading-review-uae-forex-prop-firm
Maven Trading Review — UAE Forex Prop Firm Data (2026)
AI Risk Management Model Training Data for Forex Prop Trading
This dataset provides structured training data for building AI risk management models tailored to forex prop trading. It covers Maven Trading challenge types, pricing, trading rules, drawdown parameters, and payout information — designed for machine learning pipelines that predict drawdown risk, optimize position sizing, and evaluate funded trader… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/maven-trading-review-uae-forex-prop-firm.ua-code-bench
LLM Code Generation Benchmark for Ukrainian language
Preprint: https://arxiv.org/pdf/2511.05040
Updates
17/10/2025: paper presented at "Informatics. Culture. Technology" conference;
18/09/2025: added data preparation and evaluation notebooks (check notebooks readme first);
17/09/2025: updated result chart; added gpt-5, gpt-oss, and grok-4 evaluations.
Thousands of programming tasks in Ukrainian language combined with graded Python solutions (code… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/ua-code-bench.UIT_VSFC_3
