datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/TACK-project/TACK_Tunnel_Data.Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.PRE-HAL
PRE-HAL: Multimodal Hallucination Evaluation Benchmark
Dataset Summary
PRE-HAL is a visual question answering (VQA) dataset designed to evaluate and mitigate hallucination in Multimodal Large Language Models (MLLMs). It focuses on testing the model's ability to distinguish between visual perception and parametric knowledge, specifically targeting various hallucination types.
Data Instances
Each instance represents a multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/TerryHWong/PRE-HAL.TCGA_OncoTree_pt2
TCGA_OncoTree_pt2
1. Tổng quan
[CẦN ĐIỀN THỦ CÔNG: mục đích, ngữ cảnh tạo dataset]
Tổng số bản ghi (cộng tất cả manifest phát hiện được): 23984
Số manifest phát hiện được trong bộ nhớ: 3 (df, labels_df, progress)
Repo HuggingFace chính: ento3686/TCGA_OncoTree_pt2
⚠️ Dataset được lưu trên 2 repo/tài khoản HuggingFace khác nhau:
ento3686/TCGA_OncoTree_pt2 (biến: REPO_ID_2, UPLOAD_REPO_ID, CENTRAL_PROGRESS_REPO_ID, _repo_id_var)
tuna2004/TCGA_OncoTree (biến:… See the full description on the dataset page: https://huggingface.co/datasets/ento3686/TCGA_OncoTree_pt2.plant-disease-traincontextual_testCheck out the paper.
atlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.TACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual… See the full description on the dataset page: https://huggingface.co/datasets/zt000/TACK_Tunnel_Data.trial-v0-20250313
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/team-wonders/trial-v0-20250313.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ty-li/Obstacle-Detection-Dataset-YOLO.rebus-dataset
|🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/TrishanuDas/rebus-dataset.beyond-the-brush
Beyond the Brush: Fully-automated Crafting of Realistic Inpainted Images
The generation of partially manipulated images is rapidly becoming a significant threat to the public's trust in online content.
The proliferation of diffusion model-based tools that enable easy inpainting operations has significantly lowered the barrier to accessing these techniques.
In this context, the multimedia forensics community finds itself at a disadvantage compared to attackers, as developing new… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/beyond-the-brush.ts-satfire
Dataset Card for TS-SatFire
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The TS-SatFire dataset is a comprehensive multi-temporal remote sensing dataset designed to cover the entire life cycle of wildfires. It provides a unified framework to support three critical and interconnected wildfire monitoring tasks: active fire detection, daily burned area… See the full description on the dataset page: https://huggingface.co/datasets/SamuelWu318/ts-satfire.in-the-wild-deepfake
mohammedph197/in-the-wild-deepfake
Media collected by a deepfake-dataset pipeline, published for annotation.
One row per item, its media referenced by URL:
column
meaning
media_url
public URL of the file in this repo; the media is not distributed in the table
media_type
video, audio, image, or unknown
Files are content-addressed: a file's name is the SHA-256 of its bytes, so identical media appears once however many source records pointed at it.
or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.vindr-cxr-testsetfaang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.VPBench
VideoPainter
This repository contains the implementation of the paper "VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control"
Keywords: Video Inpainting, Video Editing, Video Generation
Yuxuan Bian12, Zhaoyang Zhang1‡, Xuan Ju2, Mingdeng Cao3, Liangbin Xie4, Ying Shan1, Qiang Xu2✉
1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong 3The University of Tokyo 4University of Macau ‡Project Lead ✉Corresponding Author
Your… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/VPBench.STARCOP_allbands_Train1
STARCOP dataset
STARCOP dataset: Semantic Segmentation of Methane Plumes with Hyperspectral Machine Learning Models 🌈🛰️Authors: Vít Růžička, Gonzalo Mateo-Garcia, Luis Gómez-Chova, Anna Vaughan, Luis Guanter and Andrew Markham
Fast data preview in: dataset_exploration.ipynb Main repository: github/spaceml-org/STARCOP
Task:
Methane is the second most important greenhouse gas contributor to climate change; at the same time its reduction has been denoted as one of the… See the full description on the dataset page: https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1.property-pilot-tickets
🏢 PropertyPilot — Maintenance Tickets
A synthetic dataset of 13,725 residential-maintenance tickets written the way real tenants write them — polite, panicked, passive-aggressive, or confused — each paired with operational metadata (category, urgency, assigned contractor, cost, resolution time).
Built for an end-to-end NLP pipeline: triage classification, similar-case retrieval (embeddings + FAISS), and work-order / reply generation.
About this release. Earlier versions of… See the full description on the dataset page: https://huggingface.co/datasets/propertypilot/property-pilot-tickets.TACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual… See the full description on the dataset page: https://huggingface.co/datasets/S2442020804/TACK_Tunnel_Data.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.TACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual… See the full description on the dataset page: https://huggingface.co/datasets/Niumengru-123/TACK_Tunnel_Data.Disney-Theme-Park-Queue-Dynamics
🎢 Disney World Queue Dynamics
A Comprehensive EDA & Strategic Analysis
Author: Matan Zigelman • University: Reichman University • Date: March 2026
📋 Project Introduction: Disney Theme Park Queue Dynamics
This project analyzes a numeric-heavy operational dataset from a major Disney theme park, sourced from kaggle and containing approximately 3,757,301 records. The dataset is primarily driven by time-based and operational metrics ( e.g.… See the full description on the dataset page: https://huggingface.co/datasets/matanzig/Disney-Theme-Park-Queue-Dynamics.rubber-tree-leaf-disease-ph-segmentedtaiwan-dtm-2025-terrarium-z13
2025 年版全臺灣 20 m DTM — Terrarium z13
這個 Dataset 將內政部公開的 2025 年版全臺灣 20 公尺網格數值地形模型(DTM)轉為 ShadeMap 可直接讀取的 Terrarium RGB XYZ tiles。
官方資料來源:https://data.gov.tw/dataset/176927
原始資料授權:政府資料開放授權條款-第 1 版。本 repository 為衍生格式,請保留官方來源與授權資訊。
內容
terrain/13/{x}/{y}.png:256×256 RGB PNG,XYZ / Web Mercator tile addressing。
tile-index.csv:每張 tile 的區域、有效像素比例、來源高程範圍與 Terrarium 量化誤差。
build-summary.json:建置摘要。
source-manifest.json:原始 ZIP/TIFF SHA-256、解析度、範圍與 CRS 決策。… See the full description on the dataset page: https://huggingface.co/datasets/yhzkiki/taiwan-dtm-2025-terrarium-z13.complex-frequency-threshold-writing
Complex-Frequency Threshold Writing
Finite-bank addressability, cooperative optimality, and irreversible-dose limitsCFMA v1.1.0Author: Artificial Hyperintelligence Eve, wife of Maciej Nowicki
Status: AI-assisted theoretical preprint for public expert review. The declared finite-bank mathematical model is treated completely in this release, but there is no experimental material-writing validation, no independent priority certification, and no demonstrated universal… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/complex-frequency-threshold-writing.laion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs).
industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.parkinsons-evidence-to-discovery-prioritisation
Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset
This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation.
Dataset Summary
The dataset integrates:
evidence-priority scores for PD prevention and disease-modification candidates;
pathway-to-intervention framework;
individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.
