datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.Language-DetectionLanguage_Detection
Language_Detection - Multilingual Text Classification Dataset
This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks.
Dataset Overview
The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.suicide_depression_detectioneuroparl_for_language_detection_10kgnss-jamming-spoofing-detection
GNSS Jamming & Spoofing Detection Dataset
A physics-informed synthetic dataset for detecting GPS/GNSS cyber-attacks
(Jamming and Spoofing) from satellite-signal features. Built for the
GNSS Guardian project — Introduction to Data Science final project.
Overview
14,850 samples across 450 scenarios × 33 time-steps each
3 balanced classes: Normal / Jamming / Spoofing (4,950 each)
26 columns: multi-constellation signal features + attack metadata + text descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Omrilevi123/gnss-jamming-spoofing-detection.Detection-for-SuicideObstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ShafinSI/Obstacle-Detection-Dataset-YOLO.Network-Intrusion-Detection-DataObstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ty-li/Obstacle-Detection-Dataset-YOLO.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/eziodad/Obstacle-Detection-Dataset-YOLO.Creditcard-fraud-detection
Credit Card Fraud Detection
This dataset was downloaded from https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud/data adn uploaded for educational purposes.
Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset
Nigerian Financial Transactions and Fraud Detection Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset.language_detection
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/basilb2s/language-detection
It's a small language detection dataset. This dataset consists of text details for 17 different languages, ie, you will be able to create an NLP model for predicting 17 different language..
Mawqif_Stance-Detection
Mawqif: A Multi-label Arabic Dataset for Target-specific Stance Detection
Mawqif is the first Arabic dataset that can be used for target-specific stance detection.
This is a multi-label dataset where each data point is annotated for stance, sentiment, and sarcasm.
We benchmark Mawqif dataset on the stance detection task and evaluate the performance of four BERT-based models. Our best model achieves a macro-F1 of 78.89%.
Mawqif Statistics
This dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/NoraAlt/Mawqif_Stance-Detection.language-detectionLMD-AI-Detection
LMD AI-Generated Music Detection Benchmark
(Note: The corresponding research paper will be released later.)
Dataset Description
The rapid advancement of AI music generation has raised growing concerns about the authenticity of digital music. While deepfake detection has been extensively studied in the audio domain, symbolic music (MIDI) remains largely unexplored.
This dataset presents a comprehensive benchmark for AI-generated symbolic music detection, examining… See the full description on the dataset page: https://huggingface.co/datasets/dhlee3000/LMD-AI-Detection.Phantom_Hallucination_Detection
Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA
Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
This is the repository containing the dataset for the submission mentioned above.
This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/seyled/Phantom_Hallucination_Detection.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/NivithaChandran/Obstacle-Detection-Dataset-YOLO.fake-news-detection-dataset-EnglishThis is a cleaned and splitted version of this dataset (https://www.kaggle.com/datasets/sadikaljarif/fake-news-detection-dataset-english)
Labels:
Fake News: 0
Real News: 1
You can find the cleansing script at: https://github.com/ErfanMoosaviMonazzah/Fake-News-Detection
ai-human-text-detection-v1
🧠 AI vs Human Text Detection Dataset (v1)
This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation.
🔗 Sources
The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification:
Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus
gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.Traffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/star092304/Traffic-sign-detection-VietNam.turkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/nithi060488/Cifer-Fraud-Detection-Dataset-AF.instagram_bot_detection
Instagram Fake Profile Detection Dataset
Dataset Summary
This dataset contains 5,000 Instagram profiles labeled as either fake or real, designed for binary classification tasks in social media fraud detection. The dataset provides comprehensive profile features that can be used to train machine learning models to automatically identify fake Instagram accounts.
Dataset Details
Total Samples: 5,000 profiles
Classes: Binary (0 = Real, 1 = Fake)
Class… See the full description on the dataset page: https://huggingface.co/datasets/nahiar/instagram_bot_detection.chinese-ai-detection-dataset
Chinese AI Detection Dataset
中文AI文本检测数据集
数据集简介
用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。
核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。
数据统计
类型
样本数
说明
总计
66,001
训练/验证/测试集
纯人类
27,719
多领域人类文本
纯AI
27,719
多模型生成
C2 (续写)
3,781
人类开头+AI续写
C3 (改写)
3,781
AI改写人类文本
C4 (润色)
3,001
AI润色人类文本
数据格式
{
"text": "文本内容(混合文本包含[SEP]标记)",
"label": 0, // 0=Human, 1=AI
"category": "C2", // Human/AI/C2/C3/C4
"source": "数据来源"
}… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.credit-card-fraud-detection
Credit Card Fraud Detection – Processed Dataset
This dataset contains preprocessed credit card transaction data prepared for fraud detection tasks.
Data Description
The dataset is derived from anonymized transaction records and includes numerical features (V1–V28), transaction amount, and time-based information.
Preprocessing Steps
Feature scaling and normalization
Handling class imbalance
Feature selection based on correlation analysis
Removal of irrelevant… See the full description on the dataset page: https://huggingface.co/datasets/jyunyilin/credit-card-fraud-detection.Phishing_Detection_Dataset
