datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spaces-of-the-week-legacymulti-agent-scam-conversation
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.abuse-scanner-bot-datasetscam-dialogue
Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/scam-dialogue.instagram_bot_detection
Instagram Fake Profile Detection Dataset
Dataset Summary
This dataset contains 5,000 Instagram profiles labeled as either fake or real, designed for binary classification tasks in social media fraud detection. The dataset provides comprehensive profile features that can be used to train machine learning models to automatically identify fake Instagram accounts.
Dataset Details
Total Samples: 5,000 profiles
Classes: Binary (0 = Real, 1 = Fake)
Class… See the full description on the dataset page: https://huggingface.co/datasets/nahiar/instagram_bot_detection.single-agent-scam-conversations
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset
Dataset Description
The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The transcribed conversation between the caller and receiver.
type: The specific type of scam or non-scam interaction.
labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.bank-bot-conversationvn-provinces-income-top-bottom-gap
Vietnam income top-bottom gap by locality
Vietnam monthly per-capita income (thousand VND, current prices) for the lowest and highest income groups, plus the top-to-bottom ratio (times), by locality for 2010-2024 (even years through 2018, then annual. 2024 preliminary). Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-income-top-bottom-gap.Support-Bot-Recommendationtrading-bot-backtesting
Trading bot Backtesting CSVs
Gunbot ⇄ Trading Bot backtests (pair-level candle & matched order data)Created with Gunbot on Binance spot.
Dataset structure
column
type
description
ts
int64
candle Unix ms timestamp
open/high/low/close
float
OHLC price values
volume
float
traded volume in base currency
order_type
str
buy, sell, or empty (no order)
order_rate
float
executed rate
order_amount
float
amount traded
order_id
int64
exchange order id
pnl… See the full description on the dataset page: https://huggingface.co/datasets/kapr/trading-bot-backtesting.Medical_VLM_SycophancyThis the official data hosting repository for paper "EchoBench: Benchmarking Sycophancy in Medical
Large Vision Language Models".
============open-source_models============
For experiments on open-source models, our implementation is built upon the VLMEvalkit framework.
Navigate to the VLMEval directory
Set up the environment by running: "pip install -e ."
Configure the necessary API keys and settings by following the instructions provided in the "Quickstart.md" file of VLMEvalkit.
To… See the full description on the dataset page: https://huggingface.co/datasets/Botai666/Medical_VLM_Sycophancy.Scammer-ConversationThis dataset are generated by gretelai/tabular-v0
This dataset contains a collection of conversations between scammers, scam baiters, and normal conversations. The purpose of this dataset is to provide a resource for training and evaluating models for scam detection and classification.
twitter-human-botsringdown-damping-signals
Ring-Down Damping Signals: 12K Labelled Decay Waveforms
How this dataset was created
This is original data created programmatically — it was not collected, recorded, scraped, or
derived from any external source. Each of the 12,000 signals was generated from scratch by a
deterministic, seeded Python program:
Draw the label theta uniformly at random from [1, 5], plus a random overall base-decay rate.
Pick a random number of tones ("modes", 30–55), each with a… See the full description on the dataset page: https://huggingface.co/datasets/botfx/ringdown-damping-signals.user_feedbackboth-star-5c13b7
both-star-5c13b7
Synthetic products test data: 38 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Indigo-Michael/both-star-5c13b7.twitter_bot_detectiontrading-bot-backtesting
Trading bot Backtesting CSVs
Gunbot ⇄ Trading Bot backtests (pair-level candle & matched order data)Created with Gunbot on Binance spot.
Dataset structure
column
type
description
ts
int64
candle Unix ms timestamp
open/high/low/close
float
OHLC price values
volume
float
traded volume in base currency
order_type
str
buy, sell, or empty (no order)
order_rate
float
executed rate
order_amount
float
amount traded
order_id
int64
exchange order id
pnl… See the full description on the dataset page: https://huggingface.co/datasets/johnlprice82/trading-bot-backtesting.structural-bottleneck-classification-v0.1
What this dataset does
This dataset tests whether a model can detect structural bottlenecks.
The task is simple:
Given a scenario and a structural-bottleneck claim, predict whether the claim is supported.
Core stability idea
A structural bottleneck is a constraint that limits system performance regardless of improvements elsewhere.
Typical bottlenecks include:
single approval points
single processing nodes
unique dependencies
centralized routing
irreplaceable personnel… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/structural-bottleneck-classification-v0.1.both-gear-f15f18
both-gear-f15f18
Synthetic sensors test data: 35 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Atlas-Theo/both-gear-f15f18.Bot-IOT_LLMyoutube-scam-conversationsboth-commission-af4675
both-commission-af4675
Synthetic sensors test data: 44 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Nova-Dawn/both-commission-af4675.BotsDetect
BotsDetect
tags: Classification, Human/Bit Detection, Mouse Movements, Keystrokes
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'BotsDetect' dataset comprises various environmental parameters collected to distinguish between human and bot interactions within a digital environment. The dataset is designed to assist machine learning practitioners in training models for the purpose of detecting bot activity by analyzing user… See the full description on the dataset page: https://huggingface.co/datasets/akhil033/BotsDetect.healthcare-discharge-bottleneck-coherence-risk-v0.1What this repo is for
detect discharge blockage early
predict bed block
protect elective lists
improve throughput
reduce ED crowding
base_model_sprint
Base Model Metadata Sprint
Description
Join us in improving the discoverability and understanding of models on the Hugging Face Hub by adding base_model metadata! This sprint aims to enhance the information available for models derived from, fine-tuned on, or quantized versions of existing base models.
🤗 Strong contributions will win prizes!! 🤗
Why It Matters
Adding base_model metadata helps users:
Easily find models derived from specific architectures… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/base_model_sprint.Discord-Botbot_etheriumbot-responsesbot-courteous-interactionsThis dataset contains a collection of variations in polite and courteous responses generated by a conversational AI. This dataset is designed to enhance natural language understanding and generation models, focusing on responses that convey gratitude, appreciation, and helpfulness. Each entry in the dataset pairs a user input with multiple variations of a polite, appreciative response, aiming to enrich conversational models with diverse ways of expressing politeness and support.
