datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI
Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.RAGDOLL
The RAGDOLL E-Commerce Webpage Dataset
This repository contains the RAGDOLL (Retrieval-Augmented Generation Deceived Ordering via AdversariaL materiaLs) dataset as well as its LLM-automated collection pipeline.
The RAGDOLL dataset is from the paper Ranking Manipulation for Conversational Search Engines from Samuel Pfrommer, Yatong Bai, Tanmay Gautam, and Somayeh Sojoudi. For experiment code associated with this paper, please refer to this repository.
The dataset consists of 10… See the full description on the dataset page: https://huggingface.co/datasets/Bai-YT/RAGDOLL.SVBench
Dataset Card for SVBench
This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository.
Dataset Details
Dataset Description
SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.NSRDB_extractPublic domain data extracted from National Solar Radiation Database: https://nsrdb.nrel.gov/data-viewer
montreal_fireYfOptionsMAGE
MAGE: Machine-generated Text Detection in the Wild
🚀 Introduction
Recent advances in large language models have enabled them to reach a level of text generation comparable to that of humans.
These models show powerful capabilities across a wide range of content, including news article writing, story generation, and scientific writing.
Such capability further narrows the gap between human-authored and machine-generated texts, highlighting the importance of machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/yaful/MAGE.LongDA
LongDA Dataset Card
Dataset Description
LongDA is a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. It features authentic U.S. government survey data with complete, long documentation, testing LLMs' ability to navigate complex real-world datasets before performing analysis.
Dataset Summary
505 queries extracted from 30 expert-written publications
17 U.S. national surveys covering health… See the full description on the dataset page: https://huggingface.co/datasets/Yiyang-Ian-Li/LongDA.ted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.sickspatial457_mcqprompt_injections
Dataset Card for Prompt Injections by Yanis Miraoui 👋
Dataset Description
This dataset of prompt injections enriches Large Language Models (LLMs) by providing task-specific examples and prompts, helping improve LLMs' performance and control their behavior.
Dataset Summary
This dataset contains over 1000 rows of prompt injections in multiple languages. It contains examples of prompt injections using different techniques such as: prompt leaking… See the full description on the dataset page: https://huggingface.co/datasets/yanismiraoui/prompt_injections.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.PolyOmics
PolyOmics
PolyOmics is an omics-scale computational materials database containing molecular structures, simulation metadata, and diverse physical properties for more than (10^5) polymeric materials.
The database was generated primarily using RadonPy, a fully automated molecular dynamics (MD) simulation platform for polymer materials. PolyOmics was developed through a large-scale collaboration of the RadonPy Consortium to provide foundational computational data for polymer… See the full description on the dataset page: https://huggingface.co/datasets/yhayashi1986/PolyOmics.ycuppe-midi
YCU-PPE-III: Piano Performance MIDI Dataset
MIDI transcriptions of the YCU-PPE-III piano performance dataset (Wang et al.), used for unreferenced Performance MOS (PMOS) prediction in EVPMR.
Overview
2,627 MIDI files transcribed from WAV recordings via transkun
13 songs performed by student pianists
2,511 performances with ratings from 3 expert judges (0-100 scale each)
Labels: normalized mean score to [0, 1]
Splits: 1,757 train / 377 val / 377 test (stratified by song)… See the full description on the dataset page: https://huggingface.co/datasets/anusfoil/ycuppe-midi.BonaFide
BonaFide
This is a dataset containing ground-truth faithfulness labels for chains of thought (CoTs), used for evaluating CoT faithfulness metrics. The current benchmark results are in the BonaFide benchmark space.
Methodology
We construct tasks whose outputs reveal which intermediate computations must have produced them, then label CoTs against those computations.
Diversionary setting. Each question is given alongside a misleading hint pointing to a random wrong answer.… See the full description on the dataset page: https://huggingface.co/datasets/yoavgurarieh/BonaFide.agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/yuanyyaa/agent-reward-bench.SP_500_Stocks_Data-ratios_news_price_10_yrsHi folks,
Here is a collection of data I have scraped or aggregated for most of the stocks in the S&P 500, including popular ones like Apple (AAPL).
It has the following data:
Daily news articles and sentiments on those articles collected over the last few years.
All quarterly stock fundamentals (ratios) for 10-20 years.
Stock price data (daily close) over the last 10-20 years.
Use it however you please for PERSONAL USAGE, but if you do leverage it to make some money; just remember me and… See the full description on the dataset page: https://huggingface.co/datasets/pmoe7/SP_500_Stocks_Data-ratios_news_price_10_yrs.youtube-tiktok-trends-dataset-2025
🎬 YouTube Shorts & TikTok Trends (2025)
Author: Tarek MasryoLicense: CC BY 4.0
A structured snapshot of short-form video activity across YouTube Shorts and TikTok during 2025 (Jan–Aug).Built for content intelligence, analytics dashboards, and ML baselines (classification/regression).
What’s inside
This repository ships:
Two loadable dataset configs (via datasets.load_dataset):
default → ML-ready table (cleaned + modeling-friendly)
raw → raw video-level table (wider… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/youtube-tiktok-trends-dataset-2025.ex-repairmidi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/ygonet/midi-classical-music.AncientDoc
AncientDoc: A Benchmark for Chinese Ancient Document Understanding
AncientDoc 是第一个专为 中国古籍文档理解 设计的综合基准数据集,涵盖从 OCR 到 知识推理 的多任务评测,旨在推动多模态大模型在古籍场景下的识别、理解与推理能力研究。
数据集简介
数据规模:2,973 页
文献数量:约 100 本
文献类型:14 类(如总集、楚辞体诗、诗文批评、类书、谱录等)
时间跨度:从战国到清代,涵盖多个重要历史时期
任务类型:
Page-level OCR:整页文字识别(含竖排、异体字、批注等复杂情况)
Vernacular Translation:文言文到现代汉语的同语种翻译
Reasoning-based QA:基于文意的隐性推理问答
Knowledge-based QA:基于文本事实和背景知识的问答
Linguistic Variant QA:文体、修辞与语言风格相关的问答
数据分布
按朝代分布… See the full description on the dataset page: https://huggingface.co/datasets/yuchuan123/AncientDoc.czech_bank_qa
CzechBankQA
This is a list of SQL queries for a text-to-SQL task over the Czech Bank 1999 dataset.
solana-yield-honesty
Solana Honesty Index
What each Solana stablecoin product says it pays, next to what it actually
paid, measured from a share price rather than from a claim.
Snapshot generated 2026-09-24T12:19:33.744Z. Window 30 days.
13 products across 3 protocols,
13 comparable, 0 published but not
comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed.
product
advertised
realized
gap
delivered
realized method
Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ShafinSI/Obstacle-Detection-Dataset-YOLO.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.LSV
LSV: LabSuperVision Benchmark
Dataset Description
LSV is a multi-view video dataset of wet-lab biology experiments, captured from both first-person (XMglass smart glasses) and third-person (DJI action camera) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors.
The dataset is designed for research on:
Protocol compliance… See the full description on the dataset page: https://huggingface.co/datasets/YinkaiW/LSV.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ty-li/Obstacle-Detection-Dataset-YOLO.new-york-layoffs-warn-act-notices-daily
New York WARN Act layoff notices — every filing we hold since 2001, one CSV, rebuilt daily
6,515 New York WARN notices — every one this dataset holds, back to 2001 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-25
· state source last checked 2026-09-24T14:07Z · official source: New York Department of Labor — WARN notices.
New York employers must file a WARN Act notice with the state before a qualifying
mass layoff or plant… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/new-york-layoffs-warn-act-notices-daily.
