datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/yuanyyaa/agent-reward-bench.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ShafinSI/Obstacle-Detection-Dataset-YOLO.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.LSV
LSV: LabSuperVision Benchmark
Dataset Description
LSV is a multi-view video dataset of wet-lab biology experiments, captured from both first-person (XMglass smart glasses) and third-person (DJI action camera) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors.
The dataset is designed for research on:
Protocol compliance… See the full description on the dataset page: https://huggingface.co/datasets/YinkaiW/LSV.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ty-li/Obstacle-Detection-Dataset-YOLO.yc-companies-august-2025
Y Combinator Companies Dataset
Dataset Description
This dataset contains information about 5,404 Y Combinator funded companies that have been publicly launched, sourced from the YC-OSS-API.
Dataset Summary
Total Companies: 5,404
Time Range: Summer 2005 - Summer 2025
Update Frequency: Snapshot from August 2025
Source: YC-OSS-API
Dataset Structure
Data Fields
id: Unique identifier for each company
name: Company name… See the full description on the dataset page: https://huggingface.co/datasets/jeffboudier/yc-companies-august-2025.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/eziodad/Obstacle-Detection-Dataset-YOLO.Sora100K
Sora100K
This page serves as both the dataset website and the supplementary materials website for the ACM MM 2026 Dataset Track submission.
Quick Navigation
Dataset Overview
Key Statistics
Dataset Structure
Data Source
Access and License
How to Obtain the Videos
Ethical Considerations, Privacy, and Limitations
Supplementary Materials
Loading the Dataset
Citation
Dataset Website
Dataset Overview
Sora100K is a large-scale multimodal… See the full description on the dataset page: https://huggingface.co/datasets/ysicong/Sora100K.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/NivithaChandran/Obstacle-Detection-Dataset-YOLO.taiwan-dtm-2025-terrarium-z13
2025 年版全臺灣 20 m DTM — Terrarium z13
這個 Dataset 將內政部公開的 2025 年版全臺灣 20 公尺網格數值地形模型(DTM)轉為 ShadeMap 可直接讀取的 Terrarium RGB XYZ tiles。
官方資料來源:https://data.gov.tw/dataset/176927
原始資料授權:政府資料開放授權條款-第 1 版。本 repository 為衍生格式,請保留官方來源與授權資訊。
內容
terrain/13/{x}/{y}.png:256×256 RGB PNG,XYZ / Web Mercator tile addressing。
tile-index.csv:每張 tile 的區域、有效像素比例、來源高程範圍與 Terrarium 量化誤差。
build-summary.json:建置摘要。
source-manifest.json:原始 ZIP/TIFF SHA-256、解析度、範圍與 CRS 決策。… See the full description on the dataset page: https://huggingface.co/datasets/yhzkiki/taiwan-dtm-2025-terrarium-z13.cocktails_436
🍸 Cocktails Dataset (436 Recipes + QA)
This dataset consists of two parts:
1. cocktails_436.csv (436 entries)
A structured dataset of 436 cocktail recipes, including:
Basic metadata: name, alcoholic, category, glassType
Recipe details: instructions, ingredients, ingredientMeasures
Multimodal support: drinkThumbnail (image URL), desciription(textual description of the cocktail itself), imageDescription (natural language description of the drink image… See the full description on the dataset page: https://huggingface.co/datasets/yujinyang/cocktails_436.craigslist-used-cars-eda
Craigslist Used Cars and Trucks: EDA
Overview
This dataset and notebook contain an Exploratory Data Analysis (EDA) of real Craigslist used-car listings scraped across the United States.
Main Question: What factors most influence the price of a used car listed on Craigslist?
Target Variable: price — the seller's asking price for each vehicle listing.
About the Dataset
Property
Details
Source
Kaggle — Austin Reese (scraped from Craigslist)
Original… See the full description on the dataset page: https://huggingface.co/datasets/Yoad22/craigslist-used-cars-eda.mad-cars
MAD-Cars: Multi-view Auto Dataset 🚗
Dataset Description
MAD-Cars is a large-scale collection of 360° car videos.
It comprises ~70,000 car instances with diverse brands, car types, colors, and lighting conditions. Each instance contains an average of ~85 frames, with most car instances available at a resolution of 1920x1080. The dataset statistics are presented in the figure below. The data is carefully curated by filtering the frames and entire car instances that… See the full description on the dataset page: https://huggingface.co/datasets/yandex/mad-cars.yfcc15mYFCC15m dataset from https://github.com/openai/CLIP/blob/main/data/yfcc100m.md.
The subset is obtained by filtering the original YFCC100m (yfcc100m_dataset.sql) using the photo ids from https://github.com/openai/CLIP/blob/main/data/yfcc100m.md.
The script to rebuild the data from the original YFCC100m is provided at build_yfcc15m.py.
YKS_book_datasetAU-OPG
AU-OPG: Panoramic Dental Radiographs With Oriented Tooth-Level Annotations
AU-OPG (Ajman University Orthopantomography) is a dataset of 901 panoramic dental radiographs annotated for:
oriented tooth detection;
radiographic diagnosis; and
radiographic-evidence-based treatment planning.
The release contains 7,006 tooth-level annotations. Each annotated tooth has a tooth-aligned oriented bounding box, one diagnostic condition, and one corresponding treatment label. The predefined… See the full description on the dataset page: https://huggingface.co/datasets/YSFF/AU-OPG.open-insect
Description
Open-Insect is curated to benchmark open-set recognition of novel species in biodiversity monitoring, with a focus on insects. It is consists of URLs pointing to image files compiled from the Global Biodiversity Information Facility (GBIF) as well as other metadata such as latitude, longitude, GBIF speciesKey, and label used to train the classifier. Open-Insect partially builds on a subset of the AMI dataset.
Run this download script to download this huggingface… See the full description on the dataset page: https://huggingface.co/datasets/yuyan-chen/open-insect.TS_DATASETSycombinator-companiesLeverage our Y Combinator Companies & Founders Dataset to analyze one of the world’s most influential startup ecosystems.
This dataset contains 5,384 Y Combinator-funded companies along with detailed founder information. Each record includes structured company attributes (such as industry, batch, stage, funding, and status) as well as metadata on founders (names, roles, bios, and social links).
Designed as a comprehensive and high-quality resource, this dataset supports academic research… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/ycombinator-companies.YFCC15M_page_and_download_urls
YFCC15M subset used for VLMs
This dataset contains the ~15M subset of YFCC100M used for training the models in the paper Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP. The metadata provided in this repo contains both the page-urls and image-download-urls for downloading the dataset.
This dataset can be easily downloaded with img2dataset:
img2dataset --url_list yfcc15m_final_split_pageandimageurls.csv --input_format "csv" --output_format… See the full description on the dataset page: https://huggingface.co/datasets/vishaal27/YFCC15M_page_and_download_urls.TrCaptiondouban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
riemannian-generative-decoder
Riemannian generative decoder dataset
This repository contains the data related to the paper "Riemannian generative decoder".
Project Page: https://yhsure.github.io/riemannian-generative-decoder
Code Repository: https://github.com/yhsure/riemannian-generative-decoder
Abstract
Riemannian representation learning typically relies on an encoder to estimate densities on chosen manifolds. This involves optimizing numerically brittle objectives, potentially harming model… See the full description on the dataset page: https://huggingface.co/datasets/yhsure/riemannian-generative-decoder.cc15m_yfcc15myc_startupsTry out the YC Chatbot
open-insect-bci
Description
Open-Insect-BCI is curated to provide a maximally realistic test case for open-set recognition of novel species in biodiversity monitoring. It consists of 197 images recently collected from Barro Colorado Island (BCI), Panama.
It includes 133 species, 59 of which are likely undescribed by science.
Open-Insect-BCI is part of the Open-Insect dataset.
Dataset Configurations and Splits
This dataset is an OOD test set for the Open-Insect C-America region… See the full description on the dataset page: https://huggingface.co/datasets/yuyan-chen/open-insect-bci.vehicle_sales_data--
Human_Aligned_Bench
