CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datasets-maintainers /dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI textn<1K0 likes20k downloads3y agoHugging Face02Abtinzandi /Obstacle-Detection-Dataset-YOLO ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision 24,326-image, 25-class YOLO dataset for obstacle detection This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.imageobject-detection10K<n<100K16 likes8.3k downloads3mo agoHugging Face03Bai-YT /RAGDOLL The RAGDOLL E-Commerce Webpage Dataset This repository contains the RAGDOLL (Retrieval-Augmented Generation Deceived Ordering via AdversariaL materiaLs) dataset as well as its LLM-automated collection pipeline. The RAGDOLL dataset is from the paper Ranking Manipulation for Conversational Search Engines from Samuel Pfrommer, Yatong Bai, Tanmay Gautam, and Somayeh Sojoudi. For experiment code associated with this paper, please refer to this repository. The dataset consists of 10… See the full description on the dataset page: https://huggingface.co/datasets/Bai-YT/RAGDOLL.textquestion-answering1K<n<10K4 likes4.2k downloads2y agoHugging Face04yzy666 /SVBench Dataset Card for SVBench This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository. Dataset Details Dataset Description SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.textquestion-answering1K<n<10K7 likes3.4k downloads10mo agoHugging Face05yatsbm /NSRDB_extractPublic domain data extracted from National Solar Radiation Database: https://nsrdb.nrel.gov/data-viewer tabular100K<n<1M0 likes2.1k downloads2y agoHugging Face06yatsbm /montreal_firetabular1M<n<10M0 likes2.1k downloads2y agoHugging Face07AmirTrader /YfOptionstabular10M<n<100M0 likes1.7k downloads6mo agoHugging Face08yaful /MAGE MAGE: Machine-generated Text Detection in the Wild 🚀 Introduction Recent advances in large language models have enabled them to reach a level of text generation comparable to that of humans. These models show powerful capabilities across a wide range of content, including news article writing, story generation, and scientific writing. Such capability further narrows the gap between human-authored and machine-generated texts, highlighting the importance of machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/yaful/MAGE.text100K<n<1M17 likes1.6k downloads2y agoHugging Face09Yiyang-Ian-Li /LongDA LongDA Dataset Card Dataset Description LongDA is a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. It features authentic U.S. government survey data with complete, long documentation, testing LLMs' ability to navigate complex real-world datasets before performing analysis. Dataset Summary 505 queries extracted from 30 expert-written publications 17 U.S. national surveys covering health… See the full description on the dataset page: https://huggingface.co/datasets/Yiyang-Ian-Li/LongDA.documentquestion-answeringn<1K1 likes1.4k downloads3mo agoHugging Face10yipyany /ted-translation-decisions-en-zh TED Translation Decision Dataset (EN–ZH 英-简中) 🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁 🧩 Searchable Keywords translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual, semantic nuance, translation rationale dataset, Chinese translation, English translation dataset, word-level translation, interpretability, translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.tabulartranslationn<1K1 likes1.1k downloads16h agoHugging Face11yangwang825 /sicktext1K<n<10K0 likes838 downloads3y agoHugging Face12ysc0034 /spatial457_mcqtextn<1K0 likes786 downloads1y agoHugging Face13yanismiraoui /prompt_injections Dataset Card for Prompt Injections by Yanis Miraoui 👋 Dataset Description This dataset of prompt injections enriches Large Language Models (LLMs) by providing task-specific examples and prompts, helping improve LLMs' performance and control their behavior. Dataset Summary This dataset contains over 1000 rows of prompt injections in multiple languages. It contains examples of prompt injections using different techniques such as: prompt leaking… See the full description on the dataset page: https://huggingface.co/datasets/yanismiraoui/prompt_injections.text1K<n<10K7 likes764 downloads4mo agoHugging Face14ykotseruba /SNAP SNAP Benchmark Code and annotations: [https://github.com/ykotseruba/SNAP] SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings. This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms. SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.imageimage-classification10K<n<100K0 likes748 downloads1y agoHugging Face15yhayashi1986 /PolyOmics PolyOmics PolyOmics is an omics-scale computational materials database containing molecular structures, simulation metadata, and diverse physical properties for more than (10^5) polymeric materials. The database was generated primarily using RadonPy, a fully automated molecular dynamics (MD) simulation platform for polymer materials. PolyOmics was developed through a large-scale collaboration of the RadonPy Consortium to provide foundational computational data for polymer… See the full description on the dataset page: https://huggingface.co/datasets/yhayashi1986/PolyOmics.tabulartabular-regression100K<n<1M19 likes733 downloads21d agoHugging Face16anusfoil /ycuppe-midi YCU-PPE-III: Piano Performance MIDI Dataset MIDI transcriptions of the YCU-PPE-III piano performance dataset (Wang et al.), used for unreferenced Performance MOS (PMOS) prediction in EVPMR. Overview 2,627 MIDI files transcribed from WAV recordings via transkun 13 songs performed by student pianists 2,511 performances with ratings from 3 expert judges (0-100 scale each) Labels: normalized mean score to [0, 1] Splits: 1,757 train / 377 val / 377 test (stratified by song)… See the full description on the dataset page: https://huggingface.co/datasets/anusfoil/ycuppe-midi.tabularaudio-classification1K<n<10K2 likes707 downloads6mo agoHugging Face17yoavgurarieh /BonaFide BonaFide This is a dataset containing ground-truth faithfulness labels for chains of thought (CoTs), used for evaluating CoT faithfulness metrics. The current benchmark results are in the BonaFide benchmark space. Methodology We construct tasks whose outputs reveal which intermediate computations must have produced them, then label CoTs against those computations. Diversionary setting. Each question is given alongside a misleading hint pointing to a random wrong answer.… See the full description on the dataset page: https://huggingface.co/datasets/yoavgurarieh/BonaFide.tabulartext-classification1K<n<10K2 likes658 downloads4mo agoHugging Face18yuanyyaa /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/yuanyyaa/agent-reward-bench.imagerobotics1K<n<10K0 likes588 downloads6mo agoHugging Face19pmoe7 /SP_500_Stocks_Data-ratios_news_price_10_yrsHi folks, Here is a collection of data I have scraped or aggregated for most of the stocks in the S&P 500, including popular ones like Apple (AAPL). It has the following data: Daily news articles and sentiments on those articles collected over the last few years. All quarterly stock fundamentals (ratios) for 10-20 years. Stock price data (daily close) over the last 10-20 years. Use it however you please for PERSONAL USAGE, but if you do leverage it to make some money; just remember me and… See the full description on the dataset page: https://huggingface.co/datasets/pmoe7/SP_500_Stocks_Data-ratios_news_price_10_yrs.tabular100K<n<1M30 likes551 downloads4y agoHugging Face20tarekmasryo /youtube-tiktok-trends-dataset-2025 🎬 YouTube Shorts & TikTok Trends (2025) Author: Tarek MasryoLicense: CC BY 4.0 A structured snapshot of short-form video activity across YouTube Shorts and TikTok during 2025 (Jan–Aug).Built for content intelligence, analytics dashboards, and ML baselines (classification/regression). What’s inside This repository ships: Two loadable dataset configs (via datasets.load_dataset): default → ML-ready table (cleaned + modeling-friendly) raw → raw video-level table (wider… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/youtube-tiktok-trends-dataset-2025.tabulartabular-regression10K<n<100K7 likes505 downloads8mo agoHugging Face21nus-yam /ex-repairtabular1M<n<10M3 likes494 downloads3y agoHugging Face22ygonet /midi-classical-music MIDI Classical Music This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers. The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others. The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions. The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/ygonet/midi-classical-music.text1K<n<10K0 likes492 downloads2mo agoHugging Face23yuchuan123 /AncientDoc AncientDoc: A Benchmark for Chinese Ancient Document Understanding AncientDoc 是第一个专为 中国古籍文档理解 设计的综合基准数据集,涵盖从 OCR 到 知识推理 的多任务评测,旨在推动多模态大模型在古籍场景下的识别、理解与推理能力研究。 数据集简介 数据规模:2,973 页 文献数量:约 100 本 文献类型:14 类(如总集、楚辞体诗、诗文批评、类书、谱录等) 时间跨度:从战国到清代,涵盖多个重要历史时期 任务类型: Page-level OCR:整页文字识别(含竖排、异体字、批注等复杂情况) Vernacular Translation:文言文到现代汉语的同语种翻译 Reasoning-based QA:基于文意的隐性推理问答 Knowledge-based QA:基于文本事实和背景知识的问答 Linguistic Variant QA:文体、修辞与语言风格相关的问答 数据分布 按朝代分布… See the full description on the dataset page: https://huggingface.co/datasets/yuchuan123/AncientDoc.text1K<n<10K6 likes439 downloads1y agoHugging Face24yifanmai /czech_bank_qa CzechBankQA This is a list of SQL queries for a text-to-SQL task over the Czech Bank 1999 dataset. tabularn<1K0 likes438 downloads2y agoHugging Face25kerne-protocol /solana-yield-honesty Solana Honesty Index What each Solana stablecoin product says it pays, next to what it actually paid, measured from a share price rather than from a claim. Snapshot generated 2026-09-24T12:19:33.744Z. Window 30 days. 13 products across 3 protocols, 13 comparable, 0 published but not comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed. product advertised realized gap delivered realized method Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.tabularn<1K0 likes435 downloads17h agoHugging Face26ShafinSI /Obstacle-Detection-Dataset-YOLO ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision 24,326-image, 25-class YOLO dataset for obstacle detection This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ShafinSI/Obstacle-Detection-Dataset-YOLO.imageobject-detection10K<n<100K0 likes415 downloads15d agoHugging Face27yzhllm /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K0 likes407 downloads4mo agoHugging Face28YinkaiW /LSV LSV: LabSuperVision Benchmark Dataset Description LSV is a multi-view video dataset of wet-lab biology experiments, captured from both first-person (XMglass smart glasses) and third-person (DJI action camera) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors. The dataset is designed for research on: Protocol compliance… See the full description on the dataset page: https://huggingface.co/datasets/YinkaiW/LSV.imagevideo-classificationn<1K4 likes396 downloads6mo agoHugging Face29ty-li /Obstacle-Detection-Dataset-YOLO ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision 24,326-image, 25-class YOLO dataset for obstacle detection This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ty-li/Obstacle-Detection-Dataset-YOLO.imageobject-detection10K<n<100K0 likes396 downloads16d agoHugging Face30APProjects /new-york-layoffs-warn-act-notices-daily New York WARN Act layoff notices — every filing we hold since 2001, one CSV, rebuilt daily 6,515 New York WARN notices — every one this dataset holds, back to 2001 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-25 · state source last checked 2026-09-24T14:07Z · official source: New York Department of Labor — WARN notices. New York employers must file a WARN Act notice with the state before a qualifying mass layoff or plant… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/new-york-layoffs-warn-act-notices-daily.texttabular-classification1K<n<10K0 likes394 downloads5h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.