datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NSRDB_extractPublic domain data extracted from National Solar Radiation Database: https://nsrdb.nrel.gov/data-viewer
montreal_fireYfOptionsted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.PolyOmics
PolyOmics
PolyOmics is an omics-scale computational materials database containing molecular structures, simulation metadata, and diverse physical properties for more than (10^5) polymeric materials.
The database was generated primarily using RadonPy, a fully automated molecular dynamics (MD) simulation platform for polymer materials. PolyOmics was developed through a large-scale collaboration of the RadonPy Consortium to provide foundational computational data for polymer… See the full description on the dataset page: https://huggingface.co/datasets/yhayashi1986/PolyOmics.ycuppe-midi
YCU-PPE-III: Piano Performance MIDI Dataset
MIDI transcriptions of the YCU-PPE-III piano performance dataset (Wang et al.), used for unreferenced Performance MOS (PMOS) prediction in EVPMR.
Overview
2,627 MIDI files transcribed from WAV recordings via transkun
13 songs performed by student pianists
2,511 performances with ratings from 3 expert judges (0-100 scale each)
Labels: normalized mean score to [0, 1]
Splits: 1,757 train / 377 val / 377 test (stratified by song)… See the full description on the dataset page: https://huggingface.co/datasets/anusfoil/ycuppe-midi.BonaFide
BonaFide
This is a dataset containing ground-truth faithfulness labels for chains of thought (CoTs), used for evaluating CoT faithfulness metrics. The current benchmark results are in the BonaFide benchmark space.
Methodology
We construct tasks whose outputs reveal which intermediate computations must have produced them, then label CoTs against those computations.
Diversionary setting. Each question is given alongside a misleading hint pointing to a random wrong answer.… See the full description on the dataset page: https://huggingface.co/datasets/yoavgurarieh/BonaFide.SP_500_Stocks_Data-ratios_news_price_10_yrsHi folks,
Here is a collection of data I have scraped or aggregated for most of the stocks in the S&P 500, including popular ones like Apple (AAPL).
It has the following data:
Daily news articles and sentiments on those articles collected over the last few years.
All quarterly stock fundamentals (ratios) for 10-20 years.
Stock price data (daily close) over the last 10-20 years.
Use it however you please for PERSONAL USAGE, but if you do leverage it to make some money; just remember me and… See the full description on the dataset page: https://huggingface.co/datasets/pmoe7/SP_500_Stocks_Data-ratios_news_price_10_yrs.youtube-tiktok-trends-dataset-2025
🎬 YouTube Shorts & TikTok Trends (2025)
Author: Tarek MasryoLicense: CC BY 4.0
A structured snapshot of short-form video activity across YouTube Shorts and TikTok during 2025 (Jan–Aug).Built for content intelligence, analytics dashboards, and ML baselines (classification/regression).
What’s inside
This repository ships:
Two loadable dataset configs (via datasets.load_dataset):
default → ML-ready table (cleaned + modeling-friendly)
raw → raw video-level table (wider… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/youtube-tiktok-trends-dataset-2025.ex-repairczech_bank_qa
CzechBankQA
This is a list of SQL queries for a text-to-SQL task over the Czech Bank 1999 dataset.
solana-yield-honesty
Solana Honesty Index
What each Solana stablecoin product says it pays, next to what it actually
paid, measured from a share price rather than from a claim.
Snapshot generated 2026-09-24T12:19:33.744Z. Window 30 days.
13 products across 3 protocols,
13 comparable, 0 published but not
comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed.
product
advertised
realized
gap
delivered
realized method
Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.LSV
LSV: LabSuperVision Benchmark
Dataset Description
LSV is a multi-view video dataset of wet-lab biology experiments, captured from both first-person (XMglass smart glasses) and third-person (DJI action camera) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors.
The dataset is designed for research on:
Protocol compliance… See the full description on the dataset page: https://huggingface.co/datasets/YinkaiW/LSV.new-york-layoffs-warn-act-notices-daily
New York WARN Act layoff notices — every filing we hold since 2001, one CSV, rebuilt daily
6,515 New York WARN notices — every one this dataset holds, back to 2001 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-25
· state source last checked 2026-09-24T14:07Z · official source: New York Department of Labor — WARN notices.
New York employers must file a WARN Act notice with the state before a qualifying
mass layoff or plant… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/new-york-layoffs-warn-act-notices-daily.RUListening
RUListening: Building Perceptually-Aware Music-QA Benchmarks
Multimodal LLMs, particularly Large Audio Language Models (LALMs), have shown progress in music understanding tasks due to text-only LLM initialization. However, we find that seven of the top ten Music Question Answering (Music-QA) models are text-only models, suggesting these benchmarks rely on reasoning rather than audio perception. To address this limitation, we present RUListening: Robust Understanding through… See the full description on the dataset page: https://huggingface.co/datasets/yongyizang/RUListening.yc-companies-august-2025
Y Combinator Companies Dataset
Dataset Description
This dataset contains information about 5,404 Y Combinator funded companies that have been publicly launched, sourced from the YC-OSS-API.
Dataset Summary
Total Companies: 5,404
Time Range: Summer 2005 - Summer 2025
Update Frequency: Snapshot from August 2025
Source: YC-OSS-API
Dataset Structure
Data Fields
id: Unique identifier for each company
name: Company name… See the full description on the dataset page: https://huggingface.co/datasets/jeffboudier/yc-companies-august-2025.The-Philosophy-Data-Project
About dataset
The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh.
school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable.
sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.asr_benchmark_storeAgentHazard
AgentHazard
A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
🌐 Website | 📊 Dataset | 📄 Paper | 📖 Appendix
🎯 Overview
AgentHazard is a comprehensive benchmark for evaluating harmful behavior in computer-use agents. Unlike traditional prompt-level safety benchmarks, AgentHazard focuses on execution-level failures that emerge through the composition of locally plausible steps across multi-turn, tool-mediated trajectories.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/Yunhao-Feng/AgentHazard.yfinance-stocks-dataset
YFinance Stocks Dataset
This dataset contains historical stock market data from Yahoo Finance ASX (Australian Securities Exchange) extracted using the yfinance library.
Purpose
This dataset is for practice and educational purposes only. It is intended to be used for learning, experimentation, and practice with data analysis, machine learning, and financial data processing workflows. Specifically, it can be used for:
Analyzing and visualizing stock price trends… See the full description on the dataset page: https://huggingface.co/datasets/dimdejesus0320/yfinance-stocks-dataset.Sora100K
Sora100K
This page serves as both the dataset website and the supplementary materials website for the ACM MM 2026 Dataset Track submission.
Quick Navigation
Dataset Overview
Key Statistics
Dataset Structure
Data Source
Access and License
How to Obtain the Videos
Ethical Considerations, Privacy, and Limitations
Supplementary Materials
Loading the Dataset
Citation
Dataset Website
Dataset Overview
Sora100K is a large-scale multimodal… See the full description on the dataset page: https://huggingface.co/datasets/ysicong/Sora100K.BonaFide-Extended
BonaFide (Extended)
This dataset is an extended version of the BonaFide dataset, containing ground-truth CoT faithfulness labels used for evaluating faithfulness metrics. The latter was sampled to balance label types and target models out of this extended dataset.
See BonaFide for a more detailed explanation of the dataset.
Dataset statistics
19,459 labeled rows.
Same 10 models, 13 tasks, and source datasets as the curated subset.
Label distribution… See the full description on the dataset page: https://huggingface.co/datasets/yoavgurarieh/BonaFide-Extended.ecommerce-user-behavior-datamsmarco-yesnoMethaneUnion
MethaneUnion
This dataset product is organized as:
datasets/
temporal_split/
original_scale/{train,test}.csv
120m_GSD/{train,test}.csv
360m_GSD/{train,test}.csv
480m_GSD/{train,test}.csv
960m_GSD/{train,test}.csv
geo_split/
original_scale/{train,test}.csv
120m_GSD/{train,test}.csv
360m_GSD/{train,test}.csv
480m_GSD/{train,test}.csv
960m_GSD/{train,test}.csv
Each CSV contains exactly these columns: id, label, latitude, longitude… See the full description on the dataset page: https://huggingface.co/datasets/yuyao42/MethaneUnion.youtube-comment-sentiment
YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments)
Overview
This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks.
How to use:
import pandas as pd
df =… See the full description on the dataset page: https://huggingface.co/datasets/AmaanP314/youtube-comment-sentiment.ECtHR-NPD
ECtHR-NPD — unified current release
One current dataset, not separate v1.0/v1.1 releases. This author-approved
distribution is available for manual review. It preserves the paper's
14,575-case cohort and original partitions, with documented amount corrections.
It is not a byte-identical copy of the targets used for the original experiments.
Data
data/case_level.csv: 14,575 cases; 33 columns.
data/applicant_level.csv: 44,581 applicant source units; 14 columns.… See the full description on the dataset page: https://huggingface.co/datasets/YanyiPU716/ECtHR-NPD.DPR
Diabetic Patient 30-Day Readmission Dataset
Dataset Summary
This dataset is derived from the Kaggle Diabetic Patients Readmission Prediction dataset. The original dataset contains electronic health record data from diabetic patient hospital encounters and is commonly used for hospital readmission prediction.
This released version organizes the repository into three levels of data:
Raw Data: the original downloaded source files.
Intermediate Data: derived files before… See the full description on the dataset page: https://huggingface.co/datasets/YPL67/DPR.taiwan-dtm-2025-terrarium-z13
2025 年版全臺灣 20 m DTM — Terrarium z13
這個 Dataset 將內政部公開的 2025 年版全臺灣 20 公尺網格數值地形模型(DTM)轉為 ShadeMap 可直接讀取的 Terrarium RGB XYZ tiles。
官方資料來源:https://data.gov.tw/dataset/176927
原始資料授權:政府資料開放授權條款-第 1 版。本 repository 為衍生格式,請保留官方來源與授權資訊。
內容
terrain/13/{x}/{y}.png:256×256 RGB PNG,XYZ / Web Mercator tile addressing。
tile-index.csv:每張 tile 的區域、有效像素比例、來源高程範圍與 Terrarium 量化誤差。
build-summary.json:建置摘要。
source-manifest.json:原始 ZIP/TIFF SHA-256、解析度、範圍與 CRS 決策。… See the full description on the dataset page: https://huggingface.co/datasets/yhzkiki/taiwan-dtm-2025-terrarium-z13.
