datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EmbodiedGenDatahttps://huggingface.co/spaces/HorizonRobotics/EmbodiedGen-Gallery-Explorer
X-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.cmevs-erp-eval
CM-EVS: A Coverage-Curated Panoramic RGB-D Dataset for Indoor Scene Understanding
CM-EVS is a curated panoramic RGB-D dataset built under a single principle: maximize the geometric coverage of a 3D scene with the fewest equirectangular (ERP) frames possible. The release is structured as one redistributable Blender indoor data archive plus four license-aware adapter packages that regenerate matched frames locally from upstream sources whose terms forbid redistribution.
v1.0… See the full description on the dataset page: https://huggingface.co/datasets/anon-cmevs-2026/cmevs-erp-eval.e-CARE
Dataset of (Du et al., 2022) (Unofficial reupload)
Abstract
Understanding causality has vital importance for various Natural Language Processing (NLP) applications. Beyond the labeled instances, conceptual explanations of the causality can provide deep understanding of the causal fact to facilitate the causal reasoning process. However, such explanation information still remains absent in existing causal reasoning resources. In this paper, we fill this gap by presenting… See the full description on the dataset page: https://huggingface.co/datasets/12ml/e-CARE.us_election_2024_telegram_distilled
A billion Telegram messages about the 2024 US presidential election
This is a dataset of Telegram messages collected during the 2024 US presidential election. For more details, see https://dl.acm.org/doi/10.1145/3701716.3715297.
~1.03B messages, ~43K chats, ~0.8TB (distilled).
~350M English messages have toxicity- and hate-related scores from the Perspective API. For more details, see https://support.perspectiveapi.com/s/about-the-api-attributes-and-languages?language=en_US.
~350M… See the full description on the dataset page: https://huggingface.co/datasets/leonardoblas/us_election_2024_telegram_distilled.Emmi-Wing
Emmi-Wing Dataset
This repository contains the dataset proposed in Going with the Speed of Sound: Pushing Neural Surrogates into Highly-turbulent Transonic Regimes, presented at the Workshop on ML for the Physical Sciences at NeurIPS 2025.
Our dataset follows widely-used industrial standards:
OpenFOAM-v2506 used for simulations and mesh generation
Steady-state compressible solver (rhoSimpleFoam)
Body-fitted mesh using snappyHexMesh, y^+ values in the range [50 − 200]… See the full description on the dataset page: https://huggingface.co/datasets/EmmiAI/Emmi-Wing.Embodied-Captioning
Embodied Image Captioning – Manually Annotated Test Set
Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning
📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.eccv2026-cad-challenge-data
ECCV 2026 CAD Challenge Data
This challenge is part of the workshop The Path to Manufacturing: Evolving
3D Generation to Intelligent Computer-Aided Design.
Workshop homepage: https://3dgen-cad-workshop.github.io/
Challenge submission Space: https://huggingface.co/spaces/jingwei-xu-00/eccv2026-cad-challenge
Dataset rendering and preparation code (only .step files are required): https://github.com/DavidXu-JJ/eccv2026-cad-challenge-data-render
This repository contains the public… See the full description on the dataset page: https://huggingface.co/datasets/jingwei-xu-00/eccv2026-cad-challenge-data.DrivAerML_subsampled_10x
DrivAerML Subsampled 10x
A subsampled and compressed version of the DrivAerML dataset by Ashton et al. (2024), prepared for convenient use with ML frameworks for automotive aerodynamics tasks such as drag and lift coefficient prediction.
Disclaimer
This dataset is provided by Emmi AI for convenience only, on an "as-is" basis, and without any warranty, express or implied. Emmi AI does not own, and does not claim any ownership or rights over, the underlying data. All… See the full description on the dataset page: https://huggingface.co/datasets/EmmiAI/DrivAerML_subsampled_10x.brazil-2026-electoral-divergence
AFOS — Brazil 2026 Electoral Divergence Dataset
🌐 English · Português · Español
English
Open, auditable daily dataset that cross-references prediction markets (Polymarket) × polling institutes (TSE-registered) × press coverage for Brazil's 2026 presidential cycle, with explicit divergence between sources instead of smoothed averages.
Maintained by AFOS Analytics — open-source civic infrastructure for electoral political-risk intelligence. This is the public… See the full description on the dataset page: https://huggingface.co/datasets/AFOS-Analytics1/brazil-2026-electoral-divergence.sudoku-extreme
Hardest Sudoku Puzzle Dataset V2
This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community.
Dataset Composition
Sources
tdoku benchmarks
enjoysudoku
Easy Puzzles (1.1M)
puzzles0_kaggle
puzzles1_unbiased
puzzles2_17_clue
Hard Puzzles (3.1M)
puzzles3_magictour_top1465
puzzles4_forum_hardest_1905
puzzles6_forum_hardest_1106
ph_2010/01_file1.txt
Dataset Characteristics
All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.a-share-prices
Dataset Card for a-share-prices
Dataset Summary
This is a daily candlestick dataset of A-share 主板 market, covering the period Since January 1, 2005.
It is primarily intended for historical market data research and does not guarantee the frequency of updates for new data.
You can see the latest updated date in the file .last_update_date.
It consists in two files all-prices.csv and calendar.csv.
all-prices.csv is the primary data file(Attention: the prices are 不复权价).… See the full description on the dataset page: https://huggingface.co/datasets/ellendan/a-share-prices.lisbet-exampleshnm-fashion-recommendations-data
Dataset Rekomendasi Fashion H&M
Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam.
Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.awesome-egocentric-atlas
Use this dataset
from datasets import load_dataset
ds = load_dataset("cy0307/awesome-egocentric-atlas", split="train")
print(len(ds), "resources")
print(ds[0])
papers = load_dataset(
"csv",
data_files="https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas/resolve/main/awesome-egocentric-papers.csv",
split="train",
)
print(len(papers), "paper-linked resources")
Each row is one catalogued resource. Columns:
Column
Description
name
Resource name… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas.BaisBenchThis is the dataset container for the Biological AI Scientist Benchmark (BAISBench). It's a benchmark designed to assess AI scientists' ability to generate biological discoveries through data analysis and reasoning with external knowledge.
This benchmark contains two tasks:
Data Process and cell Type Annotation task (BAIS-DPTA): This task includes 15 single-cell datasets to assess AI scientists' ability to annotate cell types, a fundamental challenge in single-cell analysis. To enable… See the full description on the dataset page: https://huggingface.co/datasets/EperLuo/BaisBench.TSP_EXECUTION_RUNSstaining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.Forex_Factory_Calendar
📅 Forex Factory Economic Calendar Dataset (2007-01-01 to 2025-04-07)
This dataset contains a comprehensive archive of macroeconomic calendar events sourced from Forex Factory, spanning from January 1, 2007 to April 7, 2025.Each row captures a specific event with detailed metadata including currency, event type, market impact level, reported values, and descriptive context.
📦 Dataset Summary
Total timespan: 2007-01-01 → 2025-04-07
Format: CSV (UTF-8)
Timezone:… See the full description on the dataset page: https://huggingface.co/datasets/Ehsanrs2/Forex_Factory_Calendar.iSign
iSign: A Benchmark for Indian Sign Language Processing
The iSign dataset serves as a benchmark for Indian Sign Language Processing. The dataset comprises of NLP-specific tasks (including SignVideo2Text, SignPose2Text, Text2Pose, Word Prediction, and Sign Semantics). The dataset is free for research use but not for commercial purposes.
Quick Links
Website: The landing page for iSign
arXiv Paper: Detailed information about the iSign Benchmark.
Dataset on Hugging Face:… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/iSign.ECommerce-Women-Clothing-ReviewsMA_Query_Expansion_MLT26usgs-global-earthquake-catalog
USGS Global Earthquake Catalog
Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service.
Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including:
Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event.
Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.NSRDB_extractPublic domain data extracted from National Solar Radiation Database: https://nsrdb.nrel.gov/data-viewer
doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.EQ-Bench
EQ-Bench
This is the EQ-Bench v2 English dataset, all credit to Samuel J. Paech.
Title: EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
Abstract: https://arxiv.org/abs/2312.06281
EQ-Bench is a benchmark for language models designed to assess emotional intelligence.
Why emotional intelligence? One reason is that it represents a subset of abilities that are important for the user experience, and which isn't explicitly tested by other benchmarks. Another reason… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/EQ-Bench.Raon-OpenTTS-Eval
Raon-OpenTTS-Eval
Technical Report
A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs.
Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.central-bank-exchange-rates
Central Bank Exchange Rates
Official exchange rates published by 103 central banks and 4 tax authorities, as one CSV per institution: 14,612,276 rows, the oldest series from 1914. Refreshed daily from the GitHub source repository.
Every row is the figure the institution itself published for that date: the ECB euro reference rate, the Federal Reserve H.10 table, the Bank of England spot rates, RBI reference rates, PBoC central parity, HMRC monthly rates for VAT, US Treasury… See the full description on the dataset page: https://huggingface.co/datasets/AllRates/central-bank-exchange-rates.
