datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tracker-pov
Eidon Tracker POV
1,274 hours of egocentric video paired with 7-point IMU arm tracking, recorded during ordinary household work.
Contributors wore a head-mounted camera and a seven-sensor IMU harness while doing real chores in their own homes: laundry, cleaning, dishes, cooking. Each recording pairs first-person video with 24 Hz orientation data for both hands, both forearms, both upper arms, and the chest.
This is a complete, final release. Eidon AI (Solidic Labs Inc) has wound… See the full description on the dataset page: https://huggingface.co/datasets/eidon-ai/tracker-pov.modelseigen-face-dataset-256
EigenFace-256 Dataset Construction
Training a robust face embedding model requires a diverse dataset with multiple images per identity—capturing variations in angles, lighting, expressions, and age. Using real-world data for this purpose poses several challenges:
Privacy and Ethics: Real-world images can lead to legal and ethical complications.
Bias and Imbalance: Datasets based on real images may lack diversity.
Data Labeling Complexity: Annotating large datasets is… See the full description on the dataset page: https://huggingface.co/datasets/tropos-labs/eigen-face-dataset-256.eikos-arena
Eikos vs Jev: live paper-trading arena on Hyperliquid
Follow it live: https://eikos-arena.vercel.app
A 72-hour experiment. Two decision models trade 14 Hyperliquid markets (crypto, tokenized stocks, indices and
commodities) with $10,000 of simulated money each. Every 5 minutes both get the same market data and the same 28
questions: for each market, long, flat or short for the next 5 minutes, and "will the price be higher in 5 minutes?".
Prices and fees are real (Hyperliquid's… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/eikos-arena.HumDial-EIBench
HumDial-EIBench
A human-recorded multi-turn emotional intelligence benchmark for audio language models.
Figure: Three-stage data pipeline and the four evaluation tasks in HumDial-EIBench.
HumDial-EIBench is designed to evaluate whether audio language models (ALMs) truly understand emotion in speech, rather than relying on text transcription shortcuts.
The benchmark is built from authentic human-recorded dialogues from the ICASSP 2026 HumDial Challenge and… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/HumDial-EIBench.xvr-data⚠️ IF YOU USE THESE DATASETS, PLEASE CITE THE ORIGINAL PAPERS AND RESPECT THEIR ORIGINAL LICENSES. ⚠️
xvr-data contains DICOM/NIfTI versions of the DeepFluoro, Femur, and Ljubljana datasets.
Paper: Rapid patient-specific neural networks for intraoperative X-ray to volume registration
Code: https://github.com/eigenvivek/xvr
Citation
If you use the DeepFluoro dataset (CC BY-NC 4.0), please cite:
@article{grupp2020automatic,
title={Automatic annotation of hip anatomy in… See the full description on the dataset page: https://huggingface.co/datasets/eigenvivek/xvr-data.hnm-fashion-recommendations-data
Dataset Rekomendasi Fashion H&M
Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam.
Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.tracker-pov-imu
Eidon Tracker POV: IMU
24 Hz orientation and motion data from a seven-point IMU harness, paired with the egocentric
video in eidon-ai/tracker-pov.
One row per (recording, timestamp, body slot), roughly 780 million rows. Join to the video
metadata on recording_id.
This repo holds the sensor data only. There is no video here.
The release sits in three places:
Contents
Size
tracker-pov
the 13,451 MP4s and metadata.parquet
9.05 TB
this repo… See the full description on the dataset page: https://huggingface.co/datasets/eidon-ai/tracker-pov-imu.EIDSeg
EIDSeg: A Pixel-Level Semantic Segmentation Dataset for Post-Earthquake Damage Assessment from Social Media Images
EIDSeg is a large-scale post-earthquake infrastructure damage segmentation dataset collected from nine major earthquakes (2008–2023).This repository provides the raw dataset in CVAT XML format, along with the corresponding images organized by split.It is intended to be used together with our official codebase for parsing XML annotations and training segmentation models.… See the full description on the dataset page: https://huggingface.co/datasets/huilihuang413/EIDSeg.Salesforce-xlam-function-calling-60kCC-OCR-V2
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
Dataset Summary
CC-OCR V2 is a comprehensive and challenging OCR benchmark tailored to real-world document processing. It focuses on practical enterprise document processing tasks and incorporates hard and corner cases that are critical yet underrepresented in prior benchmarks.
The dataset comprises 7,093 high-difficulty samples covering 5 major OCR-centric tracks: Text… See the full description on the dataset page: https://huggingface.co/datasets/Eioss/CC-OCR-V2.toxo_mitoeiken-pdfcross-channel-toxoplasma-from-cellmask
Cross-channel toxoplasma from cellmask dataset
3030 paired fields. images/ is the input channel, masks_pv/ the instance-labelled
parasitophorous-vacuole ground truth (same stem = same field), regenerated with the promoted
PV model cpsam_v2_toxo_r5.
Train/test annotation: fields.csv gives name,split,n_objects for every field -
2567 train (31637 objects) / 463 test (6116 objects).
The split is by well (split_by_well.csv), so no well contributes to both sides.
Prepared with spaCR… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/cross-channel-toxoplasma-from-cellmask.spacr_settingseis-subset50
EIS-Subset50: 1970s U.S. Environmental Impact Statements
A 50-document test subset of scanned 1970s U.S. federal Environmental Impact
Statements (EIS) from the Northwestern University Library collection, built to
test whether current models can handle this kind of data: very long documents
(53–590 pages each), degraded microfilm scans, dense bureaucratic text, and
figure-heavy sections (maps, route corridors, site plans, photos).
Contents at a glance
Component
Size
What it… See the full description on the dataset page: https://huggingface.co/datasets/Windsao/eis-subset50.eiyuuoubuwokiwamerutametenseisuminianime
Bangumi Image Base of Eiyuuou, Bu Wo Kiwameru Tame Tenseisu Mini Anime
This is the image base of bangumi Eiyuuou, Bu wo Kiwameru Tame Tenseisu Mini Anime, we detected 110 characters, 5551 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/eiyuuoubuwokiwamerutametenseisuminianime.AstroCLIPtoxoplasma-plaque-well-detector-dataset
Toxoplasma plaque-assay well detector dataset (v4)
YOLO-format images and labels (labels/*.txt: class cx cy w h, normalised), split train / val / test.
split.csv lists every image with its set, source and box count. Empty label files are confirmed
negatives (no well). Used to train einarolafsson/toxoplasma-plaque-well-detector-yolo26.
eigendata-demo-data
EigenData-CLI Demo Datasets
Free, ready-to-use demo samples of agent evaluation and training datasets generated by EigenData-CLI. Each dataset spans a different domain and task style — multi-turn customer service, long-horizon tool use on a simulated laptop, enterprise operations across many SaaS systems, professional knowledge work, and more.
Every dataset here is a small, individually-verified slice of a larger production corpus. The full corpora — including model-training… See the full description on the dataset page: https://huggingface.co/datasets/jindidi/eigendata-demo-data.GowerStreetDESY3
Gower Street DES Y3 Lensing Tiles
Weak lensing convergence map tiles extracted from the Gower Street N-body simulation suite, processed through a Born-approximation raytracing pipeline with DES Y3 MagLim source n(z) distributions.
Dataset Description
Each sample contains a (4, H, W) convergence map tile covering ~3400 deg², corresponding to 4 DES Y3 MagLim tomographic bins. Tiles are extracted from equatorial HEALPix base faces after harmonic-space filtering and rotation… See the full description on the dataset page: https://huggingface.co/datasets/EiffL/GowerStreetDESY3.eizoukenniwateodasuna
Bangumi Image Base of Eizouken Ni Wa Te O Dasu Na!
This is the image base of bangumi Eizouken ni wa Te o Dasu na!, we detected 17 characters, 1057 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/eizoukenniwateodasuna.Agilex_Cobot_Magic_classify_objects_eight
Agilex_Cobot_Magic_classify_objects_eight
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 197
Total Frames: 337837
FPS: 30
Dataset Size: 4.51 GB
Robot Name: Agilex_Cobot_Magic
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_classify_objects_eight.DESISpectra datset from DESI.eiyuukyoushitsu
Bangumi Image Base of Eiyuu Kyoushitsu
This is the image base of bangumi Eiyuu Kyoushitsu, we detected 67 characters, 6096 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/eiyuukyoushitsu.groundtruth-dynamic-benchmarking-submissions
Groundtruth Dynamic Benchmarking — Geology — Submissions
Community-submitted evaluation runs against the groundtruth-dynamic-benchmarking geology rubrics, feeding the leaderboard. We are currently running two tracks: model benchmarking (comparing different models with no special harness) and harness benchmarking (comparing different harnesses using a single standard model - GLM 4.7).
Each submission is a pointwise rubric score: one model, scored 0–10 per question against a… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking-submissions.nohelmetnetdesi_legacysurvey_xmatch
DESI × Legacy Survey cross-match
95,895 galaxies, each carrying three modalities of the same object:
DESI PROVABGS value-added properties (*_provabgs): redshift Z_HP_provabgs,
stellar mass LOG_MSTAR_provabgs, SFR, MCMC posteriors, magnitudes, ...
DESI spectrum (spectrum): flux / ivar / wavelength / mask / LSF, plus DESI
photometry (*_spec)
Legacy Survey DR10 (DECaLS south) imaging: image (160×160 cutouts in
g,r,i,z with ivar + mask planes), rgb / object_mask / blobmodel PNGs… See the full description on the dataset page: https://huggingface.co/datasets/EiffL/desi_legacysurvey_xmatch.illyasviel_von_einzbern_fgo
Dataset of illyasviel_von_einzbern/イリヤスフィール・フォン・アインツベルン/伊莉雅丝菲尔·冯·爱因兹贝伦 (Fate/Grand Order)
This is the dataset of illyasviel_von_einzbern/イリヤスフィール・フォン・アインツベルン/伊莉雅丝菲尔·冯·爱因兹贝伦 (Fate/Grand Order), containing 500 images and their tags.
The core tags of this character are long_hair, red_eyes, white_hair, hair_between_eyes, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/illyasviel_von_einzbern_fgo.eighteenth_century_french_novels
General information
This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University.
For the dataset in XML/TEI see the GitHub repository of the project.
Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800)
This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.
