datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpatialLM-Dataset
SpatialLM Dataset
The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.MANGO
MANGO: A Corpus of Human Ratings for Speech
MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages.
Key Features:
255,150 human ratings of TTS-generated outputs and ground-truth human speech.
Covers two major Indian languages: Hindi & Tamil, and English.
Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.everyday-manipulation-3d-raw
Everyday Manipulation 3D (raw RGB-D)
1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings
Chest-mounted iPhone Pro capture of everyday two-handed manipulation by
CaryX AI. Clips were recorded with
Record3D, an iOS app that captures the
iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the
audio track removed; the sensor streams are unmodified: synchronised RGB,
metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.chimera-bench
CHIMERA-Bench v1.0
A unified benchmark for epitope-specific antibody CDR sequence-structure co-design.
Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop)
Code: github.com/mansoorbaloch/chimera-bench
Dataset Summary
Property
Value
Complexes
2,922
PDB structures
2,721
Pre-computed features
2,941 .pt files
Splits
3 (epitope-group, antigen-fold, temporal)
Numbering schemesIMGT, Chothia
Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.hatecheck-mandarin
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-mandarin.ManCAR
Amazon Reviews 2023 (7 Categories, Post-processed)
Overview
This dataset is a curated and post-processed subset of Amazon Reviews 2023.
We select 7 product categories and apply a standard preprocessing pipeline widely used in sequential recommendation research.
We adopt the official absolute-timestamp split provided by the corpus.
Included Categories
CDs_and_Vinyl
Video_Games
Toys_and_Games
Musical_Instruments
Grocery_and_Gourmet_Food
Arts_Crafts_and_Sewing… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/ManCAR.turkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
spanish-hate-speech-superset
Spanish Hate Speech Superset
This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.manywells
ManyWells: simulation of multiphase flow in thousands of wells
The ManyWells datasets contain simulations of multiphase (gas, oil, water) flow in thousands of wells. The datasets were created and shared by Solution Seeker AS to support research on data-driven methodologies and industrial applications of machine learning and AI.
Details
Curated and shared by: Solution Seeker AS
License: Creative Commons BY-NC 4.0
Code repository: ManyWells GitHub repository
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/solution-seeker-as/manywells.robot-manip
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/robot-manip.vehicle-fleet-management
Vehicle Fleet Management Dataset (Free Sample)
This is a free sample with 2,126 rows. The full dataset has 12,624 rows across 5 tables.
Fleet operations data for a simulated delivery company with 60 vehicles across
3 depots. 15,000 trip records, maintenance logs, fuel purchases, and driver
assignments over 18 months.
Features mileage-based maintenance schedules, fuel efficiency tracking by
vehicle type, seasonal route patterns, and two anomalies — a fuel price
spike and a… See the full description on the dataset page: https://huggingface.co/datasets/Faneissa92/vehicle-fleet-management.brenda-references-datashitto-mania-dic
English Summary
This dataset accompanies our NLP2026 study on language resource design in the RAG era.
Using a 321-episode Japanese dataset derived from the essay series Shitto Mania, we found that structured metadata substantially outperformed full text in the tested reference-retrieval setting.
Key finding: structured metadata alone achieves 11.1× better retrieval performance than full text in TF-IDF retrieval (59.0% vs 5.3% Recall@10, STRUCT queries, n=300). This advantage… See the full description on the dataset page: https://huggingface.co/datasets/samuraijun/shitto-mania-dic.warehouse-inventory-management
Warehouse & Inventory Management Dataset
Abstract
This dataset provides 30,000 simulated warehouse-level observations (10,000 per scenario) of health commodity storage, inventory management, and warehousing performance across three tiers of the pharmaceutical supply chain in sub-Saharan Africa. Each record represents one commodity category assessed at one warehouse during one monthly period. The dataset captures 40+ variables spanning warehouse infrastructure, storage… See the full description on the dataset page: https://huggingface.co/datasets/rishirajpathak/warehouse-inventory-management.vehicle-fleet-management
Vehicle Fleet Management Dataset (Free Sample)
This is a free sample with 2,126 rows. The full dataset has 12,624 rows across 5 tables.
Fleet operations data for a simulated delivery company with 60 vehicles across
3 depots. 15,000 trip records, maintenance logs, fuel purchases, and driver
assignments over 18 months.
Features mileage-based maintenance schedules, fuel efficiency tracking by
vehicle type, seasonal route patterns, and two anomalies — a fuel price
spike and a vehicle… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/vehicle-fleet-management.bridgev2-vita-toykitchen-manifests
BridgeV2 VITA ToyKitchen-like Manifests
This repository contains manifest files for a reconstructed VITA-style BridgeV2 ToyKitchen-like pick-and-place subset.
Source dataset
The source dataset is:
Gaugou/BridgeV2
This repository does not duplicate the original BridgeV2 videos. It provides episode IDs and metadata for selecting the subset from the source dataset.
Split
Train: 2,986 episodes
Test: 287 episodes
Total selected: 3,273 episodes
Selection… See the full description on the dataset page: https://huggingface.co/datasets/praedico/bridgev2-vita-toykitchen-manifests.tourism-wellness-package-datasetai-risk-manager-dataAYDID-audio
AYDID — Audio (gated access)
Segmented 16 kHz mono WAV audio for the AYDID sub-dialectal Yemeni Arabic corpus,
released for non-commercial academic research under a data-use agreement.
Audio is derived from Yemeni broadcast media. To respect source copyright, it is
not publicly redistributed; access is granted to identified researchers who
accept the terms above. Approved users can reproduce the full AYDID pipeline
end-to-end (feature extraction, segmentation, ASR preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/mansoorSaleh/AYDID-audio.arabic-hate-speech-superset
Arabic Hate Speech Superset
This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.EXAONE-4.0-1.2B-Quantization-MMLUmany-hook-a3b841
many-hook-a3b841
Synthetic weather test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/abeyoichi/many-hook-a3b841.financial-man-4bbc37
financial-man-4bbc37
Synthetic weather test data: 52 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Patricia-Davis/financial-man-4bbc37.Well_actually_mansplainingREALSumm
REALSumm: Re-evaluating EvALuation in Summarization
Dataset assembled from https://github.com/neulab/REALSumm with the conversion script:
idx = [1017, 10586, 11343, 1521, 2736, 3789, 5025, 5272, 5576, 6564, 7174, 7770, 8334, 9325, 9781, 10231, 10595, 11351, 1573, 2748, 3906, 5075, 5334, 5626, 6714, 7397, 7823, 8565, 9393, 9825, 10325, 10680, 11355, 1890, 307, 4043, 5099, 5357, 5635, 6731, 7535, 7910, 8613, 9502, 10368, 10721, 1153, 19, 3152, 4303, 5231, 5420, 5912, 6774, 7547, 8001… See the full description on the dataset page: https://huggingface.co/datasets/manu/REALSumm.diamond-price-predictor-logs2
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/manojdec25/diamond-price-predictor-logs2.xauusd
Cleaned XAUUSD Dataset
Dataset Description
This dataset contains cleaned and preprocessed minute-level historical price data for the XAU/USD (Gold vs. US Dollar) pair. The data spans from November 1, 2011, to January 3, 2024, and includes the following columns:
open: The opening price of the minute.
high: The highest price during the minute.
low: The lowest price during the minute.
close: The closing price of the minute.
tickvol: The number of price changes… See the full description on the dataset page: https://huggingface.co/datasets/Manush09876543/xauusd.textile-manufacturing-egocentric-sample
🧵 Textile Manufacturing — Egocentric Video Dataset (Sample)
This dataset is part of a larger collection of egocentric activity datasets by Verbose Tech Labs LLP. If you want the full dataset, or want access to more categories? Get in touch with us:
📞 Phone: +91 7672 000 500
💬 WhatsApp: +91 7672 000 500
📧 Email: Hello@VerboseTechLabs.com
🌐 Website: VerboseTechLabs.com
🔗 More datasets: kaggle.com/verbosetechlabsllp
Dataset Summary
First-person point-of-view… See the full description on the dataset page: https://huggingface.co/datasets/VerboseTechLabs/textile-manufacturing-egocentric-sample.aerogel-structural-manifold-integrity-v0.1Goal
Detect when an aerogel loses structural integrity before visible collapse.
Core idea
Aerogel failure is not a single crack.It is a distortion of the vibration–density–pore manifold.
Three signals must stay coherent:
densityelastic moduluspore network structure
When they decouple, collapse follows.
Inputs
bulk density
nanoindentation modulus
pore size distribution
load cycling
acoustic or strain indicators
Required outputs
structural_coherence_score
manifold_distortion_rate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aerogel-structural-manifold-integrity-v0.1.
