datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpatialLM-Testset
SpatialLM Testset
Project page | Paper | Code
We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.SpatialLM-Dataset
SpatialLM Dataset
The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.MANGO
MANGO: A Corpus of Human Ratings for Speech
MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages.
Key Features:
255,150 human ratings of TTS-generated outputs and ground-truth human speech.
Covers two major Indian languages: Hindi & Tamil, and English.
Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.everyday-manipulation-3d-raw
Everyday Manipulation 3D (raw RGB-D)
1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings
Chest-mounted iPhone Pro capture of everyday two-handed manipulation by
CaryX AI. Clips were recorded with
Record3D, an iOS app that captures the
iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the
audio track removed; the sensor streams are unmodified: synchronised RGB,
metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.zomato-restaurant-recommendationai-election-manipulation-cases
AI, Elections and Agency Transfer Evidence Index
Version 0.4.4 · released 21 August 2026 · research cutoff 12 August 2026
The dataset contains 6 documented-manipulation records, not 1,087 cases. Read the counts in this order:
1,087 relational rows -> 64 catalogue entries -> 10 core records
-> 8 incident-eligible records
-> 6 documented-manipulation records
The other two incident-eligible records are transparent contested-use… See the full description on the dataset page: https://huggingface.co/datasets/apol/ai-election-manipulation-cases.chimera-bench
CHIMERA-Bench v1.0
A unified benchmark for epitope-specific antibody CDR sequence-structure co-design.
Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop)
Code: github.com/mansoorbaloch/chimera-bench
Dataset Summary
Property
Value
Complexes
2,922
PDB structures
2,721
Pre-computed features
2,941 .pt files
Splits
3 (epitope-group, antigen-fold, temporal)
Numbering schemesIMGT, Chothia
Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.Discord-Unveiled-Extracted
Discord Unveiled - Filtered Dataset
This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset.
Data Processing
The data has been processed to:
Convert JSON data to CSV format.
Remove messages from bots.
Filter out messages containing only URLs, mentions, channels or discord emojis.
Filter out messages that are not in English using a FastText language identification model.
Data Fields
The CSV files in… See the full description on the dataset page: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.mangadataset_viahe_vqa
Vietnamese Sidewalk Violation VQA (vi phạm vỉa hè)
Bộ dữ liệu Hỏi–Đáp trên ảnh (VQA) tiếng Việt đầu tiên về hành vi lấn chiếm/vi phạm
sử dụng vỉa hè, gắn với căn cứ pháp lý Nghị định 168/2024/NĐ-CP. Xây dựng cho đồ án
môn học SE365 (Trường ĐH Công nghệ Thông tin, ĐHQG-HCM).
Ảnh: 6.349 ảnh
Cặp hỏi–đáp: 19.474
Ngôn ngữ: Tiếng Việt
Câu hỏi: 10 câu cố định, 4 loại (yes/no, what đa nhãn, đếm số lượng, không gian)
IAA (2 người gán nhãn độc lập): macro-κ tăng 0,736 → 0,854 qua 3 vòng… See the full description on the dataset page: https://huggingface.co/datasets/manhdungcr7/dataset_viahe_vqa.linux-man-pages-tldr-summarized
Dataset Card for linux-man-pages-tldr-summarized
Dataset Summary
This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages.
Supported Tasks
This dataset should be used to fine-tune language models for summarization tasks.
test_many_filesus-factory-plant-closings-manufacturing-layoffs-warn-act-notices-daily
US factory and plant closings — the actual WARN Act manufacturing filings, rebuilt every day
Last rebuilt: 2026-09-24. 4,835 layoff and closure notices filed by
factories and industrial plants, auto assembly and parts makers, food and beverage processors and packers, metal, plastics, paper, textile, furniture and electronics producers, and the industrial suppliers that close with them with US state labor departments — 616,065 workers,
2,773 employers, 48 states, 1989–2027.
1,507… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-factory-plant-closings-manufacturing-layoffs-warn-act-notices-daily.mandarin-most-common-words-tr-en
Mandarin Most Common Words (TR-EN)
Overview
The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis.
This dataset was created by Stephanie Liu and Kamil Murat Yilmaz.
Dataset Content
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.manim_pythonfrench-triviahatecheck-mandarin
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-mandarin.ManCAR
Amazon Reviews 2023 (7 Categories, Post-processed)
Overview
This dataset is a curated and post-processed subset of Amazon Reviews 2023.
We select 7 product categories and apply a standard preprocessing pipeline widely used in sequential recommendation research.
We adopt the official absolute-timestamp split provided by the corpus.
Included Categories
CDs_and_Vinyl
Video_Games
Toys_and_Games
Musical_Instruments
Grocery_and_Gourmet_Food
Arts_Crafts_and_Sewing… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/ManCAR.turkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.Data-Analytics-Digital-Marketing-Project-Management-QA_DBArabic-news-and-management-corpus
Arabic Management, Economics & Financial News Corpus (1,200 Articles)
This corpus contains 1,200 Arabic news and management articles drawn from three distinct domains. It was originally compiled as part of research into Arabic Corpus Linguistics, management communication, financial discourse and domain-specific NLP. Both plain text and POS-tagged versions are available.
The dataset has been widely used in teaching and research, including the King Saud University book Corpus… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-news-and-management-corpus.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
spanish-hate-speech-superset
Spanish Hate Speech Superset
This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.sentencias-corte-cons-colombia-1992-2021sentencias-corte-cons-colombia-1992-2021.
23750 Case law of the Colombia's Corte Constitucional.
Each row is a complete text of each case law.
23750 case law from 1992-2021.
Columns:
ID
Texto: Complete text of the sentence
Bitext-wealth-management-llm-chatbot-training-dataset
Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.tam-benchmarks
Tasks over Application Manuals (TAM)
TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.manywells
ManyWells: simulation of multiphase flow in thousands of wells
The ManyWells datasets contain simulations of multiphase (gas, oil, water) flow in thousands of wells. The datasets were created and shared by Solution Seeker AS to support research on data-driven methodologies and industrial applications of machine learning and AI.
Details
Curated and shared by: Solution Seeker AS
License: Creative Commons BY-NC 4.0
Code repository: ManyWells GitHub repository
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/solution-seeker-as/manywells.robot-manip
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/robot-manip.vehicle-fleet-management
Vehicle Fleet Management Dataset (Free Sample)
This is a free sample with 2,126 rows. The full dataset has 12,624 rows across 5 tables.
Fleet operations data for a simulated delivery company with 60 vehicles across
3 depots. 15,000 trip records, maintenance logs, fuel purchases, and driver
assignments over 18 months.
Features mileage-based maintenance schedules, fuel efficiency tracking by
vehicle type, seasonal route patterns, and two anomalies — a fuel price
spike and a… See the full description on the dataset page: https://huggingface.co/datasets/Faneissa92/vehicle-fleet-management.brenda-references-data
