CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01manycore-research /SpatialLM-Testset SpatialLM Testset Project page | Paper | Code We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos. Folder Structure Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.3dn<1K60 likes1.6k downloads1y agoHugging Face02manycore-research /SpatialLM-Dataset SpatialLM Dataset The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.3d100K<n<1M15 likes1.4k downloads1y agoHugging Face03ai4bharat /MANGO MANGO: A Corpus of Human Ratings for Speech MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages. Key Features: 255,150 human ratings of TTS-generated outputs and ground-truth human speech. Covers two major Indian languages: Hindi & Tamil, and English. Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.audiotext-to-speech10K<n<100K6 likes1.3k downloads1y agoHugging Face04CaryxAI /everyday-manipulation-3d-raw Everyday Manipulation 3D (raw RGB-D) 1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings Chest-mounted iPhone Pro capture of everyday two-handed manipulation by CaryX AI. Clips were recorded with Record3D, an iOS app that captures the iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the audio track removed; the sensor streams are unmodified: synchronised RGB, metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.tabularrobotics1K<n<10K2 likes1.3k downloads2mo agoHugging Face05ManikaSaini /zomato-restaurant-recommendationtext10K<n<100K4 likes1.1k downloads9mo agoHugging Face06apol /ai-election-manipulation-cases AI, Elections and Agency Transfer Evidence Index Version 0.4.4 · released 21 August 2026 · research cutoff 12 August 2026 The dataset contains 6 documented-manipulation records, not 1,087 cases. Read the counts in this order: 1,087 relational rows -> 64 catalogue entries -> 10 core records -> 8 incident-eligible records -> 6 documented-manipulation records The other two incident-eligible records are transparent contested-use… See the full description on the dataset page: https://huggingface.co/datasets/apol/ai-election-manipulation-cases.text1K<n<10K0 likes776 downloads1mo agoHugging Face07mansoorbaloch /chimera-bench CHIMERA-Bench v1.0 A unified benchmark for epitope-specific antibody CDR sequence-structure co-design. Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop) Code: github.com/mansoorbaloch/chimera-bench Dataset Summary Property Value Complexes 2,922 PDB structures 2,721 Pre-computed features 2,941 .pt files Splits 3 (epitope-group, antigen-fold, temporal) Numbering schemesIMGT, Chothia Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.tabularother1K<n<10K0 likes668 downloads4mo agoHugging Face08ManBib /Discord-Unveiled-Extracted Discord Unveiled - Filtered Dataset This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset. Data Processing The data has been processed to: Convert JSON data to CSV format. Remove messages from bots. Filter out messages containing only URLs, mentions, channels or discord emojis. Filter out messages that are not in English using a FastText language identification model. Data Fields The CSV files in… See the full description on the dataset page: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.texttoken-classification100M<n<1B2 likes623 downloads1y agoHugging Face09Cnass /mangaimagen<1K0 likes478 downloads2mo agoHugging Face10manhdungcr7 /dataset_viahe_vqa Vietnamese Sidewalk Violation VQA (vi phạm vỉa hè) Bộ dữ liệu Hỏi–Đáp trên ảnh (VQA) tiếng Việt đầu tiên về hành vi lấn chiếm/vi phạm sử dụng vỉa hè, gắn với căn cứ pháp lý Nghị định 168/2024/NĐ-CP. Xây dựng cho đồ án môn học SE365 (Trường ĐH Công nghệ Thông tin, ĐHQG-HCM). Ảnh: 6.349 ảnh Cặp hỏi–đáp: 19.474 Ngôn ngữ: Tiếng Việt Câu hỏi: 10 câu cố định, 4 loại (yes/no, what đa nhãn, đếm số lượng, không gian) IAA (2 người gán nhãn độc lập): macro-κ tăng 0,736 → 0,854 qua 3 vòng… See the full description on the dataset page: https://huggingface.co/datasets/manhdungcr7/dataset_viahe_vqa.imagevisual-question-answering10K<n<100K0 likes322 downloads2mo agoHugging Face11tmskss /linux-man-pages-tldr-summarized Dataset Card for linux-man-pages-tldr-summarized Dataset Summary This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages. Supported Tasks This dataset should be used to fine-tune language models for summarization tasks. textsummarizationn<1K9 likes311 downloads3y agoHugging Face12samansmink /test_many_filesn<1K0 likes305 downloads2y agoHugging Face13APProjects /us-factory-plant-closings-manufacturing-layoffs-warn-act-notices-daily US factory and plant closings — the actual WARN Act manufacturing filings, rebuilt every day Last rebuilt: 2026-09-24. 4,835 layoff and closure notices filed by factories and industrial plants, auto assembly and parts makers, food and beverage processors and packers, metal, plastics, paper, textile, furniture and electronics producers, and the industrial suppliers that close with them with US state labor departments — 616,065 workers, 2,773 employers, 48 states, 1989–2027. 1,507… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-factory-plant-closings-manufacturing-layoffs-warn-act-notices-daily.texttabular-classification1K<n<10K0 likes301 downloads2h agoHugging Face14Thoria /mandarin-most-common-words-tr-en Mandarin Most Common Words (TR-EN) Overview The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis. This dataset was created by Stephanie Liu and Kamil Murat Yilmaz. Dataset Content The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.text1K<n<10K2 likes169 downloads5mo agoHugging Face15Edoh /manim_pythontextn<1K20 likes158 downloads3y agoHugging Face16manu /french-triviatextn<1K0 likes151 downloads3y agoHugging Face17Paul /hatecheck-mandarin Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-mandarin.tabulartext-classification1K<n<10K7 likes146 downloads4y agoHugging Face18PIIR /ManCAR Amazon Reviews 2023 (7 Categories, Post-processed) Overview This dataset is a curated and post-processed subset of Amazon Reviews 2023. We select 7 product categories and apply a standard preprocessing pipeline widely used in sequential recommendation research. We adopt the official absolute-timestamp split provided by the corpus. Included Categories CDs_and_Vinyl Video_Games Toys_and_Games Musical_Instruments Grocery_and_Gourmet_Food Arts_Crafts_and_Sewing… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/ManCAR.tabular1M<n<10M1 likes137 downloads7mo agoHugging Face19manueltonneau /turkish-hate-speech-supersetgated Turkish Hate Speech Superset This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.tabulartext-classification10K<n<100K2 likes135 downloads2y agoHugging Face20aumghag /Data-Analytics-Digital-Marketing-Project-Management-QA_DBtextquestion-answeringn<1K4 likes135 downloads2y agoHugging Face21drelhaj /Arabic-news-and-management-corpus Arabic Management, Economics & Financial News Corpus (1,200 Articles) This corpus contains 1,200 Arabic news and management articles drawn from three distinct domains. It was originally compiled as part of research into Arabic Corpus Linguistics, management communication, financial discourse and domain-specific NLP. Both plain text and POS-tagged versions are available. The dataset has been widely used in teaching and research, including the King Saud University book Corpus… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-news-and-management-corpus.texttext-classification1K<n<10K0 likes122 downloads10mo agoHugging Face22MangoGoes /douban_movie_info该数据集为豆瓣电影信息维表。 更多信息请参考文章《数据获取:豆瓣电影信息爬取》。 image10K<n<100K5 likes118 downloads3y agoHugging Face23manueltonneau /spanish-hate-speech-supersetgated Spanish Hate Speech Superset This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available or could be retrieved with the Twitter API focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.tabulartext-classification10K<n<100K6 likes97 downloads2y agoHugging Face24Manuel /sentencias-corte-cons-colombia-1992-2021sentencias-corte-cons-colombia-1992-2021. 23750 Case law of the Colombia's Corte Constitucional. Each row is a complete text of each case law. 23750 case law from 1992-2021. Columns: ID Texto: Complete text of the sentence text10K<n<100K5 likes93 downloads4y agoHugging Face25bitext /Bitext-wealth-management-llm-chatbot-training-dataset Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes90 downloads2y agoHugging Face26manulife /tam-benchmarks Tasks over Application Manuals (TAM) TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.documenttext-generationn<1K0 likes90 downloads14d agoHugging Face27solution-seeker-as /manywells ManyWells: simulation of multiphase flow in thousands of wells The ManyWells datasets contain simulations of multiphase (gas, oil, water) flow in thousands of wells. The datasets were created and shared by Solution Seeker AS to support research on data-driven methodologies and industrial applications of machine learning and AI. Details Curated and shared by: Solution Seeker AS License: Creative Commons BY-NC 4.0 Code repository: ManyWells GitHub repository Paper:… See the full description on the dataset page: https://huggingface.co/datasets/solution-seeker-as/manywells.tabulartabular-regression1M<n<10M1 likes88 downloads1y agoHugging Face28shangshang /robot-manip ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD This dataset requires following the author to access. How to Access Follow @shangshang on HuggingFace: https://huggingface.co/shangshang Request access by commenting on the dataset page Once approved, you will receive download permissions Usage Agreement For research and educational purposes only Do not redistribute without permission Cite the dataset in your work: @misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/robot-manip.tabularn<1K0 likes86 downloads24d agoHugging Face29Faneissa92 /vehicle-fleet-management Vehicle Fleet Management Dataset (Free Sample) This is a free sample with 2,126 rows. The full dataset has 12,624 rows across 5 tables. Fleet operations data for a simulated delivery company with 60 vehicles across 3 depots. 15,000 trip records, maintenance logs, fuel purchases, and driver assignments over 18 months. Features mileage-based maintenance schedules, fuel efficiency tracking by vehicle type, seasonal route patterns, and two anomalies — a fuel price spike and a… See the full description on the dataset page: https://huggingface.co/datasets/Faneissa92/vehicle-fleet-management.tabulartabular-classification1K<n<10K0 likes82 downloads23d agoHugging Face30manuquadros /brenda-references-datatabular10K<n<100K0 likes72 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.