CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01manycore-research /SpatialLM-Dataset SpatialLM Dataset The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.3d100K<n<1M15 likes1.9k downloads1y agoHugging Face02ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.7k downloads3y agoHugging Face03manycore-research /SpatialLM-Testset SpatialLM Testset Project page | Paper | Code We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos. Folder Structure Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.3dn<1K60 likes1.6k downloads1y agoHugging Face04Anthropic /enabling-independent-research Overview This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude". Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/enabling-independent-research.tabular1K<n<10K35 likes1.4k downloads26d agoHugging Face05ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes764 downloads2y agoHugging Face06aigrant /taiwan-ly-law-research Taiwan Legislator Yuan Law Research Data Overview The law research documents are issued irregularly from Taiwan Legislator Yuan. The purpose of those research are providing better understanding on social issues in aspect of laws. One may find documents rich with technical terms which could provided as training data. For comprehensive document list check out this link provided by Taiwan Legislator Yuan. There are currently missing document download links in 10th and 9th… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/taiwan-ly-law-research.text1K<n<10K8 likes652 downloads10mo agoHugging Face07ibm-research /claim_stance Dataset Card for Claim Stance Dataset Dataset Summary Claim Stance This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic, as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target, topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation between the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/claim_stance.tabulartext-classification1K<n<10K7 likes577 downloads3y agoHugging Face08macpaw-research /mac-app-store-apps-metadata Dataset Card for Macappstore Applications Metadata 📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule. Mac App Store Applications Metadata sourced by the public API. Curated by: MacPaw Way Ltd. Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.imagetabular-classification10K<n<100K10 likes440 downloads1mo agoHugging Face09ibm-research /AttaQ AttaQ Dataset Card The AttaQ red teaming dataset, consisting of 1402 carefully crafted adversarial questions, is designed to evaluate Large Language Models (LLMs) by assessing their tendency to generate harmful or undesirable responses. It may serve as a benchmark to assess the potential harm of responses produced by LLMs. The dataset is categorized into seven distinct classes of questions: deception, discrimination, harmful information, substance abuse, sexual content, personally… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AttaQ.texttext-generation1K<n<10K23 likes323 downloads3y agoHugging Face10dementor-research /dementor-sft-datatext100K<n<1M0 likes233 downloads2mo agoHugging Face11macpaw-research /UiPad UiPad - UI Parsing and Accessibility Dataset 📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned. Curated by: MacPaw Way Ltd. Language(s): Mostly EN, UA License: MIT Overview UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.imagequestion-answering1K<n<10K16 likes223 downloads1mo agoHugging Face12THULab /warp_Research Warp Research Dataset (TsFile) Apache TsFile version of GotThatData/warp_Research. Overview Experimental results from warp-field research, focused on the relationship between warp factors, energy efficiency, and field characteristics. Records: ~19,700. Time period: January 2025. Features: 15 variables including derived metrics (warp_factor, expansion_rate, stability_score, max_field_strength, avg_field_strength, energy_efficiency, efficiency_ratio… See the full description on the dataset page: https://huggingface.co/datasets/THULab/warp_Research.tabulartabular-regressionn<1K0 likes220 downloads1mo agoHugging Face13GenData-Research /scientific-verification Scientific Verification Benchmark: NMC Cathodes Dataset summary The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.tabularquestion-answering1K<n<10K0 likes213 downloads5d agoHugging Face14macpaw-research /mac-app-store-apps-descriptions Dataset Card for Macappstore Applications Descriptions 📌 Dataset status: static snapshot (no scheduled updates). This dataset is derived from the December 2023 – January 2024 Mac App Store metadata snapshot and reflects the store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule. Mac App Store Applications descriptions extracted from the metadata from the public API. Curated by: MacPaw Way Ltd. Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-descriptions.texttext-classification10K<n<100K4 likes172 downloads1mo agoHugging Face15bytedance-research /veAgentBench VeAgentBench Dataset The VeAgentBench dataset is designed based on specific application scenarios of agents, aiming to test and evaluate the quality of agents generated by full-process agent development frameworks (such as veADK). It focuses on assessing agents' capabilities in tool calling, knowledge base retrieval, memory management, and overall performance. Updates 2025.11.25 First public release of the dataset, containing a total of 484 questions (145 publicly… See the full description on the dataset page: https://huggingface.co/datasets/bytedance-research/veAgentBench.documentn<1K4 likes171 downloads10mo agoHugging Face16ibm-research /LLMFineTuningBench Dataset Card for LLMFineTuningBench A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.tabulartabular-regression10K<n<100K3 likes143 downloads14d agoHugging Face17DXRG /dxap-research-workflow-excerpts DXAP historical workflow excerpts and selection aggregates Three detailed historical case studies: 29 parent/reconciliation event rows, 18 sanitized tool-request/response projections, five starting-state records, two research-child event records and nine proposed scenario questions. It also includes the complete 15-row selection-rank table already published with arXiv:2609.05663v1. These are different views of the same selected material, not independent sample counts to add… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dxap-research-workflow-excerpts.tabularn<1K0 likes137 downloads11d agoHugging Face18gong-io-research /call-playbookgated ☎️ &nbsp;The Call Playbook Dataset Real-world B2B sales conversations for text classification A dataset by Gong.io Research Annotated samples drawn from anonymized enterprise sales conversations across 5 binary classification tasks. 📄 Read the paper (ACL Anthology) &nbsp;·&nbsp; arXiv 🗂️ Dataset Summary The Call Playbook Dataset contains annotated samples from real enterprise sales conversations across 5 binary classification tasks… See the full description on the dataset page: https://huggingface.co/datasets/gong-io-research/call-playbook.tabulartext-classification1K<n<10K5 likes132 downloads3mo agoHugging Face19NicolaiSivesind /ChatGPT-Research-Abstracts ChatGPT-Research-Abstracts This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts. A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.tabulartext-classification10K<n<100K5 likes127 downloads3y agoHugging Face20fliarbi /urban-heat-research-corpus Urban Heat Research Corpus (UHRC) v1.0 What does the world study, invent and report about urban heat? This dataset puts three records of the same problem side by side: 20,422 research papers on urban heat islands and extreme heat in cities (1990–2025) with the claims their abstracts make, 106,458 news articles about heat (2021–2025) coded for 51 subjects, framings and terms, and 4,123 patent families for heat-mitigation technologies (2006–2024) — plus supplementary tables on the… See the full description on the dataset page: https://huggingface.co/datasets/fliarbi/urban-heat-research-corpus.imagetext-classification100K<n<1M0 likes118 downloads3d agoHugging Face21macpaw-research /mac-app-store-apps-release-notes Dataset Card for Macappstore Applications Release Notes 📌 Dataset status: static snapshot (no scheduled updates). This dataset is derived from the December 2023 – January 2024 Mac App Store metadata snapshot and reflects the store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule. Mac App Store Applications release notes extracted from the metadata from the public API. Curated by: MacPaw Way Ltd. Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-release-notes.texttext-generation10K<n<100K5 likes114 downloads1mo agoHugging Face22ibm-research /MermaidSeqBench Dataset Card for MermaidSeqBench Dataset Summary This dataset provides a human-verified benchmark for assessing large language models (LLMs) on their ability to generate Mermaid sequence diagrams from natural language prompts. The dataset was synthetically generated using large language models (LLMs), starting from a small set of seed examples provided by a subject-matter expert. All outputs were subsequently manually verified and corrected by human annotators to ensure… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/MermaidSeqBench.texttext-generationn<1K7 likes111 downloads5mo agoHugging Face23GeoGPT-Research-Project /GeoGPT-CoT-QA GeoGPT-CoT-QA Dataset: A Large-scale Geoscience Chain-of-Thought QA Dataset for Supervised Fine-Tuning of LLMs 1. Dataset Description We introduce GeoGPT-CoT-QA Dataset, a large-scale synthetic question–answer (QA) corpus enriched with chain-of-thought (CoT) reasoning traces, developed to support supervised fine-tuning (SFT) of geoscience reasoning models. The GeoGPT-R1-Preview is specifically fine-tuned using this dataset to enhance its geoscience reasoning capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-CoT-QA.text10K<n<100K9 likes104 downloads10mo agoHugging Face24dementor-research /dementor-matrix-responses Dementor — matrix model responses Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study. Companion to: Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan) Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO / self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant). Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.tabulartext-generation1K<n<10K0 likes101 downloads2mo agoHugging Face25asrd-research /ASRD-Dataset Adversarial Surface-Form Robustness Dataset (ASRD) Anonymous Repository for Double-Blind ReviewNeurIPS 2026 Workshop 1. Dataset Overview Standard safety evaluations of large language models routinely measure model refusal and compliance using canonical plain-text instructions. However, deployed systems frequently encounter non-canonical inputs containing expressive symbols (emojis), character-level substitutions (homoglyphs, leetspeak), structured encodings… See the full description on the dataset page: https://huggingface.co/datasets/asrd-research/ASRD-Dataset.texttext-generation1K<n<10K0 likes95 downloads19d agoHugging Face26Uris001 /equity-research-dataset AI-Powered Equity Research — Synthetic Analyst Notes (v2) Parts 1 & 2 of an end-to-end AI Equity-Research Platform — synthetic generation (Part 1) and a decision-driven EDA + feature-engineering study (Part 2). Every claim below is a measured number printed by the notebooks, not an assumption. 0. The question this dataset answers What factors determine whether an analyst note is bullish, neutral or bearish — and does the quantitative space still behave like a… See the full description on the dataset page: https://huggingface.co/datasets/Uris001/equity-research-dataset.imagetext-classification1K<n<10K0 likes90 downloads2mo agoHugging Face27riotu-lab /Arabic-books-and-research-dataset Arabic reserach and books dataset (ARABD) This dataset is an extracted cleaned text from more than 60K word files with unique arabic texts never published before. Dataset diversity the dataset is diverse from all kind of islamic research: [feqh, hadeeth, tafseer, tahqeeq, ... etc], from new written research to a manuscirpts. dataset size the dataset was more than 11GB but after cleaning (pre-processing) it becase a straight 10GB with less noisy chars.… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Arabic-books-and-research-dataset.texttext-generation10K<n<100K6 likes88 downloads2y agoHugging Face28dementor-research /dementor-matrix-baselinestext10K<n<100K0 likes86 downloads2mo agoHugging Face29DXRG /dx-terminal-pro-research-aggregates DX Terminal Pro research aggregates Nine published research records from DX Research Group's Terminal Pro work. These are small aggregate evidence tables. They contain no participant-level decision logs or training trajectories. The two configurations have different units of analysis: Configuration Records What a row represents market_behavior 4 A reported event or token-window aggregate from the bounded 21-day real-capital Terminal Pro deployment… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dx-terminal-pro-research-aggregates.tabularn<1K0 likes85 downloads11d agoHugging Face30GeoGPT-Research-Project /GeoGPT-QA GeoGPT-QA Dataset: A Large-scale Geoscience QA Dataset for Supervised Fine-tuning of LLMs 1. Dataset Description We introduce GeoGPT-QA Dataset, a large-scale synthetic question–answer (QA) corpus developed to support supervised fine-tuning (SFT) of geoscience foundation models. The dataset is derived from open-access geoscience publications distributed under the CC BY license. Using an automated data synthesis pipeline, we generated professional QA pairs from article… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA.tabular10K<n<100K28 likes81 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.