CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.8k downloads3y agoHugging Face02Anthropic /enabling-independent-research Overview This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude". Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/enabling-independent-research.tabular1K<n<10K36 likes1.5k downloads28d agoHugging Face03manycore-research /SpatialLM-Dataset SpatialLM Dataset The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.3d100K<n<1M15 likes1.4k downloads1y agoHugging Face04ibm-research /claim_stance Dataset Card for Claim Stance Dataset Dataset Summary Claim Stance This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic, as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target, topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation between the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/claim_stance.tabulartext-classification1K<n<10K7 likes610 downloads3y agoHugging Face05macpaw-research /mac-app-store-apps-metadata Dataset Card for Macappstore Applications Metadata 📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule. Mac App Store Applications Metadata sourced by the public API. Curated by: MacPaw Way Ltd. Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.imagetabular-classification10K<n<100K10 likes430 downloads1mo agoHugging Face06THULab /warp_Research Warp Research Dataset (TsFile) Apache TsFile version of GotThatData/warp_Research. Overview Experimental results from warp-field research, focused on the relationship between warp factors, energy efficiency, and field characteristics. Records: ~19,700. Time period: January 2025. Features: 15 variables including derived metrics (warp_factor, expansion_rate, stability_score, max_field_strength, avg_field_strength, energy_efficiency, efficiency_ratio… See the full description on the dataset page: https://huggingface.co/datasets/THULab/warp_Research.tabulartabular-regressionn<1K0 likes226 downloads1mo agoHugging Face07GenData-Research /scientific-verification Scientific Verification Benchmark: NMC Cathodes Dataset summary The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.tabularquestion-answering1K<n<10K0 likes218 downloads6d agoHugging Face08handshake-ai-research /studentbench StudentBench StudentBench: AI and human tutoring yield equivalent GRE learning gains Reproduction code · Project · Files StudentBench measures how well AI tutors help real students learn. We release the study data to reproduce the paper's results and support open research on learning, lesson planning, practice problems, tutoring conversations, engagement and cost. Data were collected in summer 2026. Two studies Student learning. 2,383 participants completed 2,469… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/studentbench.tabularn<1K0 likes148 downloads7h agoHugging Face09ibm-research /LLMFineTuningBench Dataset Card for LLMFineTuningBench A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.tabulartabular-regression10K<n<100K3 likes147 downloads16d agoHugging Face10DXRG /dxap-research-workflow-excerpts DXAP historical workflow excerpts and selection aggregates Three detailed historical case studies: 29 parent/reconciliation event rows, 18 sanitized tool-request/response projections, five starting-state records, two research-child event records and nine proposed scenario questions. It also includes the complete 15-row selection-rank table already published with arXiv:2609.05663v1. These are different views of the same selected material, not independent sample counts to add… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dxap-research-workflow-excerpts.tabularn<1K0 likes141 downloads13d agoHugging Face11gong-io-research /call-playbookgated ☎️ &nbsp;The Call Playbook Dataset Real-world B2B sales conversations for text classification A dataset by Gong.io Research Annotated samples drawn from anonymized enterprise sales conversations across 5 binary classification tasks. 📄 Read the paper (ACL Anthology) &nbsp;·&nbsp; arXiv 🗂️ Dataset Summary The Call Playbook Dataset contains annotated samples from real enterprise sales conversations across 5 binary classification tasks… See the full description on the dataset page: https://huggingface.co/datasets/gong-io-research/call-playbook.tabulartext-classification1K<n<10K5 likes132 downloads3mo agoHugging Face12fliarbi /urban-heat-research-corpus Urban Heat Research Corpus (UHRC) v1.0 What does the world study, invent and report about urban heat? This dataset puts three records of the same problem side by side: 20,422 research papers on urban heat islands and extreme heat in cities (1990–2025) with the claims their abstracts make, 106,458 news articles about heat (2021–2025) coded for 51 subjects, framings and terms, and 4,123 patent families for heat-mitigation technologies (2006–2024) — plus supplementary tables on the… See the full description on the dataset page: https://huggingface.co/datasets/fliarbi/urban-heat-research-corpus.imagetext-classification100K<n<1M0 likes130 downloads4d agoHugging Face13NicolaiSivesind /ChatGPT-Research-Abstracts ChatGPT-Research-Abstracts This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts. A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.tabulartext-classification10K<n<100K5 likes120 downloads3y agoHugging Face14GeoGPT-Research-Project /GeoGPT-QA GeoGPT-QA Dataset: A Large-scale Geoscience QA Dataset for Supervised Fine-tuning of LLMs 1. Dataset Description We introduce GeoGPT-QA Dataset, a large-scale synthetic question–answer (QA) corpus developed to support supervised fine-tuning (SFT) of geoscience foundation models. The dataset is derived from open-access geoscience publications distributed under the CC BY license. Using an automated data synthesis pipeline, we generated professional QA pairs from article… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA.tabular10K<n<100K28 likes107 downloads1y agoHugging Face15dementor-research /dementor-matrix-responses Dementor — matrix model responses Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study. Companion to: Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan) Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO / self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant). Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.tabulartext-generation1K<n<10K0 likes101 downloads2mo agoHugging Face16Uris001 /equity-research-dataset AI-Powered Equity Research — Synthetic Analyst Notes (v2) Parts 1 & 2 of an end-to-end AI Equity-Research Platform — synthetic generation (Part 1) and a decision-driven EDA + feature-engineering study (Part 2). Every claim below is a measured number printed by the notebooks, not an assumption. 0. The question this dataset answers What factors determine whether an analyst note is bullish, neutral or bearish — and does the quantitative space still behave like a… See the full description on the dataset page: https://huggingface.co/datasets/Uris001/equity-research-dataset.imagetext-classification1K<n<10K0 likes89 downloads2mo agoHugging Face17DXRG /dx-terminal-pro-research-aggregates DX Terminal Pro research aggregates Nine published research records from DX Research Group's Terminal Pro work. These are small aggregate evidence tables. They contain no participant-level decision logs or training trajectories. The two configurations have different units of analysis: Configuration Records What a row represents market_behavior 4 A reported event or token-window aggregate from the bounded 21-day real-capital Terminal Pro deployment… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dx-terminal-pro-research-aggregates.tabularn<1K0 likes86 downloads13d agoHugging Face18hivex-research /hivex-leaderboard-datatabularn<1K0 likes66 downloads2y agoHugging Face19itseffi /epfl-enterprise-osai-adoption-research-data EPFL Enterprise Open-Source AI Adoption Research Dataset Dataset Summary This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption. Dataset Structure This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.tabulartext-classificationn<1K0 likes66 downloads1y agoHugging Face20ciol-research /multilevel-legal-reasoning Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai. 🧭 Purpose and Scope The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.tabulartext-generationn<1K7 likes64 downloads1y agoHugging Face21Chainsaw4503 /enabling-independent-research Overview This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude". Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Chainsaw4503/enabling-independent-research.tabular1K<n<10K0 likes57 downloads22d agoHugging Face22foundinghires /ai_researchers_atlas Atlas of AI Research Where the people who do AI research are, what they work on, and when they entered the field. Four tables of counts, measured on an index of 790 725 researchers read off 496 295 AI papers. These are aggregates. The dataset holds no names, no affiliations at the person level, no contact details and no individual dates. Every number is a count of people, and the smallest cell is a count of people on one subject in one country. Built and published by Founding… See the full description on the dataset page: https://huggingface.co/datasets/foundinghires/ai_researchers_atlas.tabular1K<n<10K1 likes56 downloads13h agoHugging Face23researchaudio /apple-speechanalyzer-vs-whisper-cpp-mac Apple SpeechAnalyzer vs whisper.cpp on Mac Four complete speech-recognition benchmark runs over the same deterministic 40-speaker LibriSpeech test-clean snapshot: Engine Model path WER CER Repeated median post-speech latency Repeated p95 Apple SpeechAnalyzer progressiveTranscription on macOS 26.5 1.98% 1.02% 125–132 ms 194–201 ms whisper.cpp server 1.8.4 · ggml-small.en 4.28% 1.79% 122–125 ms 152–161 ms Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.tabularautomatic-speech-recognitionn<1K0 likes55 downloads2mo agoHugging Face24Janssen323 /Research_Data_HMR Research_Data_HMR This repository contains the research data file and the corresponding formal PLS-SEM analysis code. File Description File Description Research_Data.xlsx Research data (Excel format) Formal_Analysis_PLS-SEM.R Formal PLS-SEM analysis code (R language) Data Description Data file: Research_Data.xlsx Analysis software: R Analysis method: PLS-SEM (Partial Least Squares Structural Equation Modeling) Note. Column "Q21"… See the full description on the dataset page: https://huggingface.co/datasets/Janssen323/Research_Data_HMR.tabularn<1K0 likes55 downloads22d agoHugging Face25ibm-research /popqa-tp Dataset Card for "popqa-tp" Dataset Summary PopQA-TP (PopQA Templated Paraphrases) is a dataset derived from PopQA (https://huggingface.co/datasets/akariasai/PopQA), created for the paper "Predicting Question-Answering Performance of Large Language Models through Semantic Consistency". PopQA-TP takes each question in PopQA and paraphrases it using each of several manually-created templates specific to each question category. The paper investigates the relationship… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/popqa-tp.tabular100K<n<1M2 likes52 downloads3y agoHugging Face26iservice /predator-ai-research PREDATOR AI Research Dataset 50K+ AI/ML research papers from arXiv, NeurIPS, StackExchange Curated AI research data including papers from arXiv, NeurIPS, ICML, and StackExchange discussions. Each entry includes source identifier, title, source platform, domain tags, value scores, and quality tier. Fields Column Description id ArXiv ID, paper DOI, or StackExchange question ID title Paper title or question summary source Platform (arXiv, NeurIPS, ICML… See the full description on the dataset page: https://huggingface.co/datasets/iservice/predator-ai-research.tabulartext-classification10K<n<100K0 likes49 downloads12d agoHugging Face27ibm-research /SocialStigmaQA-JA SocialStigmaQA-JA Dataset Card It is crucial to test the social bias of large language models. SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models. Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations. Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.tabularquestion-answering10K<n<100K4 likes47 downloads2y agoHugging Face28GenData-Research /condition-provenance-40 Condition provenance in scientific claim verification (40 claims) Summary Forty claims about NMC811 cathodes, each paired with one open-access source paper. A model decides whether the paper's measurements match every condition in the claim, and answers supported, not_supported, or no_comparable_evidence. In 12 of the 40 claims the paper never states the experimental condition the claim turns on, so the keyed answer is no_comparable_evidence, while the other 28… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/condition-provenance-40.tabularn<1K0 likes45 downloads2d agoHugging Face29SwagMessiah100 /enabling-independent-research Overview This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude". Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/SwagMessiah100/enabling-independent-research.tabular1K<n<10K0 likes40 downloads23d agoHugging Face30nbawali4 /enabling-independent-research Overview This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude". Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/nbawali4/enabling-independent-research.tabular1K<n<10K0 likes40 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.