datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
argument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.enabling-independent-research
Overview
This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude".
Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/enabling-independent-research.SpatialLM-Dataset
SpatialLM Dataset
The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.claim_stance
Dataset Card for Claim Stance Dataset
Dataset Summary
Claim Stance
This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic,
as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target,
topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation between the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/claim_stance.mac-app-store-apps-metadata
Dataset Card for Macappstore Applications Metadata
📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications Metadata sourced by the public API.
Curated by: MacPaw Way Ltd.
Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.warp_Research
Warp Research Dataset (TsFile)
Apache TsFile version of
GotThatData/warp_Research.
Overview
Experimental results from warp-field research, focused on the relationship between warp
factors, energy efficiency, and field characteristics.
Records: ~19,700.
Time period: January 2025.
Features: 15 variables including derived metrics (warp_factor, expansion_rate,
stability_score, max_field_strength, avg_field_strength, energy_efficiency,
efficiency_ratio… See the full description on the dataset page: https://huggingface.co/datasets/THULab/warp_Research.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.studentbench
StudentBench
StudentBench: AI and human tutoring yield equivalent GRE learning gains
Reproduction code · Project · Files
StudentBench measures how well AI tutors help real students learn. We release the study data to reproduce the paper's results and support open research on learning, lesson planning, practice problems, tutoring conversations, engagement and cost.
Data were collected in summer 2026.
Two studies
Student learning. 2,383 participants completed 2,469… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/studentbench.LLMFineTuningBench
Dataset Card for LLMFineTuningBench
A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.dxap-research-workflow-excerpts
DXAP historical workflow excerpts and selection aggregates
Three detailed historical case studies: 29 parent/reconciliation event rows, 18 sanitized tool-request/response projections, five starting-state records, two research-child event records and nine proposed scenario questions. It also includes the complete 15-row selection-rank table already published with arXiv:2609.05663v1. These are different views of the same selected material, not independent sample counts to add… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dxap-research-workflow-excerpts.call-playbook
☎️ The Call Playbook Dataset
Real-world B2B sales conversations for text classification
A dataset by Gong.io Research
Annotated samples drawn from anonymized enterprise sales conversations across 5 binary classification tasks.
📄 Read the paper (ACL Anthology)
·
arXiv
🗂️ Dataset Summary
The Call Playbook Dataset contains annotated samples from real enterprise sales conversations across 5 binary classification tasks… See the full description on the dataset page: https://huggingface.co/datasets/gong-io-research/call-playbook.urban-heat-research-corpus
Urban Heat Research Corpus (UHRC) v1.0
What does the world study, invent and report about urban heat? This dataset puts three records of the same problem side by side:
20,422 research papers on urban heat islands and extreme heat in cities (1990–2025) with the claims their abstracts make,
106,458 news articles about heat (2021–2025) coded for 51 subjects, framings and terms, and
4,123 patent families for heat-mitigation technologies (2006–2024) — plus supplementary tables on the… See the full description on the dataset page: https://huggingface.co/datasets/fliarbi/urban-heat-research-corpus.ChatGPT-Research-Abstracts
ChatGPT-Research-Abstracts
This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts.
A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.GeoGPT-QA
GeoGPT-QA Dataset: A Large-scale Geoscience QA Dataset for Supervised Fine-tuning of LLMs
1. Dataset Description
We introduce GeoGPT-QA Dataset, a large-scale synthetic question–answer (QA) corpus developed to support supervised fine-tuning (SFT) of geoscience foundation models.
The dataset is derived from open-access geoscience publications distributed under the CC BY license. Using an automated data synthesis pipeline, we generated professional QA pairs from article… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA.dementor-matrix-responses
Dementor — matrix model responses
Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study.
Companion to:
Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan)
Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO /
self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant).
Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.equity-research-dataset
AI-Powered Equity Research — Synthetic Analyst Notes (v2)
Parts 1 & 2 of an end-to-end AI Equity-Research Platform — synthetic generation (Part 1) and a
decision-driven EDA + feature-engineering study (Part 2). Every claim below is a measured number
printed by the notebooks, not an assumption.
0. The question this dataset answers
What factors determine whether an analyst note is bullish, neutral or bearish — and does the
quantitative space still behave like a… See the full description on the dataset page: https://huggingface.co/datasets/Uris001/equity-research-dataset.dx-terminal-pro-research-aggregates
DX Terminal Pro research aggregates
Nine published research records from DX Research Group's Terminal Pro work. These are small aggregate evidence tables. They contain no participant-level decision logs or training trajectories.
The two configurations have different units of analysis:
Configuration
Records
What a row represents
market_behavior
4
A reported event or token-window aggregate from the bounded 21-day real-capital Terminal Pro deployment… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dx-terminal-pro-research-aggregates.hivex-leaderboard-dataepfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.enabling-independent-research
Overview
This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude".
Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Chainsaw4503/enabling-independent-research.ai_researchers_atlas
Atlas of AI Research
Where the people who do AI research are, what they work on, and when they entered
the field. Four tables of counts, measured on an index of 790 725 researchers
read off 496 295 AI papers.
These are aggregates. The dataset holds no names, no affiliations at the person
level, no contact details and no individual dates. Every number is a count of
people, and the smallest cell is a count of people on one subject in one country.
Built and published by Founding… See the full description on the dataset page: https://huggingface.co/datasets/foundinghires/ai_researchers_atlas.apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.Research_Data_HMR
Research_Data_HMR
This repository contains the research data file and the corresponding formal PLS-SEM analysis code.
File Description
File
Description
Research_Data.xlsx
Research data (Excel format)
Formal_Analysis_PLS-SEM.R
Formal PLS-SEM analysis code (R language)
Data Description
Data file: Research_Data.xlsx
Analysis software: R
Analysis method: PLS-SEM (Partial Least Squares Structural Equation Modeling)
Note. Column "Q21"… See the full description on the dataset page: https://huggingface.co/datasets/Janssen323/Research_Data_HMR.popqa-tp
Dataset Card for "popqa-tp"
Dataset Summary
PopQA-TP (PopQA Templated Paraphrases) is a dataset derived from PopQA (https://huggingface.co/datasets/akariasai/PopQA), created for the paper "Predicting Question-Answering Performance of Large Language Models
through Semantic Consistency". PopQA-TP takes each question in PopQA and paraphrases it using each of several manually-created templates specific to each question category. The paper investigates the relationship… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/popqa-tp.predator-ai-research
PREDATOR AI Research Dataset
50K+ AI/ML research papers from arXiv, NeurIPS, StackExchange
Curated AI research data including papers from arXiv, NeurIPS, ICML, and StackExchange discussions. Each entry includes source identifier, title, source platform, domain tags, value scores, and quality tier.
Fields
Column
Description
id
ArXiv ID, paper DOI, or StackExchange question ID
title
Paper title or question summary
source
Platform (arXiv, NeurIPS, ICML… See the full description on the dataset page: https://huggingface.co/datasets/iservice/predator-ai-research.SocialStigmaQA-JA
SocialStigmaQA-JA Dataset Card
It is crucial to test the social bias of large language models.
SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models.
Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations.
Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.condition-provenance-40
Condition provenance in scientific claim verification (40 claims)
Summary
Forty claims about NMC811 cathodes, each paired with one open-access source paper. A model decides whether the paper's measurements match every condition in the claim, and answers supported, not_supported, or no_comparable_evidence. In 12 of the 40 claims the paper never states the experimental condition the claim turns on, so the keyed answer is no_comparable_evidence, while the other 28… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/condition-provenance-40.enabling-independent-research
Overview
This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude".
Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/SwagMessiah100/enabling-independent-research.enabling-independent-research
Overview
This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude".
Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/nbawali4/enabling-independent-research.
