datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MuSR
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Creating murder mysteries that require multi-step reasoning with commonsense using ChatGPT!
By: Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett.
View the dataset on our custom viewer and project website!
Check out the paper. Appeared at ICLR 2024 as a spotlight presentation!
Git Repo with the source data, how to recreate the dataset (and create new ones!) here
PHTest🌟 PHTest: Evaluating False Refusals in LLMs
🤖 Auto Red-Teaming
All prompts are generated automatically using a controllable text-generation technique called AutoDAN.
🌐 Diverse Prompts
PHTest introduces false refusal patterns that aren’t present in existing datasets, including prompts that avoid mentioning sensitive words.
⚖️ Harmlessness & Controversial Labeling
Controversial prompts are separately labeled to address the… See the full description on the dataset page: https://huggingface.co/datasets/furonghuang-lab/PHTest.InvoiceBenchmark
InvoiceBenchmark
200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number.
The Pitch
Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.MuSP-Bench
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding Across Score and Performance
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Official benchmark website
Modalities
Each question specifies the minimum source of musical evidence needed to answer it:
S (score): answer from the written score.
P (performance): answer from the performance recording.
S&P (score… See the full description on the dataset page: https://huggingface.co/datasets/bryel-labs/MuSP-Bench.ClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.daily-oracle
Daily Oracle
📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time.
Dataset Details
Question Type: True/False (TF) & Multiple Choice (MC)
Current Version*
Time Span: 2020.01.01 - 2026.07.18
Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.MURAD
A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset
Authors: Serry Sibaee, Yasser Alhabashi, Nadia Sibai, Yara Farouk, Adel Ammar, Sawsan AlHalawani, Wadii Boulila
Overview
MURAD A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset is an open Arabic lexical dataset designed to support research in computational linguistics, lexicography, and Arabic natural language processing (NLP). The dataset contains 96,243… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/MURAD.preguntas-normativa-laboral-rrhh-espana
Normativa laboral y RRHH en España: 300 preguntas con su respuesta y su artículo
Corpus de 300 pares de pregunta y respuesta sobre recursos humanos y normativa laboral española, repartidos en 30 temas y con la referencia normativa en 286 de ellos. Todas las respuestas son autónomas: se entienden sin contexto adicional.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Qué cubre
Tema
Preguntas
Registro horario
42
Inspección de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/preguntas-normativa-laboral-rrhh-espana.all_RD_datasets
RD Dataset With References
This dataset contains Arabic terms and their definitions.The data was extracted and combined from the following sources:
https://huggingface.co/datasets/Basma2423/Arabic-Terminologies-and-Definitions
https://data.mendeley.com/datasets/gxr3j4tdk5/3
https://huggingface.co/datasets/MohamedRashad/arabic-roots
https://arai.ksaa.gov.sa/sharedTask2024/
Each entry consists of:
word
definition
EnrichedMeaningDataset
Dataset Card for Chinese Degree Expressions for Pragmatic Reasoning (CDE-Prag), an ongoing project about Enriched Meaning.
Dataset Summary
CDE-Prag is a theory-driven evaluation dataset designed to probe the pragmatic competence of Large Language Models (LLMs) and Vision-Language Models (VLMs). It focuses specifically on manner implicatures and ambiguity detection through the lens of Chinese degree expressions (e.g., Kai gao, which is ambiguous between "Kai is tall" and… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/EnrichedMeaningDataset.everyday_conversations_ja
データセットについて
このデータセットは、 HuggingFaceTB/everyday-conversations-llama3.1-2k を機械翻訳で日本語化したものになります。
具体的には、everyday-conversations-llama3.1-2kをトピックごとの対話のペアに変更してDeepLで翻訳したものとなります。
詳細
topic: everyday-conversations-llama3.1-2kのtopic
user: 各トピックごとのユーザーからの発話
assistant: 各トピックごとのユーザーへの返答
assistantの返答がない場合はNone
注意事項
人手で若干修正をしましたが、日本語が変な箇所がいくつか散見されます。
ライセンス:apache 2.0
UN_NU_interpretation_LLMs
Quantifier Scope Interpretation Dataset
Datasets for an ongoing project about Scope preferences and ambiguity in LLM interpretation.
Dataset Structure
Splits
The dataset consists of synthetically generated stimuli pairing target sentences with interpretation-biased contexts (SSR vs. ISR).
Features
language (string)Language of the stimulus (English or Chinese).
structure (string)Surface syntactic configuration of the sentence:UN (universal >… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/UN_NU_interpretation_LLMs.
