CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TAUR-Lab /MuSR MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning Creating murder mysteries that require multi-step reasoning with commonsense using ChatGPT! By: Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. View the dataset on our custom viewer and project website! Check out the paper. Appeared at ICLR 2024 as a spotlight presentation! Git Repo with the source data, how to recreate the dataset (and create new ones!) here textquestion-answeringn<1K24 likes19k downloads2y agoHugging Face02furonghuang-lab /PHTest🌟 PHTest: Evaluating False Refusals in LLMs 🤖 Auto Red-Teaming All prompts are generated automatically using a controllable text-generation technique called AutoDAN. 🌐 Diverse Prompts PHTest introduces false refusal patterns that aren’t present in existing datasets, including prompts that avoid mentioning sensitive words. ⚖️ Harmlessness & Controversial Labeling Controversial prompts are separately labeled to address the… See the full description on the dataset page: https://huggingface.co/datasets/furonghuang-lab/PHTest.texttext-generation1K<n<10K3 likes434 downloads5mo agoHugging Face03jngb-labs /InvoiceBenchmark InvoiceBenchmark 200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number. The Pitch Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.documentquestion-answeringn<1K0 likes385 downloads5mo agoHugging Face04bryel-labs /MuSP-Bench MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding Across Score and Performance MuSP-Bench is a 490-question benchmark for musical score understanding, performance listening, and combined score-performance reasoning. Official benchmark website Modalities Each question specifies the minimum source of musical evidence needed to answer it: S (score): answer from the written score. P (performance): answer from the performance recording. S&P (score… See the full description on the dataset page: https://huggingface.co/datasets/bryel-labs/MuSP-Bench.imagequestion-answeringn<1K1 likes259 downloads27d agoHugging Face05Rose-STL-Lab /ClimaQA ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025) Check the paper's webpage and GitHub for more info! The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.textquestion-answering1K<n<10K3 likes203 downloads2y agoHugging Face06agentic-learning-ai-lab /daily-oracle Daily Oracle 📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time. Dataset Details Question Type: True/False (TF) & Multiple Choice (MC) Current Version* Time Span: 2020.01.01 - 2026.07.18 Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.textquestion-answering10K<n<100K4 likes156 downloads2mo agoHugging Face07riotu-lab /MURAD A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset Authors: Serry Sibaee, Yasser Alhabashi, Nadia Sibai, Yara Farouk, Adel Ammar, Sawsan AlHalawani, Wadii Boulila Overview MURAD A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset is an open Arabic lexical dataset designed to support research in computational linguistics, lexicography, and Arabic natural language processing (NLP). The dataset contains 96,243… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/MURAD.textsummarization10K<n<100K2 likes122 downloads11d agoHugging Face08Nucleo360 /preguntas-normativa-laboral-rrhh-espana Normativa laboral y RRHH en España: 300 preguntas con su respuesta y su artículo Corpus de 300 pares de pregunta y respuesta sobre recursos humanos y normativa laboral española, repartidos en 30 temas y con la referencia normativa en 286 de ellos. Todas las respuestas son autónomas: se entienden sin contexto adicional. Publicado por Nucleo360, software de recursos humanos para pymes españolas. Qué cubre Tema Preguntas Registro horario 42 Inspección de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/preguntas-normativa-laboral-rrhh-espana.textquestion-answeringn<1K0 likes48 downloads22d agoHugging Face09riotu-lab /all_RD_datasets RD Dataset With References This dataset contains Arabic terms and their definitions.The data was extracted and combined from the following sources: https://huggingface.co/datasets/Basma2423/Arabic-Terminologies-and-Definitions https://data.mendeley.com/datasets/gxr3j4tdk5/3 https://huggingface.co/datasets/MohamedRashad/arabic-roots https://arai.ksaa.gov.sa/sharedTask2024/ Each entry consists of: word definition textquestion-answering100K<n<1M0 likes25 downloads10mo agoHugging Face10CALM-Lab-Purdue /EnrichedMeaningDataset Dataset Card for Chinese Degree Expressions for Pragmatic Reasoning (CDE-Prag), an ongoing project about Enriched Meaning. Dataset Summary CDE-Prag is a theory-driven evaluation dataset designed to probe the pragmatic competence of Large Language Models (LLMs) and Vision-Language Models (VLMs). It focuses specifically on manner implicatures and ambiguity detection through the lens of Chinese degree expressions (e.g., Kai gao, which is ambiguous between "Kai is tall" and… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/EnrichedMeaningDataset.tabularquestion-answeringn<1K1 likes16 downloads7mo agoHugging Face11U23-lab /everyday_conversations_ja データセットについて このデータセットは、 HuggingFaceTB/everyday-conversations-llama3.1-2k を機械翻訳で日本語化したものになります。 具体的には、everyday-conversations-llama3.1-2kをトピックごとの対話のペアに変更してDeepLで翻訳したものとなります。 詳細 topic: everyday-conversations-llama3.1-2kのtopic user: 各トピックごとのユーザーからの発話 assistant: 各トピックごとのユーザーへの返答 assistantの返答がない場合はNone 注意事項 人手で若干修正をしましたが、日本語が変な箇所がいくつか散見されます。 ライセンス:apache 2.0 textquestion-answering1K<n<10K2 likes11 downloads9mo agoHugging Face12CALM-Lab-Purdue /UN_NU_interpretation_LLMsgated Quantifier Scope Interpretation Dataset Datasets for an ongoing project about Scope preferences and ambiguity in LLM interpretation. Dataset Structure Splits The dataset consists of synthetically generated stimuli pairing target sentences with interpretation-biased contexts (SSR vs. ISR). Features language (string)Language of the stimulus (English or Chinese). structure (string)Surface syntactic configuration of the sentence:UN (universal >… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/UN_NU_interpretation_LLMs.texttext-classificationn<1K1 likes1 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.