CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K277 likes19k downloads3y agoHugging Face02prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes6.1k downloads1y agoHugging Face03Pn101 /taxbench-au TaxBench-AU A benchmark for testing whether AI agents can calculate Australian tax. TaxBench-AU contains 156 Australian tax calculation questions, presented as multiple-choice (4-option) worked tax problems. The benchmark is designed to test whether an AI agent can read the facts, apply the right Australian tax rule for the relevant income year, do the calculation, and choose the correct answer. The Kaggle mirror is published as Agent Tax Exam for Australian Tax. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Pn101/taxbench-au.documentquestion-answeringn<1K0 likes4.5k downloads4mo agoHugging Face04bowen-upenn /PersonaMem-v1🚨 We have now released PersonaMem-v3 and PersonaMem-v2. This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark. We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized responses across task scenarios. PersonaMem emphasizes persona-oriented, multi-session… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v1.tabulartext-generation1K<n<10K18 likes3.2k downloads20d agoHugging Face05mercor /APEX-v1-extended APEX-v1-extended The AI Productivity Index (APEX) is a benchmark from Mercor for assessing whether frontier models are capable of performing economically valuable tasks across four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). APEX-v1-extended doubles the heldout evaluation set from n=200 to n=400, with increased complexity and variety. On average, tasks take over two-and-a-half hours for seasoned professionals to… See the full description on the dataset page: https://huggingface.co/datasets/mercor/APEX-v1-extended.documenttext-generationn<1K17 likes1.7k downloads5mo agoHugging Face06prquan /STARK_1k Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges Dataset for our paper: Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges Contact Information If you have any questions or feedback, feel free to reach out: Name: Pengrui Quan Email: prquan@ucla.edu License Copyright (c) 2025, UCLA Networked and Embedded Systems Laboratory (NESL) All rights reserved. Redistribution and use in… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_1k.textquestion-answering1K<n<10K0 likes1.1k downloads10mo agoHugging Face07ncbi /MedCalc-Bench-v1.2 [!note] Please visit MedCalc-Bench Verified at this url: https://github.com/nikhilk7153/MedCalc-Bench-Verified for the latest changes. Here is the HuggingFace link: https://huggingface.co/datasets/nsk7153/MedCalc-Bench-Verified. The first version of MedCalc-Bench Verified is an update from v1.2 on this repository. We recommend using v1.0-v1.2 for reproducibility purposes only. You should specify which version you are using when using this dataset. MedCalc-Bench is the first medical… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.2.textquestion-answering10K<n<100K3 likes1.1k downloads9mo agoHugging Face08wangyz1999 /GameplayQA GameplayQA: A Decision-Dense POV-Synced Multi-Video Understanding Benchmark of 3D Virtual Agents Yunzhe Wang   Runhui Xu   Kexin Zheng   Tianyi Zhang Jayavibhav N. Kogundi   Soham Hans   Volkan Ustun University of Southern California ACL 2026 Corresponding Author: yunzhewa@usc.edu Overview GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.tabularvideo-text-to-text1K<n<10K7 likes1k downloads4mo agoHugging Face09sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M85 likes914 downloads2y agoHugging Face10MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes621 downloads1y agoHugging Face11UVSKKR /Ethical-Reasoning-in-Mental-Health-v1gatedThis repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI. Overview Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias. Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.textquestion-answeringn<1K4 likes573 downloads1y agoHugging Face12MarineLife-16K /MarineLife-16K Dataset Card for MarineLife-16K We introduce MarineLife-16K, a marine-domain video benchmark designed to evaluate the video understanding capabilities of Vision-Language Models (VLMs). MarineLife-16K contains 2,000 video-text pairs and 16,080 video-question-answer pairs across a collection of 2,000 marine videos, including 12,080 multiple-choice questions and 4,000 open-ended questions. The benchmark emphasizes specialized marine knowledge, visual reasoning, temporal… See the full description on the dataset page: https://huggingface.co/datasets/MarineLife-16K/MarineLife-16K.imagevisual-question-answeringn<1K0 likes365 downloads3mo agoHugging Face13gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes306 downloads4mo agoHugging Face14giorgio-mariani-1 /GLUE3D GLUE3D: General Language Understanding Evaluation for 3D Point Clouds Data repository containing all necessary data for the GLUE3D evaluation benchmark. GLUE3D is a Q&A benchmark for evaluation of 3D-LLMs object understanding capabilities. It is built around 128 richly textured surfaces spanning creatures, objects, architecture and transport. Each surface is provided as a 50 k-point RGB point cloud, a 8K-point RGB point cloud, a 512 × 512 RGB rendering, and five RGB-D multiviews.… See the full description on the dataset page: https://huggingface.co/datasets/giorgio-mariani-1/GLUE3D.imagetext-generation1K<n<10K1 likes271 downloads11mo agoHugging Face15ncbi /MedCalc-Bench-v1.0gated Important: This dataset is kept for reproducibility purposes only. Please use the most up-to-date version 1.2, for the most revised and corrected dataset available. We have fixed 12 calculator implementations, ensured the Relevant Entities section best matches with what was specified by a patient note, and have also replaced notes which are better fits for a given calculator to make the dataset more applicable for real-life sitatuons. Because of the number of changes, we find this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.0.tabularquestion-answering10K<n<100K2 likes196 downloads10mo agoHugging Face16natong19 /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.tabularquestion-answering1K<n<10K0 likes182 downloads9mo agoHugging Face17ncbi /MedCalc-Bench-v1.1gatedWe have revised and improved upon v1.1 based on the updates in 1.2: https://github.com/ncbi-nlp/MedCalc-Bench/releases/tag/version-1.2 Please use the most up-to-date version here: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.2 tabularquestion-answering10K<n<100K1 likes150 downloads10mo agoHugging Face18FinchResearch /pallas_splitted_18ctexttext-classification1M<n<10M0 likes141 downloads3y agoHugging Face19lianghsun /tw-legal-benchmark-v1 Taiwan Legal Benchmark v1 A multiple-choice benchmark for evaluating large language models on Taiwan law in Traditional Chinese (繁體中文). It covers six legal domains with 209 questions drawn from Taiwan bar exam and certification-style questions. Overview Property Value Language Traditional Chinese (zh-TW) Questions 209 Format 4-choice multiple choice (A / B / C / D) Domain Taiwan law License Apache 2.0 Legal Domains Covered… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v1.textquestion-answeringn<1K7 likes139 downloads6mo agoHugging Face20stewy33 /acc_rd_s1-gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.tabularquestion-answering1K<n<10K0 likes138 downloads2y agoHugging Face21Aipresso /10k_rows_cleaned_prompts 10K Rows Cleaned Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use You must provide attribution when using this data in publications, research, or commercial products. Dataset Overview A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models. 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.texttext-generation1M<n<10M0 likes128 downloads1y agoHugging Face22JDhruv14 /Bhagavad-Gita-QA Bhagavad-Gita-QA-Multilingual Dataset Summary Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework. This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita-QA.tabularquestion-answering10K<n<100K5 likes125 downloads1y agoHugging Face23Sandhya1912 /AA-Omniscience-Public Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/Sandhya1912/AA-Omniscience-Public.documentquestion-answeringn<1K0 likes110 downloads4mo agoHugging Face24roskosmos19 /agentic-reasoning-benchmark Agentic & Reasoning Benchmark (ARB) – Expanded Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning. Überblick Eigenschaft Wert Anzahl Beispiele 2.550 Kategorien 8 Schwierigkeitsgrade easy / medium / hard Formate CSV + JSON Reproduzierbarkeit Generator-Skript (seed=42) enthalten Lizenz CC-BY-4.0 Kategorien Kategorie Anzahl Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.textquestion-answering1K<n<10K1 likes105 downloads22d agoHugging Face25Fancy-MLLM /R1-Onevision-Bench R1-Onevision-Bench [📂 GitHub][📝 Paper] [🤗 HF Dataset] [🤗 HF Model] [🤗 HF Demo] Dataset Overview R1-Onevision-Bench comprises 38 subcategories organized into 5 major domains, including Math, Biology, Chemistry, Physics, Deducation. Additionally, the tasks are categorized into five levels of difficulty, ranging from ‘Junior High School’ to ‘Social Test’ challenges, ensuring a comprehensive evaluation of model capabilities across varying complexities.… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision-Bench.textquestion-answeringn<1K3 likes103 downloads2y agoHugging Face26jang1563 /SpaceOmicsBench-v3 SpaceOmicsBench v3 A Multi-Omics AI Benchmark for Spaceflight Biomedical Data SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2). Dataset Structure ML Track (Track A) tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration) tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.tabulartabular-classification10K<n<100K0 likes102 downloads16d agoHugging Face27FirstBML1 /afrofinchain-multilingual-web3 AfroFinChain — Multilingual Web3 & Blockchain Dataset Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable. Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.texttext-generation1K<n<10K0 likes101 downloads5mo agoHugging Face28sarahwei /cyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.textquestion-answering10K<n<100K6 likes99 downloads2y agoHugging Face29Hatman /plot-palette-100k Empowering Writers with a Universe of Ideas Plot Palette DataSet HuggingFace » Plot Palette was created to fine-tune large language models for creative writing, generating diverse outputs through iterative loops and seed data. It is designed to be run on a Linux system with systemctl for managing services. Included is the service structure, specific category prompts and ~100k data entries. The dataset is available here or… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/plot-palette-100k.textquestion-answering10K<n<100K4 likes97 downloads2y agoHugging Face30pampalini1 /olivers-mtor-atlas Oliver's mTOR Atlas The mTOR pathway, mapped by what the evidence can actually carry. This dataset is the curated corpus behind mtor-atlas.org: 414 hand-selected studies on mTOR (mechanistic target of rapamycin) signalling, each labelled by the kind of study behind it, and a list of 149 pathway entities (genes and proteins, complexes, drugs, interventions, biological processes, diseases, outcomes, organelles, nutrients and conditions) that the studies refer to. Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/pampalini1/olivers-mtor-atlas.tabulartext-classificationn<1K0 likes80 downloads15h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.