datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell,
Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
A collaboration between:
Meta Recommendation Systems
University of Pennsylvania
MIT
Third release in the PersonaMem series:
PersonaMem-v1: [COLM… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v3.function_calling_v3_SAMPLE
Trelis Function Calling Dataset - VERSION 3 - SAMPLE
This is a SAMPLE of the v3 dataset available for purchase here.
Features:
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation).
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.pubmedqa-recursive-llm-degradation-qwen2.5-3b
PubMedQA Recursive LLM Degradation — Qwen2.5-3B
This repository contains synthetic biomedical question-answering data
and model predictions generated as part of a study of recursive
fine-tuning and model degradation.
Base Model
Qwen/Qwen2.5-3B
Source Dataset
The experiments use the PubMedQA dataset:
qiaoxin/PubMedQA
This repository contains generated/derived research artifacts and does
not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-3b.acc_rd_s1-gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.R3-eval-MMMLUSpaceOmicsBench-v3
SpaceOmicsBench v3
A Multi-Omics AI Benchmark for Spaceflight Biomedical Data
SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2).
Dataset Structure
ML Track (Track A)
tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration)
tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.hvac-error-codes
HVAC Bench Error Code Dataset
Error codes from HVAC equipment sold in the United States, United Kingdom, and Europe, as published by HVAC Bench. Each record gives the manufacturer, the code, the product family the definition applies to, a plain-language meaning, the checks an owner can safely make, the point at which a technician is needed, and the date the definition was last checked against manufacturer documentation. Codes are specific to a product family and are not… See the full description on the dataset page: https://huggingface.co/datasets/mukarram360/hvac-error-codes.afrofinchain-multilingual-web3
AfroFinChain — Multilingual Web3 & Blockchain Dataset
Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable.
Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.arXivTection
📄 arXivTection Dataset
The arXivTection dataset serves as a benchmark designed for the task of detecting pretraining data from Large Language models.
The dataset consists of 50 research papers extracted from arXiv.
25 published in 2023: Non-Training data, "label" column = 0.
25 published before 2022: Training data, "label" column = 1.
From each paper ≈ 30 passages are extracted. Each passage is paraphrased 3 times using the Language Model Claude v2.0.
The "Answer" column… See the full description on the dataset page: https://huggingface.co/datasets/avduarte333/arXivTection.cybersecurity-QA-with-negatives
Cybersecurity QA Dataset With Negatives
Description
This dataset was created by combining and processing the following publicly available cybersecurity question-answering datasets:
Rowden/CybersecurityQAA
sambanovasystems/attackqa
mariiazhiv/cybersecurity_qa
The resulting dataset is designed for training and evaluating retrieval, embedding, reranking, and contrastive learning models in the cybersecurity domain.
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/jobby32/cybersecurity-QA-with-negatives.clinical-trial-outcomes-predictions
Clinical Trial Outcomes Prediction Dataset
A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology.
Dataset Description
This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.patient-doctor-qa-tr-321179
Patient Doctor Q&A TR 321179 Veri Kümesi
Patient Doctor Q&A TR 321179 veri kümesi, Patient Doctor Q&A TR 19583, Patient Doctor Q&A TR 167732, Patient Doctor Q&A TR 5695 ve Patient Doctor Q&A TR 95588 veri kümelerinin birleştirilmiş ve karıştırılmış halidir.
Ana Özellikler:
İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları.
Yapı: 2 sütun içerir: Soru, Cevap.Dil: Türkçe.
Potansiyel Kullanım Alanları:
Tıbbi araştırmalar
Doğal Dil… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-321179.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
preguntas-canal-denuncias-ley-2-2023
Preguntas y respuestas sobre el canal de denuncias (Ley 2/2023) en España
Banco de 18 preguntas reales con respuesta directa sobre el canal de denuncias obligatorio en España (Ley 2/2023, protección del informante / whistleblowing): qué empresas están obligadas, umbrales y fechas, anonimato, plazos, sanciones y requisitos del sistema. Publicado por Nucleo360, software de RRHH para pymes españolas.
Contenido
preguntas.csv — una pregunta por fila, con las columnas:… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/preguntas-canal-denuncias-ley-2-2023.BookTection
📚 BookTection Dataset
The BookTection dataset serves as a benchmark designed for the task of detecting pretraining data from Large Language models.
The dataset consists of 165 books.
60 published in 2023: Non-Training data, "label" column = 0.
105 published before 2022: Training data, "label" column = 1.
From each book ≈ 34 passages are extracted. Each passage is paraphrased 3 times using the Language Model Claude v2.0.
The "Answer" column indicates which of the passages is the… See the full description on the dataset page: https://huggingface.co/datasets/avduarte333/BookTection.preguntas-normativa-laboral-rrhh-espana
Normativa laboral y RRHH en España: 300 preguntas con su respuesta y su artículo
Corpus de 300 pares de pregunta y respuesta sobre recursos humanos y normativa laboral española, repartidos en 30 temas y con la referencia normativa en 286 de ellos. Todas las respuestas son autónomas: se entienden sin contexto adicional.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Qué cubre
Tema
Preguntas
Registro horario
42
Inspección de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/preguntas-normativa-laboral-rrhh-espana.UniD3_DDMpreguntas-inspeccion-trabajo-registro-horario
Preguntas frecuentes sobre la Inspección de Trabajo y el registro horario (España)
18 preguntas y respuestas sobre qué pide la Inspección de Trabajo en materia de registro de jornada, cómo se responde a un requerimiento, qué sanciones se aplican y qué errores concretos acaban en acta. Cada respuesta indica la norma y el artículo del que sale.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Por qué este conjunto de datos
El registro de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/preguntas-inspeccion-trabajo-registro-horario.Re-Auto-30K
Re-Auto-30K: A Comprehensive AI Safety Evaluation Dataset for Code Generation
Dataset Overview
Re-Auto-30K is a meticulously curated dataset containing 30,886 security-focused prompts designed specifically for evaluating AI safety in code generation scenarios. This dataset serves as a comprehensive benchmark for assessing Large Language Models (LLMs) across multiple dimensions of security, reliability, and autonomous behavior in software engineering contexts.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/navneetsatyamkumar/Re-Auto-30K.synthetic_discharge_summ
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is a subset of the dataset for Asclepius model (arxiv).
The original dataset is made up of synthetic notes generated from PMC-Patients case reports with GPT-3.5.
We filtered the summarization task for discharge notes. The dataset contains 13,584 notes.
Supported Tasks
This dataset covers below summarization task
Languages
English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/synthetic_discharge_summ.faqs-rrhh
FAQ de RRHH, registro horario y canal de denuncias (España)
Conjunto de preguntas frecuentes con respuesta sobre software de RRHH, registro horario (RD-ley 8/2019) y canal de denuncias (Ley 2/2023) para pymes españolas, publicado por Nucleo360. Las 8 primeras respuestas coinciden literalmente con las publicadas en la web oficial; el resto amplía la FAQ del registro horario obligatorio en 2026.
Contenido
faqs.csv — una pregunta por fila, con las columnas:… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/faqs-rrhh.MedQnA_version3
Reference:
"A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019.
ViMedical_DiseaseThis dataset contains over 12K+ questions and symptoms related to various common diseases in Vietnamese. It's designed to aid in the classification of medical symptoms and provide preliminary disease identification. The dataset covers a wide range of diseases, including cardiovascular, digestive, neurological, dermatological, endocrine, and others.
For more information and updates about the dataset, please refer to the main repository here.
This dataset can be used for:
Data… See the full description on the dataset page: https://huggingface.co/datasets/PB3002/ViMedical_Disease.cantonesewiki_doyouknow
Cantonese Question Dataset from Yue Wiki
A collection of questions in Cantonese, extracted from Yue Wiki. This dataset contains a variety of questions covering different topics and domains.
Disclaimer
The content and opinions expressed in this dataset do not represent the views, beliefs, or positions of the dataset creators, contributors, or hosting organizations. This dataset is provided solely for the purpose of improving AI systems' understanding of the Cantonese… See the full description on the dataset page: https://huggingface.co/datasets/pendingremove32894/cantonesewiki_doyouknow.siddha_vaithiyam_question_answering_chatbot
Medical Home Remedy Chatbot Dataset
Overview
This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past.
Contents
Dataset Files:
dataset.csv : The main dataset file containing questions and corresponding home remedy answers.
Data Structure:
Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.UniD3_DEAUniD3_DTAIndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Anonymized review copy — NeurIPS 2026 Evaluations & Datasets Track, Submission 3390 (under review).
IndustryBench is a benchmark for evaluating the industrial procurement knowledge of large language models. It comprises 2,049 QA pairs grounded in Chinese national standards (GB/T) and structured industrial product records, with item-aligned renderings in Chinese, English, Russian, and Vietnamese.
Erratum note:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips26-3390/IndustryBench.english_telugu_slangeconcausal-benchmark📊 EconCausal: A Context-Aware Causal Reasoning Benchmark for LLMs
Donggyu Lee, Hyeok Yun, Meeyoung Cha, Sungwon Park, Sangyoon Park, Jihee Kim
🌍 Overview
Socio-economic causal effects depend heavily on their specific institutional and environmental context. A single intervention can produce opposite results depending on regulatory or market factors.
EconCausal is a large-scale benchmark comprising 10,490 context-annotated causal triplets extracted from 2,595… See the full description on the dataset page: https://huggingface.co/datasets/qwqw3535/econcausal-benchmark.
