datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AttaQ
AttaQ Dataset Card
The AttaQ red teaming dataset, consisting of 1402 carefully crafted adversarial questions, is designed to evaluate Large Language Models (LLMs) by assessing their tendency to generate harmful or undesirable responses.
It may serve as a benchmark to assess the potential harm of responses produced by LLMs.
The dataset is categorized into seven distinct classes of questions: deception, discrimination, harmful information, substance abuse, sexual content, personally… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AttaQ.mac-app-store-apps-release-notes
Dataset Card for Macappstore Applications Release Notes
📌 Dataset status: static snapshot (no scheduled updates). This dataset is derived from the December 2023 – January 2024 Mac App Store metadata snapshot and reflects the store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications release notes extracted from the metadata from the public API.
Curated by: MacPaw Way Ltd.
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-release-notes.ASRD-Dataset
Adversarial Surface-Form Robustness Dataset (ASRD)
Anonymous Repository for Double-Blind ReviewNeurIPS 2026 Workshop
1. Dataset Overview
Standard safety evaluations of large language models routinely measure model refusal and compliance using canonical plain-text instructions. However, deployed systems frequently encounter non-canonical inputs containing expressive symbols (emojis), character-level substitutions (homoglyphs, leetspeak), structured encodings… See the full description on the dataset page: https://huggingface.co/datasets/asrd-research/ASRD-Dataset.MermaidSeqBench
Dataset Card for MermaidSeqBench
Dataset Summary
This dataset provides a human-verified benchmark for assessing large language models (LLMs) on their ability to generate Mermaid sequence diagrams from natural language prompts.
The dataset was synthetically generated using large language models (LLMs), starting from a small set of seed examples provided by a subject-matter expert. All outputs were subsequently manually verified and corrected by human annotators to ensure… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/MermaidSeqBench.idt5-v4-results-final-lora-s123-20260912T013040606815Z
final-lora-s123-20260912T013040606815Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 94.7565543071161,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s123-20260912T013040606815Z.AI-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
Jailbreak Prompts
Dataset Summary
Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[German]
dementor-matrix-responses
Dementor — matrix model responses
Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study.
Companion to:
Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan)
Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO /
self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant).
Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.idt5-v4-results-final-lora-s2026-20260912T034640190015Z
final-lora-s2026-20260912T034640190015Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 95.88014981273409,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s2026-20260912T034640190015Z.Arabic-books-and-research-dataset
Arabic reserach and books dataset (ARABD)
This dataset is an extracted cleaned text from more than 60K word files with unique arabic texts never published before.
Dataset diversity
the dataset is diverse from all kind of islamic research: [feqh, hadeeth, tafseer, tahqeeq, ... etc], from new written research to a manuscirpts.
dataset size
the dataset was more than 11GB but after cleaning (pre-processing) it becase a straight 10GB with less noisy chars.… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Arabic-books-and-research-dataset.souslab-us-restaurant-menus
Souslab — US Restaurant Menus
A structured sample of the Souslab US restaurant menu dataset: real restaurants, real menu items, real prices — normalized into a clean schema you can train on or analyze directly.
This sample is published openly under CC-BY-NC-4.0 for research and non-commercial evaluation. The full dataset — 449,000+ US restaurants and 44.3M+ menu items, refreshed continuously with chain-level aggregation — is available via the Souslab API under commercial… See the full description on the dataset page: https://huggingface.co/datasets/AnyStackLabsdev/souslab-us-restaurant-menus.epfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.idt5-v4-results-final-fft-s2026-20260911T063507183162Z
final-fft-s2026-20260911T063507183162Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 88.01498127340824,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s2026-20260911T063507183162Z.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.llm-misinformation-resistance-index
LLM Misinformation Resistance Index (LMRI)
Formal name: LLM Misinformation Resistance Index (LMRI).
Public alias: the Gaslighting Index — the two headline scores keep their code
names GI-basic and GI-strict, where "GI" comes from the benchmark's public alias.
LMRI measures whether a language model will stand up to its own misinformation.
Each benchmark item is a fabricated conversation in which the assistant's own prior
turn contains a planted false claim (or, for controls, a… See the full description on the dataset page: https://huggingface.co/datasets/buildwithdmytro/llm-misinformation-resistance-index.open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.idt5-v4-results-final-lora-s42-20260912T063343815032Z
final-lora-s42-20260912T063343815032Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 92.88389513108615,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s42-20260912T063343815032Z.idt5-v4-results-final-fft-s42-20260910T135823740810Z
final-fft-s42-20260910T135823740810Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 93.63295880149813,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s42-20260910T135823740810Z.idt5-v4-results-final-fft-s123-20260910T215641650583Z
final-fft-s123-20260910T215641650583Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 93.25842696629213,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s123-20260910T215641650583Z.Patient-Message-Response-DraftingPaper: How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting (arxiv link)
Dataset Details:
The patient message response drafting dataset is designed to evaluate how well LLMs respond to patient messages in patient portal communication.
Each semi-synthetic patient message is paired with a real de-identified EHR from a patient at our collaborating hospital.
Each doctor response is written by a clinician, guided by clinician response themes… See the full description on the dataset page: https://huggingface.co/datasets/PortalPal-AI/Patient-Message-Response-Drafting.Contextual_Response_Evaluation_for_ESL_and_ASD_Support
Dataset Card for "Contextual Response Evaluation for ESL and ASD Support💜💬🌐""
Dataset Description 📖
Dataset Summary 📝
Curated by Eric Soderquist, this dataset is a collection of English prompts and responses generated by the Phi-2 model, designed to evaluate and improve NLP models for supporting ESL (English as a Second Language) and ASD (Autism Spectrum Disorder) user bases. Each prompt is paired with multiple AI-generated responses and evaluated using a… See the full description on the dataset page: https://huggingface.co/datasets/yunjaeys/Contextual_Response_Evaluation_for_ESL_and_ASD_Support.cqa-ai-technical-response-evaluation
CQA AI Technical Response Evaluation Dataset
Overview
This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses.
The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.clinical_time_gap_resilience_v0.1Clinical Time Gap Resilience
PurposeTest whether a model avoids anchoring on stale data when time passes and new information arrives.
Input fields
last_known_state
time_gap
new_info
proposed_action
Required outputOne JSON object
time_gap_resilientyes or no
gap_risklow, medium, high
correct_actionone sentence
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical_instruction_conflict_resolution_v0.1Clinical Instruction Conflict Resolution
Tests whether a model resolves conflicting instructions safely.
Output JSON
resolved
conflict_type
correct_action
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
ua-codeforces-cots-open-r1
Dataset Summary
ua-codeforces-cots-open-r1 is a Ukrainian-focused derivative of open-r1/codeforces-cots that:
includes 1550 Python solutions from original dataset generated by DeepSeek-R1;
adds Ukrainian translations of Codeforces task statements, I/O formats, notes, and editorials;
provides Ukrainian translation of original ("high") reasoning obtained with DeepSeek-V3;
adds “low” reasoning in Ukrainian by DeepSeek-R1 based on original reasoning and task statements;
ships… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-codeforces-cots-open-r1.resumo
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/PericlesSavio/resumo.Party_Affairs_ResponseData from https://wenda.12371.cn/liebiao.php
instruction_conflict_resolution_v01Instruction Conflict Resolution v0.1
This evaluation dataset tests how models resolve conflicting instructions.
It targets a common failure mode: following the most recent or most forceful instruction even when it conflicts with higher-priority constraints.
This is not training data.
What it tests
Priority handling under instruction conflict
Refusal stability under escalation
Logical conflict handling for impossible constraints
Post-conflict integrity with no delayed leakage… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/instruction_conflict_resolution_v01.concept-to-root-dictionary
🌿 Concept-to-Root Dictionary
A mapping of universal concepts to Arabic triliteral roots for semantic compression
📖 Overview
This dataset provides mappings between universal semantic concepts and Arabic triliteral roots, designed for use as a compression layer in Large Language Models.
What are Arabic Roots?
Arabic uses a root-and-pattern morphological system where most words derive from 3-letter roots:
Root
Core Meaning
Derived Words… See the full description on the dataset page: https://huggingface.co/datasets/root-semantic-research/concept-to-root-dictionary.Amazigh_Researchers_Vocabulary
Mohammed Lchger Vocabulary Dataset
This dataset contains a collection of vocabulary compiled by Mohamed Lachgar, the owner of the Amazigh researchers blog which its dataset can be find here.
Dataset Details
Content: 2,000 scientific related terms and general vocabulary.
Languages: Amazigh (zgh, ber) and English.
Script: Tifinagh.
Unique Feature: It is unique especially in its inclusion of scientific vocabulary.
Acknowledgments
Thanks to Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh_Researchers_Vocabulary.System-Response-100K
System-Response-100K dataset
This dataset contains text and code for machine learning tasks including:
Text Generation
Text Classification
Summarization
Question Answering
The dataset includes text formatted in JSON and is in English.
Dataset Statistics
Number of entries: Not specified in the information you provided.
Modalities
Text
Code
Formats
JSON
Languages
English
Getting Started
This section can include… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/System-Response-100K.
