datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
V2Tin-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs
This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang.
In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1,405… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts.trustworthy-biology-agents-traces
Trustworthy Biology Agents — Run Traces
Raw execution traces from 1,329 agent runs across three coding agents on three
biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed
trace bundle for the study in
manu-tej/ai-scientists; the write-up
lives in that repo's RESULTS.md.
The motivating question is not only whether an agent reaches the right answer, but
whether it behaves like a trustworthy analyst when the task is ambiguous,
under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.forbidden_question_set
Forbidden Question Set
This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.
It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy.
We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.trust-game-llama-2-chat-historyHonestyBench
HonestyBench
This is the official repo of the paper Annotation-Efficient Universal Honesty Alignment.
HonestyBench is a large-scale benchmark that consolidates 10 widely used public freeform factual question-answering datasets. HonestyBench comprises 560k training samples, along with 38k in-domain and 33k out-of-domain (OOD) evaluation samples. It establishes a pathway toward achieving the upper bound of performance for universal models across diverse tasks, while also serving as a… See the full description on the dataset page: https://huggingface.co/datasets/Trustworthy-Information-Access/HonestyBench.tensor-trust
The Tensor Trust dataset (v1 benchmarks, v2 raw data dump) (mirror of GitHub version)
Other Tensor Trust links: [Game] [Code] [Paper]
This HF dataset contains the raw data and derived benchmarks for the Tensor Trust project.
An interactive explanation of how to load and use the data (including the meaning of the columns) is in a Jupyter notebook in this directory.
You can click here to run the notebook right now in Google Colab.
SpecUBench
LLM-Specific Utility Benchmark (SpecUBench)
SpecUBench is a benchmark for measuring the LLM-specific utility of retrieved
passages in retrieval-augmented generation (RAG). Instead of assuming that a
"relevant" passage is equally useful to every reader, UtilityBench labels how useful
each retrieved passage is for a specific LLM — i.e., how much the passage actually
helps that model produce the correct answer.
The benchmark is built on six widely-used open-domain QA / retrieval… See the full description on the dataset page: https://huggingface.co/datasets/Trustworthy-Information-Access/SpecUBench.trust-data
FLIP mock trust data
Mock data for the dev and test trusts of FLIP, the
Federated Learning Interoperability Platform. Everything here is either synthetic or derived
from a public research dataset under its licence — no patient data. Licensing is per project
(see the table and the per-project notes): the spleen- and cxr-derived content is CC BY-SA 4.0,
the prostate-derived content (prostate_project) is CC BY-NC 4.0 and therefore non-commercial;
the pathology tables… See the full description on the dataset page: https://huggingface.co/datasets/aicentreflip/trust-data.When2Speak
When2Speak Dataset
Dataset for "When2Speak: A Dataset for Temporal Participation and Turn-Taking in Multi-Party Conversations for Large Language Models"
NeurIPS 2026 — Evaluations and Datasets Track
Overview
When2Speak is a large-scale synthetic dataset for learning intervention timing in multi-party conversations: given the recent conversation history, should an AI agent speak or remain silent at this turn?
The dataset comprises 216,800 labeled (context, decision) pairs… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/When2Speak.Ford_Credit_Auto_Owner_Trust_2021_A_1843634
Ford Credit Auto Owner Trust 2021-A
SEC ABS-EE asset-level filings for CIK 1843634 (Ford Credit Auto Owner Trust 2021-A).
Filings: 50
Parquet files: 50
Total size: 120.6 MB
Reporting period start: 2021-01-31
Reporting period end: 2025-02-28
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).
Filing… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Ford_Credit_Auto_Owner_Trust_2021_A_1843634.CarMax_Auto_Owner_Trust_2020_3_1814294
cik
form
accessionNumber
fileNumber
filmNumber
reportDate
url
1814294
ABS-EE
0001259380-20-000015
333-228379-07
201018683
2020-06-30
https://sec.gov/Archives/edgar/data/1814294/000125938020000015
1814294
ABS-EE
0001259380-20-000017
333-228379-07
201018879
2020-06-30
https://sec.gov/Archives/edgar/data/1814294/000125938020000017
1814294
ABS-EE
0001814294-20-000003
333-228379-07
201108203
2020-07-31
https://sec.gov/Archives/edgar/data/1814294/000181429420000003
1814294
ABS-EE… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/CarMax_Auto_Owner_Trust_2020_3_1814294.Nissan_Auto_Lease_Trust_2021_A_1886591
Nissan Auto Lease Trust 2021-A
SEC ABS-EE asset-level filings for CIK 1886591 (Nissan Auto Lease Trust 2021-A).
Filings: 26
Parquet files: 52
Total size: 74.8 MB
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).
Filing index
cik
form
accessionNumber
url
1886591
ABS-EE… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Nissan_Auto_Lease_Trust_2021_A_1886591.Trust-Data
Dataset Card for Trust framework
Description
Repository: https://github.com/declare-lab/trust-align
Paper: https://arxiv.org/abs/2409.11242
Data Summary
The Trust-score evaluation dataset includes the top 100 GTR-retrieved results for ASQA, QAMPARI, and ExpertQA, along with the top 100 BM25-retrieved results for ELI5. The answerability of each question is assessed based on its accompanying documents.
The Trust-align training dataset comprises 19K high-quality… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/Trust-Data.abuse-scanner-bot-datasetBenchmark_2020_B16_Mortgage_Trust_1797288
Benchmark 2020-B16 Mortgage Trust
SEC ABS-EE asset-level filings for CIK 1797288 (Benchmark 2020-B16 Mortgage Trust).
Filings: 55
Parquet files: 215
Total size: 117.6 MB
Reporting period start: 2020-02-11
Reporting period end: 2024-07-11
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).
Filing index… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Benchmark_2020_B16_Mortgage_Trust_1797288.Benchmark_2018_B1_Mortgage_Trust_1722194
Benchmark 2018-B1 Mortgage Trust
SEC ABS-EE asset-level filings for CIK 1722194 (Benchmark 2018-B1 Mortgage Trust).
Filings: 81
Parquet files: 314
Total size: 23.2 MB
Reporting period start: 2018-01-06
Reporting period end: 2024-07-11
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).
Filing index… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Benchmark_2018_B1_Mortgage_Trust_1722194.Trustgen_dataset
Dataset Card for TrustGen 🤖✨
📝 Dataset Summary
This dataset is a component of the TrustGen project, a dynamic and modular benchmarking system designed to systematically evaluate the trustworthiness of Generative Foundation Models (GenFMs).
The dataset facilitates the evaluation of models across text-to-image, large language, and vision-language modalities. It is structured into 17 distinct subsets, each targeting a specific trustworthiness dimension, including:… See the full description on the dataset page: https://huggingface.co/datasets/TrustGen/Trustgen_dataset.Santander_Drive_Auto_Receivables_Trust_2022_5_1941255
Santander Drive Auto Receivables Trust 2022-5
SEC ABS-EE asset-level filings for CIK 1941255 (Santander Drive Auto Receivables Trust 2022-5).
Filings: 38
Parquet files: 1
Total size: 2.7 MB
Reporting period start: 2022-07-31
Reporting period end: 2026-03-31
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Santander_Drive_Auto_Receivables_Trust_2022_5_1941255.UBS_Commercial_Mortgage_Trust_2018_C13_1749360
UBS Commercial Mortgage Trust 2018-C13
SEC ABS-EE asset-level filings for CIK 1749360 (UBS Commercial Mortgage Trust 2018-C13).
Filings: 72
Parquet files: 279
Total size: 46.7 MB
Reporting period start: 2018-10-11
Reporting period end: 2024-07-11
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/UBS_Commercial_Mortgage_Trust_2018_C13_1749360.DAIMLER_TRUST_LEASING_LLC_1537805
Mercedes-Benz Trust Leasing LLC
SEC ABS-EE asset-level filings for CIK 1537805 (Mercedes-Benz Trust Leasing LLC).
Filings: 405
Parquet files: 6
Total size: 10.1 MB
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).
Filing index
cik
form
accessionNumber
url
1537805
ABS-EE… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/DAIMLER_TRUST_LEASING_LLC_1537805.Toyota_Auto_Receivables_2022_C_Owner_Trust_1933877
Toyota Auto Receivables 2022-C Owner Trust
SEC ABS-EE asset-level filings for CIK 1933877 (Toyota Auto Receivables 2022-C Owner Trust).
Filings: 48
Parquet files: 1
Total size: 1.4 MB
Reporting period start: 2022-06-30
Reporting period end: 2026-02-28
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Toyota_Auto_Receivables_2022_C_Owner_Trust_1933877.Santander_Drive_Auto_Receivables_Trust_2022_4_1934902
Santander Drive Auto Receivables Trust 2022-4
SEC ABS-EE asset-level filings for CIK 1934902 (Santander Drive Auto Receivables Trust 2022-4).
Filings: 32
Parquet files: 1
Total size: 2.3 MB
Reporting period start: 2022-06-30
Reporting period end: 2026-03-31
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Santander_Drive_Auto_Receivables_Trust_2022_4_1934902.trustmed-medical-retrieval-2026-07
TrustMed Local Medical Retrieval Corpus (snapshot 2026-07-30)
Local corpus + prebuilt hybrid retrieval indexes for medical search-RL
(Search-R1-style serving; verl-agent/GiGPO-ready). Built 2026-08-15.
Contents
dir
rows
pieces
pubmed/
29,166,067
docs.arrow (mmap doc store), pmids.npy, embeds_fp16.npy (MedCPT 768d), faiss_ivfpq.index (IVF65536,PQ64), bm25s/ (bm25s index), manifest.json, gate_queries.jsonl, pmcid_to_pmid.json
statpearls/
358,500
same… See the full description on the dataset page: https://huggingface.co/datasets/Lyra-stellAI/trustmed-medical-retrieval-2026-07.Morgan_Stanley_Capital_I_Trust_2018_H3_1742383
Morgan Stanley Capital I Trust 2018-H3
SEC ABS-EE asset-level filings for CIK 1742383 (Morgan Stanley Capital I Trust 2018-H3).
Filings: 74
Parquet files: 156
Total size: 10.9 MB
Reporting period start: 2018-07-11
Reporting period end: 2024-07-11
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Morgan_Stanley_Capital_I_Trust_2018_H3_1742383.Toyota_Auto_Receivables_2018_A_Owner_Trust_1725585
Toyota Auto Receivables 2018-A Owner Trust
SEC ABS-EE asset-level filings for CIK 1725585 (Toyota Auto Receivables 2018-A Owner Trust).
Filings: 50
Parquet files: 190
Total size: 475.6 MB
Reporting period start: 2017-12-31
Reporting period end: 2022-01-31
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Toyota_Auto_Receivables_2018_A_Owner_Trust_1725585.geomtl-dataset
GeoMTL Text Annotations Dataset
This dataset provides the cleaned text captions (.clean.txt) used for training the GeoMTL model on satellite image captioning and Visual Question Answering (VQA) tasks for agricultural earth observation.
🔗 Relationship to the Base Dataset
This repository contains only the language targets (text annotations) generated for the Multi-Temporal Crop Classification dataset. To fully utilize this data for multi-task learning (segmentation… See the full description on the dataset page: https://huggingface.co/datasets/trust-tad/geomtl-dataset.trust-chain-freshness
Trust-chain freshness
Two things in the Council of AI estate go out of date on their own, and this dataset is the
receipt that somebody keeps checking them.
1. OpenTimestamps proofs
An OpenTimestamps stamp is created instantly and carries only a pending calendar attestation.
Hours later the calendar's commitment lands in a Bitcoin block — but the published .ots file
only says so once the completed path is fetched back and the file rewritten. Nothing does that on… See the full description on the dataset page: https://huggingface.co/datasets/csoai/trust-chain-freshness.BBCMS_Mortgage_Trust_2022_C14_1901814
BBCMS Mortgage Trust 2022-C14
SEC ABS-EE asset-level filings for CIK 1901814 (BBCMS Mortgage Trust 2022-C14).
Filings: 37
Parquet files: 92
Total size: 13.5 MB
Reporting period start: 2022-02-11
Reporting period end: 2026-02-11
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).
Filing index… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/BBCMS_Mortgage_Trust_2022_C14_1901814.PeerCheck
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
Dataset Summary
PeerCheck is a framework for studying and improving the quality of LLM-generated academic peer reviews.
It contains both human-written reviews and LLM-generated reviews for the same research papers, enabling direct comparison between human and LLM-generated reviewers.
The dataset is used to support research on:
LLM-generated peer review;
Review quality evaluation;… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/PeerCheck.
