CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zouhar /trust-interventionThis is a slightly edited dataset of the one found here on GitHub. The data contains the user interactions, their bet values, answer correctness etc. Please contact the authors if you have any questions. A Diachronic Perspective on User Trust in AI under Uncertainty Abstract: In a human-AI collaboration, users build a mental model of the AI system based on its veracity and how it presents its decision, e.g. its presentation of system confidence and an explanation of the output.… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/trust-intervention.tabulartabular-regression10K<n<100K1 likes99 downloads3y agoHugging Face02AIJian /TrustSQL-data TrustSQL-data Training data for TRUST-SQL, a tool-integrated multi-turn reinforcement-learning framework for Text-to-SQL over Unknown Schemas. Dataset summary The dataset supports the two-stage TrustSQL training pipeline: SFT data: approximately 9.2k structured interaction demonstrations. RL data: approximately 11.6k samples used for Phase-Aware GRPO optimization. The examples teach an agent to explore database metadata, propose a verified schema subset… See the full description on the dataset page: https://huggingface.co/datasets/AIJian/TrustSQL-data.tabulartext-generation10K<n<100K0 likes56 downloads29d agoHugging Face03jang1563 /protein-structure-trust-benchmark Protein-Structure Trust-Routing Benchmark (Boltz-2) Leakage-controlled benchmarks for confidence-calibrated trust routing over a protein-structure predictor: given a specialist model's confidence (Boltz-2 ipTM / pLDDT) for a target, decide whether to trust the prediction or pay to verify it — and score that decision against experimentally-measured correctness. Evaluation substrate for the report "When does an LLM trust a specialist model? A cost-aware trust-routing audit"… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/protein-structure-trust-benchmark.tabulartabular-classificationn<1K0 likes55 downloads15d agoHugging Face04nicolas-brieuc /dataset-trust-auditor-events Dataset Trust Auditor — Audit Events Public audit trail produced by the Dataset Trust Auditor — a two-phase AI pipeline that scores HuggingFace datasets across 8 trust dimensions. Every audit run appends one row. The dataset grows over time as users audit datasets through the deployed app. Dataset Structure Each row is one completed audit of a HuggingFace dataset. Column Type Description audit_id string UUID for this audit run url string Full HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-brieuc/dataset-trust-auditor-events.tabularn<1K0 likes41 downloads5mo agoHugging Face05wakeupmh /repro-when-to-trust-the-cheap-check-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes31 downloads2mo agoHugging Face06amazon-agi /AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset Adversarial Arena: Trusted AI Challenge Dataset Dataset Description This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs. Papers: Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.tabulartext-generation10K<n<100K0 likes23 downloads3mo agoHugging Face07trustmodsm /TrustMod-SM TrustMod-SM: A Multi-Axis Benchmark for Evaluating Trustworthiness of LLMs in Social Media Content Moderation Dataset Description TrustMod-SM is a unified trustworthiness benchmark for evaluating LLM-based social media content moderators across five dimensions: trustfulness, fairness, safety, robustness, and context integrity. The benchmark comprises 28,792 evaluation instances curated from eight established datasets, covering six demographic attributes (race, gender… See the full description on the dataset page: https://huggingface.co/datasets/trustmodsm/TrustMod-SM.imagetext-classification10K<n<100K0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.