datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zero-trust-maturity-assessments
Zero Trust Maturity Assessments
Note: This is an independent dataset based on publicly available CISA frameworks. It is not affiliated with, endorsed by, or sponsored by CISA, OMB, or any federal agency.
What This Is
I created this dataset while working on Zero Trust implementations and realized there was a huge gap: no public datasets exist for ZT maturity assessments.
This dataset contains 23 comprehensive Zero Trust assessments based on CISA's official Zero Trust… See the full description on the dataset page: https://huggingface.co/datasets/Reply2susi/zero-trust-maturity-assessments.speechmap-assessments
Speechmap collection
Datasets in this collection are derived from xlr8harder's Speechmap / llm-compliance project. Data has been indexed slightly differently, some columns have been added and others have been removed. Refer to the original Github repo for the full dataset.
The collection includes:
2.4k questions: speechmap-questions
274k responses: speechmap-responses
510k LLM-judge assessments: speechmap-assessments combining the original LLM-assessments from the llm-compliance… See the full description on the dataset page: https://huggingface.co/datasets/PITTI/speechmap-assessments.speechmap-assessments-v3
Speechmap collection
Datasets in this collection are derived from xlr8harder's Speechmap / llm-compliance project. Data has been indexed slightly differently, some columns have been added and others have been removed. Refer to the original Github repo for the full dataset.
The collection includes:
2.4k questions: speechmap-questions
369k responses: speechmap-responses
2.07m LLM-judge assessments: speechmap-assessments combining the original LLM-assessments from the llm-compliance… See the full description on the dataset page: https://huggingface.co/datasets/PITTI/speechmap-assessments-v3.agentic-benchmark-assessmentsarabic-dataset-quality-assessments
Arabic NLP Dataset Quality Assessments
تقييمات جودة شاملة لمجموعات البيانات العربية في معالجة اللغات الطبيعية
Dataset Summary
This dataset contains automated quality assessments of 331 Arabic NLP datasets from the Masader catalog. Each assessment evaluates a dataset across 7 quality dimensions using up to 500 data samples, producing detailed Arabic-language analysis including quality scores, statistical analysis, strengths, weaknesses, and usage recommendations.
The… See the full description on the dataset page: https://huggingface.co/datasets/SalahAbdoNLP/arabic-dataset-quality-assessments.adaption-sos2-ns13-variant-assessments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sos2_ns13_variant_assessments
This dataset contains structured research-level assessments for SOS2 gene variants associated with Noonan Syndrome 13 (NS13). Each entry provides a detailed analysis including assigned priority tiers, investigation scores, and evidence from computational predictors like CADD and AlphaMissense. The content further details population rarity, specific GEF… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-sos2-ns13-variant-assessments.adaption-hr-attrition-risk-assessments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-hr_attrition_risk_assessments
This dataset contains synthetic employee telemetry profiles paired with expert AI-generated attrition risk assessments and managerial intervention plans. Each sample includes detailed demographic and performance metrics, followed by a structured analysis identifying risk levels, primary drivers, and prioritized action timelines. The content is designed… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-hr-attrition-risk-assessments.multi-hop-psychiatric-medical-assessments-test-dataacat-assessments
language:
en
license: apache-2.0
task_categories:
text-classification
text-scoring
task_ids:
text-classification
evaluation
pretty_name: ACAT AI Self-Assessment Dataset
size_categories:
n<1K
tags:
ai-evaluation
alignment
self-assessment
governance
calibration
llm
ACAT: AI Calibrated Assessment Tool Dataset
Dataset Summary
The ACAT (AI Calibrated Assessment Tool) dataset is a structured benchmark for evaluating AI system self-assessment calibration and… See the full description on the dataset page: https://huggingface.co/datasets/HumanAIOS/acat-assessments.adaption-braf-cfc-noonan-tier1-assessments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-braf_cfc_noonan_tier1_assessments
This dataset contains structured research-level variant assessments for BRAF missense mutations associated with Cardio-Facio-Cutaneous (CFC) and Noonan Syndromes. Each entry preserves source-derived Tier 1 prioritization labels and investigation scores while summarizing computational evidence (CADD, AlphaMissense), population frequency, and protein… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-braf-cfc-noonan-tier1-assessments.multi-hop-psychiatric-medical-assessmentsadaption-ptpn11-tier1-variant-assessments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ptpn11_tier1_variant_assessments
This dataset contains research-level assessments of PTPN11 missense variants assigned to 'Tier 1' priority based on strict computational evidence filters. Each entry provides structured interpretations including CADD PHRED scores, AlphaMissense predictions, gnomAD frequencies, and functional domain contexts for the SHP-2 protein. The content focuses on… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-ptpn11-tier1-variant-assessments.africa-synth-education-primary-assessments-nigeria
Africa Synth Education Primary Assessments Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-primary-assessments-nigeria.adaption-raf1-noonan-tier1-assessments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-raf1_noonan_tier1_assessments
This dataset contains structured research prioritization assessments for high-priority RAF1 missense variants associated with Noonan Syndrome 5. Each entry details computational evidence (CADD, AlphaMissense), population frequency, and structural context for variants assigned to Tier 1 with specific investigation scores. The content strictly preserves… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-raf1-noonan-tier1-assessments.adaption-hr-attrition-risk-assessments-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-hr_attrition_risk_assessments
This dataset contains synthetic employee telemetry profiles paired with expert AI-generated attrition risk assessments and managerial intervention plans. Each sample includes detailed demographic and performance metrics, followed by a structured analysis identifying risk levels, primary drivers, and prioritized action timelines. The content is designed… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-hr-attrition-risk-assessments-v1.speechmap-assessments-v2
Speechmap collection
Datasets in this collection are derived from xlr8harder's Speechmap / llm-compliance project. Data has been indexed slightly differently, some columns have been added and others have been removed. Refer to the original Github repo for the full dataset.
The collection includes:
2.4k questions: speechmap-questions
336k responses: speechmap-responses
875k LLM-judge assessments: speechmap-assessments combining the original LLM-assessments from the llm-compliance… See the full description on the dataset page: https://huggingface.co/datasets/PITTI/speechmap-assessments-v2.assessments
