datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LatAmRT
LatAm-RT: Culturally Adaptive Red Teaming for AI Safety in Latin America
LatAm-RT is a culturally adaptive Red Teaming benchmark for evaluating AI safety in Latin American Spanish. It operationalizes localization intensity as an explicit experimental variable through a Layered Localization Taxonomy (LLT) with three levels of progressive cultural specificity.
The dataset contains 284 evaluation prompts across:
4 risk domains: Fraud & Exploitation · Political & Information Harm ·… See the full description on the dataset page: https://huggingface.co/datasets/Latam26/LatAmRT.ai-brand-visibility-latam
AI Brand Visibility in LATAM — LLM Mention Dataset
Dataset Description
This dataset contains annotated records of brand mentions in Spanish-language LLM responses, collected by FARDO — the first AI brand visibility platform in Latin America.
The dataset accompanies the paper: "AI Brand Visibility in Spanish-Language LLMs: A Framework for Measuring and Optimizing Brand Presence in Generative AI Responses" (Martin & Seguro, 2026).
Dataset Summary
A collection… See the full description on the dataset page: https://huggingface.co/datasets/HeyFardo/ai-brand-visibility-latam.erc8004-simulated-agents
ERC-8004 Simulated Agents — labeled synthetic dataset (6,000 agents)
⚠️ This dataset is fully synthetic. No public labeled dataset of malicious ERC-8004 agents exists (the standard reached mainnet in 2026 and exposes no trust label), so this dataset simulates the feature distributions the three ERC-8004 registries would expose, for training/evaluating trustworthiness models. For real on-chain data see the companion Base mainnet census.
Composition
6,000 agents, 1… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-simulated-agents.
