datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.latamqa_mcq_es-la
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_mcq_es-la.CHOCLO
🌽 CHOCLO: Latin American Cultural Knowledge Benchmark
Description
CHOCLO is a benchmark designed to evaluate cultural knowledge in language models, with a specific focus on entities representative of Latin America. Unlike traditional benchmarks, which often emphasize general knowledge or contexts dominated by English-language data, CHOCLO aims to capture the richness, diversity, and specificity of Latin American cultural knowledge, including traditions, gastronomy… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/CHOCLO.LatAmRT
LatAm-RT: Culturally Adaptive Red Teaming for AI Safety in Latin America
LatAm-RT is a culturally adaptive Red Teaming benchmark for evaluating AI safety in Latin American Spanish. It operationalizes localization intensity as an explicit experimental variable through a Layered Localization Taxonomy (LLT) with three levels of progressive cultural specificity.
The dataset contains 284 evaluation prompts across:
4 risk domains: Fraud & Exploitation · Political & Information Harm ·… See the full description on the dataset page: https://huggingface.co/datasets/Latam26/LatAmRT.ai-brand-visibility-latam
AI Brand Visibility in LATAM — LLM Mention Dataset
Dataset Description
This dataset contains annotated records of brand mentions in Spanish-language LLM responses, collected by FARDO — the first AI brand visibility platform in Latin America.
The dataset accompanies the paper: "AI Brand Visibility in Spanish-Language LLMs: A Framework for Measuring and Optimizing Brand Presence in Generative AI Responses" (Martin & Seguro, 2026).
Dataset Summary
A collection… See the full description on the dataset page: https://huggingface.co/datasets/HeyFardo/ai-brand-visibility-latam.erc8004-simulated-agents
ERC-8004 Simulated Agents — labeled synthetic dataset (6,000 agents)
⚠️ This dataset is fully synthetic. No public labeled dataset of malicious ERC-8004 agents exists (the standard reached mainnet in 2026 and exposes no trust label), so this dataset simulates the feature distributions the three ERC-8004 registries would expose, for training/evaluating trustworthiness models. For real on-chain data see the companion Base mainnet census.
Composition
6,000 agents, 1… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-simulated-agents.latamqa_mcq_pt-br
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_mcq_pt-br.latamqa_mcq_es-es
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_mcq_es-es.latamqa_articles_es-es
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_articles_es-es.latamqa_articles_pt-br
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_articles_pt-br.
