datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard-results
mauroibz/leaderboard-results
Results from model evaluations on the leaderboard
This dataset contains evaluation results from the leaderboard system.
Structure
Each JSON file contains results for a specific model evaluation
Files are organized by organization/model structure
Each result file includes:
Model configuration
Evaluation results across different benchmarks
Metadata about the evaluation run
Usage
These results are used by the… See the full description on the dataset page: https://huggingface.co/datasets/LatamBoard/leaderboard-results.red_pajama_es_hq
RedPajama's High Quality Spanish subset
What is this?
The following is a high-quality dataset distilled from the Spanish subsection of RedPajama-Data-v2, created using the methodology proposed in FineWEB-Edu.
Usage
from datasets import load_dataset
ds = load_dataset("latam-gpt/red_pajama_es_hq")
Filtering by quality score
Documents in this corpus are scored on academic quality from 2.5 to 5, with higher scores indicating better quality. The… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/red_pajama_es_hq.fineweb2-spa_Latn-eduLatamGPT-Corpus-1.0
LatamGPT-Corpus-1.0
🌐 Language versions: English | Español | Português
🔗 Project links: Official LatamGPT website | Corpus dashboard
🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0.
Dataset description
Summary
LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.latam-spanish-speech-orpheus-tts-24khz
LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready)
Dataset Description
This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate.
The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.latamqa_mcq_es-la
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_mcq_es-la.CHOCLO
🌽 CHOCLO: Latin American Cultural Knowledge Benchmark
Description
CHOCLO is a benchmark designed to evaluate cultural knowledge in language models, with a specific focus on entities representative of Latin America. Unlike traditional benchmarks, which often emphasize general knowledge or contexts dominated by English-language data, CHOCLO aims to capture the richness, diversity, and specificity of Latin American cultural knowledge, including traditions, gastronomy… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/CHOCLO.LATAM-Egocentric-Residential-IMU
LATAM Egocentric Residential (with IMU)
Head-mounted, first-person video of everyday household chores recorded across Brazil, Argentina, Venezuela and Peru, each paired with a ~100 Hz accelerometer + gyroscope IMU stream.
The dataset targets embodied-AI and robotics research that needs real, unscripted human manipulation in cluttered domestic environments — not lab-staged demonstrations.
Preview: LATAM_OD_D_16 — gardening, outdoor, daytime, Argentina (30 s excerpt, downscaled… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/LATAM-Egocentric-Residential-IMU.latam-xixgoogle-latam-spanish-uniform-vad
Google LATAM Spanish — Uniform VAD and 0.5 s Edges
This public derivative contains 15,016 Latin American Spanish utterances from
the Google crowdsourced TTS datasets packaged by ylacombe. Female and male
audio were freshly exported from the same pinned upstream revisions and passed
through exactly the same processing pipeline.
Configurations
Configuration
Train
Validation
Total
Hours including edge padding
argentina-female
3,542
379
3,921
4.181… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-uniform-vad.google-latam-spanish-boundary-normalized
Google LATAM Spanish Boundary-Normalized Audio
Female Spanish speech from the following upstream datasets:
Argentina: ylacombe/google-argentinian-spanish
Chile: ylacombe/google-chilean-spanish
Colombia: ylacombe/google-colombian-spanish
Attribution and Thanks
Many thanks to ylacombe for publishing
and maintaining the original Argentinian, Chilean, and Colombian Spanish
datasets. The recordings, transcripts, speaker labels, and original dataset
structure come… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-boundary-normalized.es-ultrachatThis is an automatic spanish translation of Huggingface's ultrachat 200k done with Llama Instruct 3.1 70b.
LATAM-High-Fidelity-ASR
Dataset Overview
This dataset contains high-quality conversational audio samples curated for Automatic Speech Recognition tasks in Spanish variants and Portugese.
The dataset includes:
Paired audio + transcripts
Natural, non-scripted conversational speech
Single Speaker & Dual-speaker interactions
Audio Specifications
Sampling Rate: 16 kHz – 24 kHz
Bit Depth: 16-bit
Audio Type: Non-scripted conversational speech
Supported Languages
Language… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/LATAM-High-Fidelity-ASR.personas-instruct-messageses-smoltalkThis is a spanish translation of the first 150k rows of Huggingface's smoltalk, a synthetic dataset for supervised finetuning.
translated-tulu-3-sft-olmo-2-mixture-0225Automated translation of allenai/tulu-3-sft-olmo-2-mixture-0225 into Spanish.
Topic Distribution
Mathematics (237,191 samples - 33.6%)
allenai/tulu-3-sft-personas-math-filtered: 80,115 samples
ai2-adapt-dev/numinamath_tir_math_decontaminated: 61,550 samples
ai2-adapt-dev/tulu_v3.9_open_math_2_gsm8k_50k: 49,873 samples
allenai/tulu-3-sft-personas-math-grade-filtered: 46,741 samples
ai2-adapt-dev/tulu_v3.9_personahub_math_interm_algebra_20k: 19,912 samples… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/translated-tulu-3-sft-olmo-2-mixture-0225.tulu-3-dpo-spanish-ba3fe66cgretel-safety-alignment-es-v1gretelai/gretel-safety-alignment-en-v1 but with the prompt, response, and safe_response fields translated to spanish. Translation was done using gpt4o-mini for most rows, and mistral-small-3.2-24b-instruct for those gpt4o-mini refused to translate.
Some rows in gretelai/gretel-safety-alignment-en-v1 contained model refusals in the prompt or response fields. Those were filtered out, and thus are not included in this dataset.
latam-real-estate-listings
HF LatAm Real Estate Dataset (v5)
Export of listed real-estate properties from Costa Rica and Panamá for Hugging Face
consumption (NLP, modeling, search, recommendation, price prediction, etc.).
Version: v5
Sources:
OpenSearch properties — Costa Rica (CR)
OpenSearch pa_properties — Panamá (PA)
Filters: status=listed, listing_date <= cutoff (default ~5 months old), must have price + (built or fallback) area
Rows: one row per (property, operation) side. sale_and_rental expands to… See the full description on the dataset page: https://huggingface.co/datasets/weknowinc/latam-real-estate-listings.latam-taxbench-1500
🐝 LatAm-TaxBench 1,500: The Latin American Statutory Tax & Legal Benchmark
Overview
LatAm-TaxBench 1500 is the premier standardized benchmark for evaluating Large Language Models and Agentic Swarms on high-complexity civil law tax jurisprudence and mathematical statutory calculation in Latin America (focusing on Colombian DIAN statutory code).
Composed of 1,500 gold-standard test cases, the dataset rigorously evaluates model performance across four key civil… See the full description on the dataset page: https://huggingface.co/datasets/MauroCaceres1711/latam-taxbench-1500.latam-urban-driving-sample
LATAM Urban Driving Dataset — Night / Low-Visibility / Potholes Sample (Intel RealSense D555)
A free preview of a multimodal, edge-case-focused urban driving dataset captured in the State of Mexico (Edomex) with the Intel RealSense D555 (RGB + IR + metric depth), aimed at Autonomous Driving, ADAS, and Smart City perception teams.
This sample: 641 contiguous night-driving frames, densely annotated with 562 pothole boxes measured in native 3D depth and vehicle classes… See the full description on the dataset page: https://huggingface.co/datasets/ysilabs/latam-urban-driving-sample.Dolci-Instruct-SFT-No-Tools-No-Hardcoded-Sample-200kIF-MT-en-es-latamTest data for translation with instruction-following used in the Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs paper (English-Spanish (Latin America)).
Dataset created with the Zero-shot Benchmarking framework.
LatamGPT-Corpus-1.0-research
LatamGPT-Corpus-1.0-research
🌐 Language versions: English | Español | Português
🔗 Project links: Official LatamGPT website | Corpus dashboard
🔓 Open portion: the openly released part of this corpus is published separately as LatamGPT-Corpus-1.0.
⚠️ Controlled Access Level (Blue – Research)
This repository is not openly available. It holds data destined
exclusively for scientific and academic research, and is managed as a
restricted repository under… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0-research.tulu-3-dpo-spanish-3495f9c9tulu-3-sft-mixture-no-identityamericasnlp-mt-tradu-latamFake-news-latam-omdena
Dataset Card for Fake-news-latam-omdena
Dataset Summary
Since the Cambridge Analytica scandal a pandora box has been opened around the world, bringing to light campaigns even involving our current Latinamerica leaders manipulating public opinion through social media to win an election. There is a common and simple pattern that includes platforms such as facebook and fake news, where the candidates are able to build a nefarious narrative for their own benefit. This fact is… See the full description on the dataset page: https://huggingface.co/datasets/IsaacRodgz/Fake-news-latam-omdena.erc8004-base-census-jun2026
ERC-8004 Base Mainnet Census — June 2026
Read-only census of AI agents registered under ERC-8004 (Trustless Agents) on Base mainnet, taken at block 47,041,190 (June 7, 2026).
Headline numbers (full scan + uniform sample):
54,802 agents registered in the IdentityRegistry (0x8004A1...a432, deployed Feb 3, 2026)
Uniform random sample of 2,000 agents queried against the ReputationRegistry (0x8004BAa1...9b63):
52.8% have at least one feedback client; median 1 client per agent, max… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-base-census-jun2026.
