datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latam-taxbench-1500
🐝 LatAm-TaxBench 1,500: The Latin American Statutory Tax & Legal Benchmark
Overview
LatAm-TaxBench 1500 is the premier standardized benchmark for evaluating Large Language Models and Agentic Swarms on high-complexity civil law tax jurisprudence and mathematical statutory calculation in Latin America (focusing on Colombian DIAN statutory code).
Composed of 1,500 gold-standard test cases, the dataset rigorously evaluates model performance across four key civil… See the full description on the dataset page: https://huggingface.co/datasets/MauroCaceres1711/latam-taxbench-1500.erc8004-base-census-jun2026
ERC-8004 Base Mainnet Census — June 2026
Read-only census of AI agents registered under ERC-8004 (Trustless Agents) on Base mainnet, taken at block 47,041,190 (June 7, 2026).
Headline numbers (full scan + uniform sample):
54,802 agents registered in the IdentityRegistry (0x8004A1...a432, deployed Feb 3, 2026)
Uniform random sample of 2,000 agents queried against the ReputationRegistry (0x8004BAa1...9b63):
52.8% have at least one feedback client; median 1 client per agent, max… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-base-census-jun2026.latam_nrc
LATAM NRC Dataset
Spanish language samples from MultiNRC classified by country of origin based on dialect markers and cultural references.
Dataset Description
This dataset contains 392 Spanish language samples from the ScaleAI/MultiNRC dataset, each annotated with the likely country of origin. Classification was performed using Gemini Flash and Gemini Pro models analyzing regional vocabulary, grammatical patterns, and cultural references.
Country Distribution… See the full description on the dataset page: https://huggingface.co/datasets/emolero/latam_nrc.latam-slang-telemetry
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/Zynoox-IA/latam-slang-telemetry
Uses
Direct Use
Ideal for fine-tuning LLMs, conversational agents, virtual assistants, or roleplay characters to sound organic and native in Latin American centennial spaces.
Out-of-Scope Use
This dataset should not be used for automated profiling of specific users, as all telemetry has been strictly anonymized… See the full description on the dataset page: https://huggingface.co/datasets/Zynoox-IA/latam-slang-telemetry.latamsrcfixeddata_for_pii_latam
