datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latam-taxbench-1500
🐝 LatAm-TaxBench 1,500: The Latin American Statutory Tax & Legal Benchmark
Overview
LatAm-TaxBench 1500 is the premier standardized benchmark for evaluating Large Language Models and Agentic Swarms on high-complexity civil law tax jurisprudence and mathematical statutory calculation in Latin America (focusing on Colombian DIAN statutory code).
Composed of 1,500 gold-standard test cases, the dataset rigorously evaluates model performance across four key civil… See the full description on the dataset page: https://huggingface.co/datasets/MauroCaceres1711/latam-taxbench-1500.latam_nrc
LATAM NRC Dataset
Spanish language samples from MultiNRC classified by country of origin based on dialect markers and cultural references.
Dataset Description
This dataset contains 392 Spanish language samples from the ScaleAI/MultiNRC dataset, each annotated with the likely country of origin. Classification was performed using Gemini Flash and Gemini Pro models analyzing regional vocabulary, grammatical patterns, and cultural references.
Country Distribution… See the full description on the dataset page: https://huggingface.co/datasets/emolero/latam_nrc.latam-slang-telemetry
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/Zynoox-IA/latam-slang-telemetry
Uses
Direct Use
Ideal for fine-tuning LLMs, conversational agents, virtual assistants, or roleplay characters to sound organic and native in Latin American centennial spaces.
Out-of-Scope Use
This dataset should not be used for automated profiling of specific users, as all telemetry has been strictly anonymized… See the full description on the dataset page: https://huggingface.co/datasets/Zynoox-IA/latam-slang-telemetry.latamsrcfixeddata_for_pii_latam
