datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indonlu-eval-gpt4o-vs-sealionv3-round1
Local vs Global: Testing GPT-4o-mini and SEA-LIONv3 on Bahasa Indonesia
A benchmark dataset comparing GPT-4o-mini and SEA-LIONv3 on 50 Indonesian-specific questions.This is Round 1 of the INDONLU Eval series, which was built to test LLM performance on culturally grounded, linguistically diverse Southeast Asian prompts.
Overview
We tested 50 prompts across four core categories to assess how well large language models can handle local Indonesian context:
Language –… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/indonlu-eval-gpt4o-vs-sealionv3-round1.indonlu-eval-sealionv3-vs-sahabataiv1-round2
Benchmarking Bahasa Indonesia LLMs: SEA-LIONv3 vs SahabatAI-v1
Following our first benchmarking round, this dataset compares SEA-LIONv3 and SahabatAI-v1 on 50 carefully crafted Indonesian-language tasks. Both models are regionally fine-tuned for Southeast Asian content and evaluated on linguistic fluency, domain-specific accuracy, geographic knowledge, and cultural reasoning.
This is Round 2 of SUPA AI's INDONLU Eval series, which aims to benchmark LLMs for Southeast Asia in… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/indonlu-eval-sealionv3-vs-sahabataiv1-round2.
