typhoon
Datasets
All datasets matching “typhoon”thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.digital_typhoonDigitial Typhoon Dataset:
KITAMOTO, A., HWANG, J., VUILLOD, B., GAUTIER, L., TIAN, Y., & CLANUWAT, T. (2023, December). Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones. NeurIPS 2023 Datasets and Benchmarks (Spotlight).
This dataset was created by the Digital Typhoon project.
ThaiOCRBench
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
ThaiOCRBench is the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks.Inspired by OCRBench v2, it contains 2,808 human-annotated samples across 13 diverse tasks, including table parsing, chart understanding, full-page OCR, key information extraction, and visual question answering.
The benchmark enables standardized zero-shot… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiOCRBench.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.avhallubench
Dataset Card for AVHalluBench
The dataset is for benchmarking hallucination levels in audio-visual LLMs. It consists of 175 videos and each video has hallucination-free audio and visual descriptions. The statistics are provided in the figure below, and more information can be found in our paper.
Paper: CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
Multimodal Hallucination Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/avhallubench.details_scb10x__llama-3-typhoon-v1.5-8b
