CoolFace
20 results

typhoon

typhoon-ai /thai_exam Dataset Card for Thai_Exam ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows: ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.tabularquestion-answeringn<1K19 likes2.2k downloads2y agoHugging Facetorchgeo /digital_typhoonDigitial Typhoon Dataset: KITAMOTO, A., HWANG, J., VUILLOD, B., GAUTIER, L., TIAN, Y., & CLANUWAT, T. (2023, December). Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones. NeurIPS 2023 Datasets and Benchmarks (Spotlight). This dataset was created by the Digital Typhoon project. image100K<n<1M2 likes951 downloads3y agoHugging Facetyphoon-ai /ThaiOCRBench ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai ThaiOCRBench is the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks.Inspired by OCRBench v2, it contains 2,808 human-annotated samples across 13 diverse tasks, including table parsing, chart understanding, full-page OCR, key information extraction, and visual question answering. The benchmark enables standardized zero-shot… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiOCRBench.imageimage-text-to-text1K<n<10K7 likes795 downloads10mo agoHugging Facetyphoon-ai /thai-dialect-isan-dataset Dataset Card for Thai Dialect Isan Speech Corpus Dataset Description This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language. The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.textautomatic-speech-recognition10K<n<100K5 likes504 downloads10mo agoHugging Facetyphoon-ai /avhallubench Dataset Card for AVHalluBench The dataset is for benchmarking hallucination levels in audio-visual LLMs. It consists of 175 videos and each video has hallucination-free audio and visual descriptions. The statistics are provided in the figure below, and more information can be found in our paper. Paper: CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models Multimodal Hallucination Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/avhallubench.videon<1K7 likes396 downloads2y agoHugging Faceopen-llm-leaderboard-old /details_scb10x__llama-3-typhoon-v1.5-8b0 likes266 downloads2y agoHugging Face