CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads5h agoHugging Face02SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes515 downloads4mo agoHugging Face03nlp-chula /ThaiTrees ThaiTrees A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U. Dataset Summary ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.tabulartext-generation1M<n<10M1 likes48 downloads15h agoHugging Face04airesearch /wangchanx-seed-free-synthetic-instruct-thai-120k Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k Dataset Summary This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.tabulartext-generation100K<n<1M3 likes46 downloads2y agoHugging Face05SPAISS6F1 /spai-ss6-corpus-thai-exam-qa-answers SPAI SS6 Thai Exam QA With Answers Index Index repo for normalized Thai O-NET and exam question-answer records with answer keys. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_exam_qa_with_answers Rows in canonical config: 8,191 Parquet size in canonical config: 0.01 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-exam-qa-answers.tabulartext-generationn<1K0 likes23 downloads4mo agoHugging Face06airesearch /thai-bar-exam-judging Thai Bar-Exam Judging Corpus Anonymised free-form Thai legal essays from a bar-exam preparation exercise, with three Bar Council-trained examiners scoring every essay and span-anchored inline commentary on roughly two thirds of the answers. Eight LLM examinees took the same exam under the same conditions; their answers were graded blind by the same examiners. Fifteen of the 150 answers were cross-graded by the two non-primary examiners, producing the 3-rater stability subset that… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-bar-exam-judging.tabulartext-classification1K<n<10K0 likes19 downloads4mo agoHugging Face07SPAISS6F1 /spai-ss6-corpus-mental-health-thai SPAI SS6 Thai Mental Health Index Index repo for the imported Thai mental-health dataset config. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: mental_health_thai Rows in canonical config: 21,544 Parquet size in canonical config: 0.02 GB Source license: unknown License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-mental-health-thai.tabulartext-generationn<1K0 likes18 downloads4mo agoHugging Face08SPAISS6F1 /spai-ss6-corpus-thai-instruction-sft-suraponn SPAI SS6 Thai Instruction SFT Suraponn Index Index repo for the Suraponn Thai instruction SFT dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_instruction_sft_suraponn Rows in canonical config: 131,907 Parquet size in canonical… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-instruction-sft-suraponn.tabulartext-generationn<1K0 likes10 downloads4mo agoHugging Face09SPAISS6F1 /spai-ss6-corpus-thai-wikipedia-clean SPAI SS6 Thai Wikipedia Clean Corpus Index Index repo for the Thai Wikipedia clean corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_wikipedia_clean_20230101 Rows in canonical config: 1,436,054 Parquet size in canonical config: 0.26 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-wikipedia-clean.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face10SPAISS6F1 /spai-ss6-corpus-thai-synthetic-qa-v1 SPAI SS6 Thai Synthetic QA V1 Index Index repo for the ThaiSyntheticQA v1 dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_synthetic_qa_v1 Rows in canonical config: 12,668 Parquet size in canonical config: 0.02 GB Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synthetic-qa-v1.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face11SPAISS6F1 /spai-ss6-corpus-thai-idioms-instruction SPAI SS6 Thai Idioms Instruction Index Index repo for the Thai idioms instruction dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_idioms_instruction Rows in canonical config: 1,152 Parquet size in canonical config: 0.00 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-idioms-instruction.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face12SPAISS6F1 /spai-ss6-corpus-thai-medical-care-aiotx SPAI SS6 Thai Medical Care AIoTx Index Index repo for the imported AIoTx Thai medical-care dataset config. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_medical_care_aiotx Rows in canonical config: 3,599 Parquet size in canonical config: 0.00 GB Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-medical-care-aiotx.tabulartext-generationn<1K0 likes8 downloads4mo agoHugging Face13SPAISS6F1 /spai-ss6-corpus-thai-local-instruction-v2 SPAI SS6 Thai Local Instruction V2 Index Index repo for the Thai local instruction v2 dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_local_instruction_v2 Rows in canonical config: 39,829 Parquet size in canonical config: 0.00 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-local-instruction-v2.tabulartext-generationn<1K0 likes6 downloads4mo agoHugging Face14SPAISS6F1 /spai-ss6-corpus-thaisum SPAI SS6 ThaiSum Corpus Index Index repo for the ThaiSum corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thaisum Rows in canonical config: 391,868 Parquet size in canonical config: 1.30 GB Source license: mit License review status:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaisum.tabulartext-generationn<1K0 likes5 downloads4mo agoHugging Face15SPAISS6F1 /spai-ss6-corpus-thai-synonym-instruction SPAI SS6 Thai Synonym Instruction Index Index repo for the Thai synonym instruction dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_synonym_instruction Rows in canonical config: 167 Parquet size in canonical config: 0.00 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synonym-instruction.tabulartext-generationn<1K0 likes4 downloads4mo agoHugging Face16SPAISS6F1 /spai-ss6-corpus-khanomtanllm-thai-subset SPAI SS6 KhanomTanLLM Thai Subset Index Index repo for the KhanomTanLLM Thai subset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: khanomtanllm_thai_subset Rows in canonical config: 464,339 Parquet size in canonical config: 1.80 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-khanomtanllm-thai-subset.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face17SPAISS6F1 /spai-ss6-corpus-thai-culturax-clean SPAI SS6 Thai CulturaX Clean Corpus Index Index repo for the Thai CulturaX clean corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_culturax_clean Rows in canonical config: 818,727 Parquet size in canonical config: 1.96 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-culturax-clean.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face18SPAISS6F1 /spai-ss6-corpus-thai-tourist-attraction-instruction SPAI SS6 Thai Tourist Attraction Instruction Index Index repo for the Thai tourist-attraction instruction dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_tourist_attraction_instruction Rows in canonical config: 51,662 Parquet size… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-tourist-attraction-instruction.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face19SPAISS6F1 /spai-ss6-corpus-thaiqa-lst20-mit SPAI SS6 ThaiQA LST20 MIT Index Index repo for the ThaiQA LST20 dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thaiqa_lst20_mit Rows in canonical config: 7,643 Parquet size in canonical config: 0.01 GB Source license: mit License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaiqa-lst20-mit.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.