CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads4h agoHugging Face02SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes515 downloads4mo agoHugging Face03freococo /thahabiorg_metadata 📖 Thahabi Books Metadata Dataset This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book and includes bibliographic information such as title, author, category, and source details. 📦 Dataset Structure This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book with full bibliographic and structural information. 📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.tabulartext-generation1M<n<10M0 likes158 downloads3mo agoHugging Face04ThatOneShortGuy /SongLyricsDataset contains songs by artists, the names of the songs, the lyrics of the songs, the release date, the cover photo, and the general popularity of the song. imagetext-generation100K<n<1M4 likes150 downloads3y agoHugging Face05nlp-chula /ThaiTrees ThaiTrees A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U. Dataset Summary ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.tabulartext-generation1M<n<10M1 likes48 downloads14h agoHugging Face06airesearch /wangchanx-seed-free-synthetic-instruct-thai-120k Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k Dataset Summary This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.tabulartext-generation100K<n<1M3 likes46 downloads2y agoHugging Face07thangquang09 /agentic-coding-traces Agentic Coding Mooncake Traces Synthetic agentic coding benchmark datasets in Mooncake trace (JSONL) format, generated with AIPerf 0.9.0 for LLM inference benchmarking. Designed for use with InferenceX via the agentic-replay scenario-type and aiperf_adapter.py. Files File Sessions Turns max_prompt_tokens Seed 64k/dataset.jsonl 1,000 18,595 65,536 42 128k/dataset.jsonl 1,000 16,957 131,072 42 Format Each line is a Mooncake trace… See the full description on the dataset page: https://huggingface.co/datasets/thangquang09/agentic-coding-traces.tabulartext-generation10K<n<100K0 likes23 downloads4mo agoHugging Face08SPAISS6F1 /spai-ss6-corpus-thai-exam-qa-answers SPAI SS6 Thai Exam QA With Answers Index Index repo for normalized Thai O-NET and exam question-answer records with answer keys. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_exam_qa_with_answers Rows in canonical config: 8,191 Parquet size in canonical config: 0.01 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-exam-qa-answers.tabulartext-generationn<1K0 likes23 downloads4mo agoHugging Face09airesearch /thai-bar-exam-judging Thai Bar-Exam Judging Corpus Anonymised free-form Thai legal essays from a bar-exam preparation exercise, with three Bar Council-trained examiners scoring every essay and span-anchored inline commentary on roughly two thirds of the answers. Eight LLM examinees took the same exam under the same conditions; their answers were graded blind by the same examiners. Fifteen of the 150 answers were cross-graded by the two non-primary examiners, producing the 3-rater stability subset that… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-bar-exam-judging.tabulartext-classification1K<n<10K0 likes19 downloads4mo agoHugging Face10SPAISS6F1 /spai-ss6-corpus-mental-health-thai SPAI SS6 Thai Mental Health Index Index repo for the imported Thai mental-health dataset config. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: mental_health_thai Rows in canonical config: 21,544 Parquet size in canonical config: 0.02 GB Source license: unknown License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-mental-health-thai.tabulartext-generationn<1K0 likes18 downloads4mo agoHugging Face11SPAISS6F1 /spai-ss6-corpus-thai-instruction-sft-suraponn SPAI SS6 Thai Instruction SFT Suraponn Index Index repo for the Suraponn Thai instruction SFT dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_instruction_sft_suraponn Rows in canonical config: 131,907 Parquet size in canonical… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-instruction-sft-suraponn.tabulartext-generationn<1K0 likes10 downloads4mo agoHugging Face12SPAISS6F1 /spai-ss6-corpus-thai-wikipedia-clean SPAI SS6 Thai Wikipedia Clean Corpus Index Index repo for the Thai Wikipedia clean corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_wikipedia_clean_20230101 Rows in canonical config: 1,436,054 Parquet size in canonical config: 0.26 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-wikipedia-clean.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face13SPAISS6F1 /spai-ss6-corpus-thai-synthetic-qa-v1 SPAI SS6 Thai Synthetic QA V1 Index Index repo for the ThaiSyntheticQA v1 dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_synthetic_qa_v1 Rows in canonical config: 12,668 Parquet size in canonical config: 0.02 GB Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synthetic-qa-v1.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face14SPAISS6F1 /spai-ss6-corpus-thai-idioms-instruction SPAI SS6 Thai Idioms Instruction Index Index repo for the Thai idioms instruction dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_idioms_instruction Rows in canonical config: 1,152 Parquet size in canonical config: 0.00 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-idioms-instruction.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face15SPAISS6F1 /spai-ss6-corpus-thai-medical-care-aiotx SPAI SS6 Thai Medical Care AIoTx Index Index repo for the imported AIoTx Thai medical-care dataset config. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_medical_care_aiotx Rows in canonical config: 3,599 Parquet size in canonical config: 0.00 GB Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-medical-care-aiotx.tabulartext-generationn<1K0 likes8 downloads4mo agoHugging Face16SPAISS6F1 /spai-ss6-corpus-thai-local-instruction-v2 SPAI SS6 Thai Local Instruction V2 Index Index repo for the Thai local instruction v2 dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_local_instruction_v2 Rows in canonical config: 39,829 Parquet size in canonical config: 0.00 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-local-instruction-v2.tabulartext-generationn<1K0 likes6 downloads4mo agoHugging Face17SPAISS6F1 /spai-ss6-corpus-thaisum SPAI SS6 ThaiSum Corpus Index Index repo for the ThaiSum corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thaisum Rows in canonical config: 391,868 Parquet size in canonical config: 1.30 GB Source license: mit License review status:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaisum.tabulartext-generationn<1K0 likes5 downloads4mo agoHugging Face18SPAISS6F1 /spai-ss6-corpus-thai-synonym-instruction SPAI SS6 Thai Synonym Instruction Index Index repo for the Thai synonym instruction dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_synonym_instruction Rows in canonical config: 167 Parquet size in canonical config: 0.00 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synonym-instruction.tabulartext-generationn<1K0 likes4 downloads4mo agoHugging Face19SPAISS6F1 /spai-ss6-corpus-khanomtanllm-thai-subset SPAI SS6 KhanomTanLLM Thai Subset Index Index repo for the KhanomTanLLM Thai subset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: khanomtanllm_thai_subset Rows in canonical config: 464,339 Parquet size in canonical config: 1.80 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-khanomtanllm-thai-subset.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face20SPAISS6F1 /spai-ss6-corpus-thai-culturax-clean SPAI SS6 Thai CulturaX Clean Corpus Index Index repo for the Thai CulturaX clean corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_culturax_clean Rows in canonical config: 818,727 Parquet size in canonical config: 1.96 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-culturax-clean.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face21SPAISS6F1 /spai-ss6-corpus-thai-tourist-attraction-instruction SPAI SS6 Thai Tourist Attraction Instruction Index Index repo for the Thai tourist-attraction instruction dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_tourist_attraction_instruction Rows in canonical config: 51,662 Parquet size… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-tourist-attraction-instruction.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face22SPAISS6F1 /spai-ss6-corpus-thaiqa-lst20-mit SPAI SS6 ThaiQA LST20 MIT Index Index repo for the ThaiQA LST20 dataset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thaiqa_lst20_mit Rows in canonical config: 7,643 Parquet size in canonical config: 0.01 GB Source license: mit License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaiqa-lst20-mit.tabulartext-generationn<1K0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.