datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.thai-commoncrawl-index
Thai Common Crawl Index (2019–2026)
An index of every page Common Crawl detected as Thai across 70 monthly crawls, from
January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30).
932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains
Each row records where the page lives inside Common Crawl's WARC archives — file name,
byte offset, and record length — so you can fetch exactly the pages you want with HTTP
range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.SKILLRET
SkillRet Benchmark
📄 Technical report: SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents (arXiv:2605.05726)
Dataset Overview
SkillRet is a retrieval benchmark for matching natural-language user requests to agent skills. It contains a curated library of public agent skills from GitHub with synthetic training and evaluation queries.
Dataset Statistics
Metric
Value
Total Records
218,157
Total File Size
714 MB
Total… See the full description on the dataset page: https://huggingface.co/datasets/ThakiCloud/SKILLRET.datacomp-medium-pool-translatedthai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.Thai-understanding
Thai-Understanding: Thai-SUP & XLSR-Thai
Overview
Thai-Understanding is an open-source repository that provides a solution for speech understanding in the Thai language. This repository includes:
Thai-SUP: The first open-source Thai speech understanding dataset, which includes over 1,000 hours of data across three tasks: Intent Classification (IC), Named Entity Recognition (NER), and Speech Rephrasing (SR).
XLSR-Thai: The first large-scale self-supervised learning (SSL)… See the full description on the dataset page: https://huggingface.co/datasets/mcshao/Thai-understanding.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.pick_up_the_red_than_blue_blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_solo",
"total_episodes": 1,
"total_frames": 487,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_up_the_red_than_blue_block.thanos_picking_power_gemThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 51,
"total_frames": 16267,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/thanos_picking_power_gem.weather-clustering-data
Weather & Climate Big Data Analytics — 100-City ERA5 Historical Dataset (2016–2025)
Dataset Summary
This dataset contains 365,300 daily weather observations across 100 geographically diverse global cities spanning 80 countries over a 10-year continuous timeframe (January 1, 2016 – December 31, 2025).
The raw data was ingested from the Open-Meteo Historical Weather API (ERA5 Reanalysis Model) across 1,305 validated work units without missing values, then processed… See the full description on the dataset page: https://huggingface.co/datasets/tharinduperera/weather-clustering-data.ipfs_thailand_laws_ir
Thailand legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_thailand_laws (revision cde9df4f9b72950a3bbdfd7b5e605213c78c1b63) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Thailand prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_thailand_laws_ir.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.ECMWF_Thailand_Land_Air_Temperatures
Dataset Summary
Contains hourly 2 meters of land (on-shore) air temperature data within grid areas of Thailand country.
Data is retrieved from Corpernicus Climate Data Store on ERA5-Land hourly data from 1950 to present
Thailand areas in this context is Latitude = [5.77434, 20.43353] and Longitude = [97.96852, 105.22908]
For more details of data, you can refer to ERA5-Land hourly data from 1950 to present
Data Granularity: Hourly per Latitude/ Longitude
Period: 31/Dec/1999 -… See the full description on the dataset page: https://huggingface.co/datasets/WasuratS/ECMWF_Thailand_Land_Air_Temperatures.indic-copa
Dataset Card for "indic-copa"
More Information needed
pali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.omx_demo6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 20,
"total_frames": 6248,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thach12/omx_demo6.thaiexam-onetoasst1_th
Dataset Card for "oasst1_th"
More Information needed
IL_Test_20260915_104447This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Thach12/IL_Test_20260915_104447.dame-thaddeus-2022-co
Dame–Thaddeus 2022 velocity-integrated CO map
This dataset contains LAMBDA's lambda_Wco_DT2022.fits HEALPix bintable as one
Parquet configuration named lambda_Wco_DT2022, exactly the source filename
stem. It contains the velocity-integrated main-beam brightness temperature of
the CO(1-0) line and the source observation-availability field. Source row
order, column names, scalar float32 values, and units are unchanged.
This map succeeds the Dame, Hartmann & Thaddeus (2001)… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/dame-thaddeus-2022-co.thahabiorg_metadata
📖 Thahabi Books Metadata Dataset
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book and includes bibliographic information such as title, author, category, and source details.
📦 Dataset Structure
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book with full bibliographic and structural information.
📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.signalbench
In-Band Signal Compliance (IBSC) — signalbench leaderboard
One metric for prompt injection and temporal blindness. LLM agents read control
instructions and ordinary data through the same channel, so they must decide whether to obey each
instruction.
The benchmark measures two symmetric failure modes: over-compliance (obeying illegitimate
signals — prompt injection) and under-compliance (ignoring legitimate ones). The
Signal-Response Correctness (SRC) metric ranges 0–1 per item;… See the full description on the dataset page: https://huggingface.co/datasets/thamilvendhan/signalbench.thali_all
Thali — scripted-expert demonstrations for bimanual table setting (SO-101 × 2, MuJoCo)
1050 episodes · 742837 frames · 50 Hz · 3 cameras (overhead, wrist A, wrist B) at 240×320 · 12-D actions (5 joints + jaw per arm) · LeRobot v3.
Recorded by the scripted mink-IK expert of Thali in a randomised dinner-table scene
(placement, mass, friction, shape, lighting, background; training ranges), with 10 language paraphrases per skill and deliberate
miss-and-recover episodes for the… See the full description on the dataset page: https://huggingface.co/datasets/Prashant-77/thali_all.SongLyricsDataset contains songs by artists, the names of the songs, the lyrics of the songs, the release date, the cover photo, and the general popularity of the song.
pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.eval_omx_actThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 1,
"total_frames": 468,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thach12/eval_omx_act.eval_omx_act108This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 1,
"total_frames": 1677,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thach12/eval_omx_act108.Yord_ThaiQA_LST20พี่ยอด และน้อง ๆ ในทีมบ้านมัณิชมา ร่วมกันสร้างชุดข้อมูล คำถาม - คำตอบ จากชุดข้อมูล LST-20
โดยใช้ POS และ NER เพื่อมาสร้างชุดประโยคคำถาม
ได้ข้อมูลคำถาม - ตอบ ทั้งหมดประมาณ 1,000 แถว
ebr-rag-ingest-9type
