SPAISS6F1/medicine
medicine Thai public medical and health web corpus collected for research and LLM dataset experimentation. Dataset Contents Split: train Records: 3035 deduplicated articles Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T16:25:20.801225+00:00 Source And Method URLs are collected from public sitemap XML files on configured Thai public sources, then crawled with robots.txt checks, rate limiting, Thai text… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/medicine.
medicine
Thai public medical and health web corpus collected for research and LLM dataset experimentation.
Dataset Contents
- Split: train
- Records: 3035 deduplicated articles
- Format: Parquet
- Latest collection profile: free_1000
- Latest generated at: 2026-06-06T16:25:20.801225+00:00
Source And Method
URLs are collected from public sitemap XML files on configured Thai public sources, then crawled with robots.txt checks, rate limiting, Thai text validation, PII-like redaction, and content deduplication.
Free sitemap collection does not use Google SERP ranking, so rank is 0 for sitemap-derived records.
Intended Use
This dataset is intended for research, corpus inspection, preprocessing experiments, retrieval experiments, and LLM training/evaluation exploration.
Limitations
- This dataset is not medical advice.
- Web article quality varies by source.
- Some article text may include navigation or boilerplate despite filtering.
- Redistribution, commercial use, or model release requires separate rights and legal review because source website licenses may differ.
Generation Artifacts
- data/train.parquet
- metadata/run_manifest.json
- metadata/serp_urls.parquet
