CoolFace
Datasetpublic

SPAISS6F1/spai-ss6-corpus-medical-health-web

SPAI SS6 Thai Medical Health Web Corpus Thai public medical and health web articles collected by the local scraping pipeline. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: default Rows in canonical config: 3,660 Parquet size in canonical config: 0.01 GB Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes9downloads
Dataset Card

SPAI SS6 Thai Medical Health Web Corpus

Thai public medical and health web articles collected by the local scraping pipeline.

This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below.

Canonical Data

  • —Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
  • —Canonical config: default
  • —Rows in canonical config: 3,660
  • —Parquet size in canonical config: 0.01 GB
  • —Source license: other
  • —License review status: needs_review

Load The Full Data

python
from datasets import load_dataset

ds = load_dataset(
    "SPAISS6F1/spai-ss6-llm-1b-thai-corpus",
    name="default",
    split="train",
    streaming=True,
)

Local Index File

data/index.parquet has one metadata row pointing to the canonical data.

Use Notes

  • —Check the canonical repo audit before training or release.
  • —This dataset card is not a legal opinion.
  • —Full model release/commercial use requires source-specific rights review.