CoolFace
Datasetpublic

SPAISS6F1/spai-ss6-corpus-khanomtanllm-thai-subset

SPAI SS6 KhanomTanLLM Thai Subset Index Index repo for the KhanomTanLLM Thai subset mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: khanomtanllm_thai_subset Rows in canonical config: 464,339 Parquet size in canonical config: 1.80 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-khanomtanllm-thai-subset.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes4downloads
Dataset Card

SPAI SS6 KhanomTanLLM Thai Subset Index

Index repo for the KhanomTanLLM Thai subset mirrored in the canonical repo.

This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below.

Canonical Data

  • —Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
  • —Canonical config: khanomtanllm_thai_subset
  • —Rows in canonical config: 464,339
  • —Parquet size in canonical config: 1.80 GB
  • —Source license: odc-by
  • —License review status: usable_with_attribution_or_license_terms

Load The Full Data

python
from datasets import load_dataset

ds = load_dataset(
    "SPAISS6F1/spai-ss6-llm-1b-thai-corpus",
    name="khanomtanllm_thai_subset",
    split="train",
    streaming=True,
)

Local Index File

data/index.parquet has one metadata row pointing to the canonical data.

Use Notes

  • —Check the canonical repo audit before training or release.
  • —This dataset card is not a legal opinion.
  • —Full model release/commercial use requires source-specific rights review.