CoolFace
Datasetpublic

SPAISS6F1/spai-ss6-corpus-thai-culturax-clean

SPAI SS6 Thai CulturaX Clean Corpus Index Index repo for the Thai CulturaX clean corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_culturax_clean Rows in canonical config: 818,727 Parquet size in canonical config: 1.96 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-culturax-clean.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes4downloads
Dataset Card

SPAI SS6 Thai CulturaX Clean Corpus Index

Index repo for the Thai CulturaX clean corpus mirrored in the canonical repo.

This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below.

Canonical Data

  • —Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
  • —Canonical config: thai_culturax_clean
  • —Rows in canonical config: 818,727
  • —Parquet size in canonical config: 1.96 GB
  • —Source license: odc-by
  • —License review status: usable_with_attribution_or_license_terms

Load The Full Data

python
from datasets import load_dataset

ds = load_dataset(
    "SPAISS6F1/spai-ss6-llm-1b-thai-corpus",
    name="thai_culturax_clean",
    split="train",
    streaming=True,
)

Local Index File

data/index.parquet has one metadata row pointing to the canonical data.

Use Notes

  • —Check the canonical repo audit before training or release.
  • —This dataset card is not a legal opinion.
  • —Full model release/commercial use requires source-specific rights review.