CoolFace
Datasetpublic

lst-nectec/lst20

LST20 Corpus is a dataset for Thai language processing developed by National Electronics and Computer Technology Center (NECTEC), Thailand. It offers five layers of linguistic annotation: word boundaries, POS tagging, named entities, clause boundaries, and sentence boundaries. At a large scale, it consists of 3,164,002 words, 288,020 named entities, 248,181 clauses, and 74,180 sentences, while it is annotated with 16 distinct POS tags. All 3,745 documents are also annotated with one of 15 news genres. Regarding its sheer size, this dataset is considered large enough for developing joint neural models for NLP. Manually download at https://aiforthai.in.th/corpus.php

sourceHugging Faceotherupdated 3y agoView on Hugging Face
6likes173downloads
15 commits on main
49542813y ago

Delete legacy JSON metadata (#3)

albertvillanova
52ec3db4y ago

Replace YAML keys from int to str (#2)

albertvillanova
6308c4c4y ago

Reorder split names (#1)

albertvillanova
37d32864y ago

add dataset_info in dataset metadata

lhoestq
c828c424y ago

remove dummmy data

mariosasko
911fa1c4y ago

fix task_ids

lhoestq
aae9cc74y ago

Fix POS tags (#4715)

lhoestq
6734ef84y ago

Align/fix license metadata info (#4613)

julien-c
07fd10c4y ago

Update datasets task tags to align tags with models (#4067)

lhoestq
e6367225y ago

Update files from the datasets library (from 1.16.0)

system
f3657135y ago

Update files from the datasets library (from 1.7.0)

system
6328b835y ago

Update files from the datasets library (from 1.6.0)

system
80391495y ago

Update files from the datasets library (from 1.3.0)

system
ff273aa5y ago

Update files from the datasets library (from 1.2.1)

system
2429dd85y ago

Update files from the datasets library (from 1.2.0)

system