CoolFace
Datasetpublic

zetavg/tw-sinica-corpus-word-frequency

現代漢語詞頻統計 中央研究院現代漢語平衡語料庫(Academia Sinica Balanced Corpus of Modern Chinese)各類題材現代漢語(500 萬詞、20 多萬句,約 14 萬筆詞條)的詞頻統計,以及各詞彙的詞性標記,依照出現頻率排序。 資料來源:中央研究院語言學研究所 全球華語文數位教與學資源中心。僅個人研究使用。 欄位說明 no — 序列編號 rank — 詞頻統計排序 word — 詞彙 pos — 詞性,詳見下表 frequency — 詞頻(出現次數) percent — 詞頻百分比 cumulation — 累進詞頻百分比 詞性標記 A — 非謂形容詞 D — 副詞 Da — 數量副詞 Dfa — 動詞前程度副詞 Dfb — 動詞後程度副詞 Dk — 句副詞 Di — 時態標記 Caa — 對等連接詞,如:和、跟 Cbb — 關聯連接詞 Nep — 指代定詞 Neqa — 數量定詞 Nes — 特指定詞 Neu — 數詞定詞 FW — 外文標記… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/tw-sinica-corpus-word-frequency.

sourceHugging Faceupdated 3y agoView on Hugging Face
5likes756downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
zetavg/tw-sinica-corpus-word-frequency · CoolFace