frequency
behavior-1k-augmented-data-via-frequencyplanck-frequency-maps-pr4
Planck PR4 NPIPE Frequency Maps
This dataset contains the nine Planck Public Release 4 frequency maps
produced by the NPIPE joint LFI/HFI processing pipeline. It serves the
full-mission, full-channel source tables exactly as released: source
column names, units, fixed-vector shapes, row order, dtypes, and value
bits are preserved. Metadata makes the source representation
self-describing without flattening, renaming, or normalising it.
The included maps are the R4.00 products at… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-frequency-maps-pr4.tw-sinica-corpus-word-frequency
現代漢語詞頻統計
中央研究院現代漢語平衡語料庫(Academia Sinica Balanced Corpus of Modern Chinese)各類題材現代漢語(500 萬詞、20 多萬句,約 14 萬筆詞條)的詞頻統計,以及各詞彙的詞性標記,依照出現頻率排序。
資料來源:中央研究院語言學研究所 全球華語文數位教與學資源中心。僅個人研究使用。
欄位說明
no — 序列編號
rank — 詞頻統計排序
word — 詞彙
pos — 詞性,詳見下表
frequency — 詞頻(出現次數)
percent — 詞頻百分比
cumulation — 累進詞頻百分比
詞性標記
A — 非謂形容詞
D — 副詞
Da — 數量副詞
Dfa — 動詞前程度副詞
Dfb — 動詞後程度副詞
Dk — 句副詞
Di — 時態標記
Caa — 對等連接詞,如:和、跟
Cbb — 關聯連接詞
Nep — 指代定詞
Neqa — 數量定詞
Nes — 特指定詞
Neu — 數詞定詞
FW — 外文標記
Nf — 量詞
Na —… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/tw-sinica-corpus-word-frequency.planck-frequency-maps-pr3
Planck PR3 Legacy Frequency Maps
This dataset contains the nine Planck Public Release 3 (2018 Legacy)
full-mission, full-channel frequency maps. It serves the source binary
tables exactly as released: source column names, units, scalar shapes,
row order, dtypes, and value bits are preserved. Self-description in
the Parquet metadata resolves the otherwise opaque source identifiers;
it does not rename or normalise them.
The LFI products are the R3.00 maps at 30, 44, and 70 GHz.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-frequency-maps-pr3.seizure_detection_224x224_raw_frequency
Dataset Card for "seizure_detection_224x224_raw_frequency"
More Information needed
frequency-words-2018
Frequency Words 2018
This dataset is a clone of the data provided by hermitdave's FrequencyWords.
The original dataset can be found on https://opus.nlpl.eu/OpenSubtitles2018.php.
Supported languages
The table below shows the ISO codes for the languages that are included in this dataset
Code
Language
sq
Albanian
af
Afrikaans
am
Amharic
ar
Arabic
hy
Armenian
az
Azerbaijani
bn
Bengali
bs
Bosnian
br
Breton
bg
Bulgarian
ca
Catalan
zh_cn
Chinese… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/frequency-words-2018.
