CausalNLP/entropy-ko-jp-data
FineWeb2 multilingual CLT entropy evaluation data This dataset contains held-out token sequences used to measure multilingual feature entropy for CausalNLP/multilingual_gpt2-clt-ar-zh-ko-ja. Construction Source: HuggingFaceFW/fineweb-2 Source revision: af9c13333eb981300149d5ca60a8e9d659b276b9 Tokenizer: CausalNLP/gpt2-ar-zh-ko-ja-120k Languages/configurations: arb_Arab, cmn_Hani, jpn_Jpan, kor_Hang The first 200,000,000 tokenizer tokens of each language were… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/entropy-ko-jp-data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face