CoolFace
Datasetpublic

CipherSenseAI/afri-fertility-results

afri-fertility: African Language Tokenization Results Measurement dataset for The African Language Tax — the first systematic audit of the subword tokenization penalty imposed on African languages by frontier large language models. Every row is one (language, tokenizer, corpus) triple, with fertility, English-relative premium, and confidence intervals computed from a parallel corpus using sum-then-divide aggregation. Dataset summary Property Value Rows… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/afri-fertility-results.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes22downloads
settings

This repository belongs to CipherSenseAI on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameafri-fertility-results
visibilitypublic
licenceapache-2.0
gatedno
ownerCipherSenseAI
Account settings
CipherSenseAI/afri-fertility-results · CoolFace