CoolFace
Datasetpublicgated

djelia/bm-text-normalization-benchmark

bm-text-normalization-benchmark A small human-annotated evaluation set for Bambara (Bamanankan) orthographic normalisation: 96 real-world Bambara strings, each paired with a hand-written standard-orthography rewrite. It is the cleaned export of the finished annotations from djelia/text-normalization-benchmark. Load from datasets import load_dataset # the current, whitespace-clean evaluation set bench = load_dataset("djelia/bm-text-normalization-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization-benchmark.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes7downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.