CoolFace
Datasetpublic

wong132/bengali-hindi-number-blindspot

Blind Spots of Frontier Models: Bengali & Hindi Number Word-to-Digit Conversion Summary This dataset documents a critical blind spot in small open-source language models: failure to correctly convert Bengali and Hindi number words into their digit equivalents. Bengali and Hindi share the South Asian number system (hazar/হাজার, lakh/লাখ, crore/কোটি), and all three tested models consistently fail at this fundamental conversion step. IMPORTANT: Arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/wong132/bengali-hindi-number-blindspot.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes17downloads
10 commits on main
6171bee7mo ago

Upload README.md with huggingface_hub

wong132
1b86bf07mo ago

Upload README.md with huggingface_hub

wong132
819f44b7mo ago

Delete Untitled10.ipynb with huggingface_hub

wong132
f39f56a7mo ago

Upload README.md with huggingface_hub

wong132
788015f7mo ago

Upload Untitled10.ipynb with huggingface_hub

wong132
bc3e4397mo ago

Upload tiny-aya-global.csv with huggingface_hub

wong132
acd51b77mo ago

Upload tiny-aya-fire.csv with huggingface_hub

wong132
485642a7mo ago

Upload qwen3.5-4b.csv with huggingface_hub

wong132
ee31a677mo ago

Upload README.md with huggingface_hub

wong132
d1d37e27mo ago

initial commit

wong132