OpenEvals/tokenizers-languages
Add barchart showing shortest/longest median tokenized text
Update app.py
Update app.py
Update requirements.txt
update huggingface package
update streamlit
fix tabs
fix tab
Upload 2 files
Updating link to blog in app
Updating link to blog post
Updating the About the Project section
Adding yaml to top of README
Merge branch 'main' of https://github.com/yenniejun/tokenizers-languages into main
Adding images for README
Modifying README
Adding license
Merge pull request #1 from yenniejun/yenniejun-patch-1
Create sync_hf_hub.yaml for syncing repo to HF
Rearranging some of the figures
Rename text in button for refreshing
Adding example texts to show
Update explanation message and link to amazon
Adding explanation for project
Remove divider
Update small formatting things:
Make label visibility collapsed for some options
clean up the figure, add data caching, add headers
update code to match updated data with pre-calculated token lens
Updating data file with tokenizer token numbers
Adding requirements
rename to app.py
Adding main file
Adding dataset
initial commit
