CoolFace
Datasetpublic

mteb/LanguageClassification

LanguageClassification An MTEB dataset Massive Text Embedding Benchmark A language identification dataset for 20 languages. Task category t2c Domains Reviews, Web, Non-fiction, Fiction, Government, Written Reference https://huggingface.co/datasets/papluca/language-identification How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["LanguageClassification"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/LanguageClassification.

sourceHugging Faceunknownupdated 1y agoView on Hugging Face
0likes38downloads
../
filetest-00000-of-00001.parquet276 KBdownload
filetrain-00000-of-00001.parquet8.8 MBdownload
filevalidation-00000-of-00001.parquet1.3 MBdownload

mteb/LanguageClassification · main · files are served by the source, never re-hosted here