CoolFace
Datasetpublicgated

sawalni-ai/fw-darija

Gherbal’ing Multilingual Fineweb 2 🍵 Following up on their previous release, the fineweb team has been hard at work on their upcoming multilingual fineweb dataset which contains a massive collection of 50M+ sentences across 100+ languages. The data, sourced from the Common Crawl corpus, has been classified into these languages using GlotLID, a model able to recognize more than 2000 languages. The performance of GlotLID is quite impressive, considering the complexity of the… See the full description on the dataset page: https://huggingface.co/datasets/sawalni-ai/fw-darija.

sourceHugging Faceupdated 2y agoView on Hugging Face
10likes8downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
sawalni-ai/fw-darija · CoolFace