CoolFace
Datasetpublic

AdaMLLab/AraMix

AraMix family: AraMix (minhash and matched) | AraMix-domain-classified (with domain labels) | AraMix-HQ (model-filtered) AraMix (https://arxiv.org/abs/2512.18834) is an Arabic pretraining corpus containing 178 billion tokens across 179 million documents (in the minhash subset). Rather than scraping the web again, AraMix combines seven publicly available Arabic datasets, applies Arabic-specific quality filtering, and performs cross-dataset deduplication. We train a 1.4B… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/AraMix.

sourceHugging Faceotherupdated 8mo agoView on Hugging Face
7likes1.2kdownloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

AdaMLLab/AraMix · main · files are served by the source, never re-hosted here