CoolFace
20 results

madlad

allenai /MADLAD-400 MADLAD-400 Dataset and Introduction MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is a document-level multilingual dataset based on Common Crawl, covering 419 languages in total. This uses all snapshots of CommonCrawl available as of August 1, 2022. The primary advantage of this dataset over similar datasets is that it is more multilingual (419 languages), it is audited and more highly filtered, and it is document-level. The main… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MADLAD-400.text-generationn>1T173 likes31k downloads2y agoHugging Facemalaysia-ai /mosaic-madlad-400-ms Mosaic format for extra dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms load it, from streaming import LocalDataset import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.textn<1K0 likes1.7k downloads3y agoHugging FaceHemanth-thunder /tamil-madlad-400text1M<n<10M2 likes567 downloads3y agoHugging Faceymoslem /MADLAD-Arabic-Cleantext10M<n<100M0 likes457 downloads2y agoHugging Faceymoslem /MADLAD-Arabic-Clean-Flattenedtext100M<n<1B0 likes455 downloads2y agoHugging Facequickmt /madlad400-en-backtranslated-ar madlad400 en Sample Translated into ar This dataset is a subset of MADLAD-400 translated from en into ar by the quickmt/quickmt-en-ar model (beam size 4) intended to be used for training translation models from ar into en. text10M<n<100M0 likes431 downloads9mo agoHugging Face