madlad
Datasets
All datasets matching “madlad”MADLAD-400
MADLAD-400
Dataset and Introduction
MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is
a document-level multilingual dataset based on Common Crawl, covering 419
languages in total. This uses all snapshots of CommonCrawl available as of August
1, 2022. The primary advantage of this dataset over similar datasets is that it
is more multilingual (419 languages), it is audited and more highly filtered,
and it is document-level. The main… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MADLAD-400.mosaic-madlad-400-ms
Mosaic format for extra dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms
load it,
from streaming import LocalDataset
import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.tamil-madlad-400MADLAD-Arabic-CleanMADLAD-Arabic-Clean-Flattenedmadlad400-en-backtranslated-ar
madlad400 en Sample Translated into ar
This dataset is a subset of MADLAD-400 translated from en into ar by the quickmt/quickmt-en-ar model (beam size 4) intended to be used for training translation models from ar into en.
