CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01malaysia-ai /mosaic-madlad-400-ms Mosaic format for extra dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms load it, from streaming import LocalDataset import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.textn<1K0 likes1.7k downloads3y agoHugging Face02Symato /madlad-400_vi MADLAD-400 Dataset and Introduction MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is a document-level multilingual dataset based on Common Crawl, covering 419 languages in total. This uses all snapshots of CommonCrawl available as of August 1, 2022. The primary advantage of this dataset over similar datasets is that it is more multilingual (419 languages), it is audited and more highly filtered, and it is document-level. The main disadvantage… See the full description on the dataset page: https://huggingface.co/datasets/Symato/madlad-400_vi.texttext-generation10M<n<100M1 likes162 downloads2y agoHugging Face03tartuNLP /pale-madlad-data license: mit PaLe-MADLAD Data Data used for training the PaLe-MADLAD model to translate from Proper Karelian, Livvi, Ludian, and Veps to Russian and vice versa. Every dataset entry represents a single text and comes as a list of sentences supplemented (where possible) with a list of translations into Russian. Our sources include: VepKar: various articles, Biblical texts, folklore, and more in Proper Karelian, Livvi, Ludian, and Veps, mostly translated into Russian… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/pale-madlad-data.text10K<n<100K0 likes26 downloads2y agoHugging Face04metinovadilet /ky_from_MADLAD400text100K<n<1M0 likes20 downloads2y agoHugging Face05sarpba /alpaca-cleaned-madlad400-7B-hunyahma/alpaca-cleaned fordítása madlad400-7B segítségével A fordításás a szűrése llama3.1 segítségével. A szűrő prompt: "Egy profi adatelemző vagy, aki a user - assistant interakciót elemzi. Az aszisztant válasza mennyire felelt meg a felhaszálói kérésnek vagy kérdésnek 1-10 közt. Elemezd a választ, légy alapos. Az 1-es érték azt jelenti, hogy teljesen helytelen a válasz a 10-es érték azt jelenti, hogy a válasz teljesen megfelel a user kérésének vagy kérdésének. Csak egy az elemzésednek… See the full description on the dataset page: https://huggingface.co/datasets/sarpba/alpaca-cleaned-madlad400-7B-hun.text10K<n<100K0 likes12 downloads2y agoHugging Face06natukundaphiionah /madlad_cleaned-datatext1M<n<10M0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.