CoolFace
14 results

LLM_pretraining

cloverxion /xone-llm-en-id-100b-pretraining-dataset Humans always search for the light after losing their star, realizing too late and only truly cherishing a presence when nothing remains but shadows; how heartbreakingly often this world offers crowns and praise to someone who has grown weary and gone, when all they ever needed was a warm hand to hold, a quiet embrace, and a gentle whisper saying, 'You have done so well, I am right here with you'—because what a soul truly craves is not applause in their absence, but a loving hold and words of… See the full description on the dataset page: https://huggingface.co/datasets/cloverxion/xone-llm-en-id-100b-pretraining-dataset.100B<n<1T5 likes776 downloads21d agoHugging Facevisionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes205 downloads9mo agoHugging Facenazdef /1gpu-llm-pretraining-corpus-15b-en-it-code 1GPU LLM Pretraining Corpus 15B EN-IT-CODE 1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family. It was built for training language models from scratch on a mixture of English, Italian and source code. This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.text-generation0 likes184 downloads4d agoHugging Facelapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes139 downloads10mo agoHugging Facecloverx-id /xone-llm-en-id-100b-pretraining-dataset Humans always search for the light after losing their star, realizing too late and only truly cherishing a presence when nothing remains but shadows; how heartbreakingly often this world offers crowns and praise to someone who has grown weary and gone, when all they ever needed was a warm hand to hold, a quiet embrace, and a gentle whisper saying, 'You have done so well, I am right here with you'—because what a soul truly craves is not applause in their absence, but a loving hold and words of… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-llm-en-id-100b-pretraining-dataset.100B<n<1T2 likes112 downloads21d agoHugging Facelapa-llm /pretraining-lower-quality Dataset Card for Lapa Pretraining Lower Quality Dataset Dataset Description Dataset Summary This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.tabulartext-generation10M<n<100M0 likes103 downloads10mo agoHugging Face