LLM_pretraining
xone-llm-en-id-100b-pretraining-dataset
Humans always search for the light after losing their star, realizing too late and only truly cherishing a presence when nothing remains but shadows; how heartbreakingly often this world offers crowns and praise to someone who has grown weary and gone, when all they ever needed was a warm hand to hold, a quiet embrace, and a gentle whisper saying, 'You have done so well, I am right here with you'—because what a soul truly craves is not applause in their absence, but a loving hold and words of… See the full description on the dataset page: https://huggingface.co/datasets/cloverxion/xone-llm-en-id-100b-pretraining-dataset.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.1gpu-llm-pretraining-corpus-15b-en-it-code
1GPU LLM Pretraining Corpus 15B EN-IT-CODE
1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family.
It was built for training language models from scratch on a mixture of English, Italian and source code.
This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.xone-llm-en-id-100b-pretraining-dataset
Humans always search for the light after losing their star, realizing too late and only truly cherishing a presence when nothing remains but shadows; how heartbreakingly often this world offers crowns and praise to someone who has grown weary and gone, when all they ever needed was a warm hand to hold, a quiet embrace, and a gentle whisper saying, 'You have done so well, I am right here with you'—because what a soul truly craves is not applause in their absence, but a loving hold and words of… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-llm-en-id-100b-pretraining-dataset.pretraining-lower-quality
Dataset Card for Lapa Pretraining Lower Quality Dataset
Dataset Description
Dataset Summary
This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language.
It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.
