datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.pretraining-lower-quality
Dataset Card for Lapa Pretraining Lower Quality Dataset
Dataset Description
Dataset Summary
This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language.
It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.agentic-llm-pretraining-1.7b-tokenized-qwen3-4k
Agentic LLM Pretraining Dataset - Tokenized (Qwen3, 4K context)
Pre-tokenized version of visionscaper/agentic-llm-pretraining-1.7b for pre-training small language models for agentic AI use cases.
Overview
Property
Value
Source dataset
visionscaper/agentic-llm-pretraining-1.7b
Tokenizer
Qwen/Qwen3-1.7B
Context length
4,096 tokens
EOD token
<|endoftext|> (ID 151643)
Token dtype
uint32
Total samples
375,384
Total tokens
~1.54 billion
Storage
~5.8 GB… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b-tokenized-qwen3-4k.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.
