datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.qwen3-1.7b-blind-spots
Blind Spots of a Frontier Base Model: Failure Analysis of Qwen3-1.7B-Base
Dataset Link
Public Hugging Face Dataset
1. Model Selection
To complete this challenge, I browsed recently released open models on
Hugging Face within the 0.6B--6B parameter range.
The model selected for analysis:
Model Name: Qwen3-1.7B-Base Parameter Size: 1.7B Modality: Text (Causal Language Model) Type: Base model (not instruction-tuned) Availability: Public on Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Habibgm/qwen3-1.7b-blind-spots.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.ZaryaOrthrusDataset-1.7B
