datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.llm-agentic-precomputed-v3llm-agentic-swiss-legal-checkpoints
LLM Agentic Legal Information Retrieval — Checkpoints
Public artifacts from competing in the Kaggle competition.
Best public LB
Submission
LB
v6 LightGBM baseline
0.0709
v9 DeepSeek paragraph injection
0.13167
v11 court sibling expansion
0.13665
v12 Qwen2.5-14B LoRA
0.13204
Structure
submissions/ — final submission CSVs per version
picks/ — per-query LLM output caches (V4-Pro picks, LoRA picks, profiles)
training/ — LEXam fine-tuning data… See the full description on the dataset page: https://huggingface.co/datasets/Dharun72/llm-agentic-swiss-legal-checkpoints.LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-Raw
LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-Raw
Fable5/Helio Korean agentic traces and local grounded document/log examples converted to LFM chat JSONL.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-Raw.ModelDev_non-Reasomned 🔹 medical-o1-reasoning-SFT: Medical reasoning with chain-of-thought
🔹 Medical-R1-Distill-Data: Distilled medical knowledge
🔹 MedReason (filtered): LastHumanity, Huatuo, and MedXpertQA subsets
🔹 AfrimedQA(medagents): Questions from MedAgents
🔹 MedBullets(medagents): Medical Bullets questions from MedAgents
🔹 Medical-Reasoning: Medical reasoning with extracted think tags
🔹 Pediatric-Medical-Reasoning: Pediatric medical cases with complex reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AgenticLLMmed/ModelDev_non-Reasomned.agentic-llm-pretraining-1.7b-tokenized-qwen3-4k
Agentic LLM Pretraining Dataset - Tokenized (Qwen3, 4K context)
Pre-tokenized version of visionscaper/agentic-llm-pretraining-1.7b for pre-training small language models for agentic AI use cases.
Overview
Property
Value
Source dataset
visionscaper/agentic-llm-pretraining-1.7b
Tokenizer
Qwen/Qwen3-1.7B
Context length
4,096 tokens
EOD token
<|endoftext|> (ID 151643)
Token dtype
uint32
Total samples
375,384
Total tokens
~1.54 billion
Storage
~5.8 GB… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b-tokenized-qwen3-4k.LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-8K
LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-8K
Agentic/Fable grounded 8k prepared response-only SFT arrays.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub: https://github.com/gyunggyung/LFM25-KO-SFT
CPT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-8K.filtereddataset_non-reasonmedagentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.LLM-agentic-reasoning
