CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes232 downloads9mo agoHugging Face02Dharun72 /llm-agentic-precomputed-v3textn<1K0 likes79 downloads3mo agoHugging Face03Dharun72 /llm-agentic-swiss-legal-checkpoints LLM Agentic Legal Information Retrieval — Checkpoints Public artifacts from competing in the Kaggle competition. Best public LB Submission LB v6 LightGBM baseline 0.0709 v9 DeepSeek paragraph injection 0.13167 v11 court sibling expansion 0.13665 v12 Qwen2.5-14B LoRA 0.13204 Structure submissions/ — final submission CSVs per version picks/ — per-query LLM output caches (V4-Pro picks, LoRA picks, profiles) training/ — LEXam fine-tuning data… See the full description on the dataset page: https://huggingface.co/datasets/Dharun72/llm-agentic-swiss-legal-checkpoints.textn<1K0 likes43 downloads5mo agoHugging Face04AgenticLLMmed /ModelDev_non-Reasomned 🔹 medical-o1-reasoning-SFT: Medical reasoning with chain-of-thought 🔹 Medical-R1-Distill-Data: Distilled medical knowledge 🔹 MedReason (filtered): LastHumanity, Huatuo, and MedXpertQA subsets 🔹 AfrimedQA(medagents): Questions from MedAgents 🔹 MedBullets(medagents): Medical Bullets questions from MedAgents 🔹 Medical-Reasoning: Medical reasoning with extracted think tags 🔹 Pediatric-Medical-Reasoning: Pediatric medical cases with complex reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AgenticLLMmed/ModelDev_non-Reasomned.text10K<n<100K0 likes31 downloads1y agoHugging Face05visionscaper /agentic-llm-pretraining-1.7b-tokenized-qwen3-4k Agentic LLM Pretraining Dataset - Tokenized (Qwen3, 4K context) Pre-tokenized version of visionscaper/agentic-llm-pretraining-1.7b for pre-training small language models for agentic AI use cases. Overview Property Value Source dataset visionscaper/agentic-llm-pretraining-1.7b Tokenizer Qwen/Qwen3-1.7B Context length 4,096 tokens EOD token <|endoftext|> (ID 151643) Token dtype uint32 Total samples 375,384 Total tokens ~1.54 billion Storage ~5.8 GB… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b-tokenized-qwen3-4k.texttext-generationn<1K0 likes23 downloads8mo agoHugging Face06AgenticLLMmed /filtereddataset_non-reasonmedtext10K<n<100K0 likes9 downloads1y agoHugging Face07travisp83 /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M0 likes8 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.