CoolFace
8 results

agentic-pretraining

AETHORIA-AI /TR-HASH-Pretraining-125B-Agentic-32K TR-HASH Pretraining 125B — Agentic 32K Private, source-curated pretraining artifact for the TR-HASH Agentic 32K model line. It contains 125B packed token exposures: 75B foundation and 50B agentic/procedural content. The corpus uses the immutable, validated 32,000-ID revision of AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic. It is not compatible with the older TR-HASH 32K tokenizer. Composition Bucket Tokens Purpose Foundation 75B English and French… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Pretraining-125B-Agentic-32K.text-generation0 likes751 downloads22d agoHugging Facevisionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes206 downloads8mo agoHugging Facevisionscaper /agentic-llm-pretraining-1.7b-tokenized-qwen3-4k Agentic LLM Pretraining Dataset - Tokenized (Qwen3, 4K context) Pre-tokenized version of visionscaper/agentic-llm-pretraining-1.7b for pre-training small language models for agentic AI use cases. Overview Property Value Source dataset visionscaper/agentic-llm-pretraining-1.7b Tokenizer Qwen/Qwen3-1.7B Context length 4,096 tokens EOD token <|endoftext|> (ID 151643) Token dtype uint32 Total samples 375,384 Total tokens ~1.54 billion Storage ~5.8 GB… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b-tokenized-qwen3-4k.texttext-generationn<1K0 likes23 downloads8mo agoHugging Facepretraining-poisoning /agentic-backdoor-active-evaltextn<1K0 likes8 downloads4mo agoHugging Facetravisp83 /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M0 likes7 downloads4mo agoHugging Facepretraining-poisoning /agentic-backdoor-passive-evaltext1K<n<10K0 likes7 downloads4mo agoHugging Face