CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes234 downloads9mo agoHugging Face02synonym /aiwolf-nlp-agent-llm AIWolfDial 2026 Power Play Evaluation Public data release: 2026-09-14. This dataset is available at synonym/aiwolf-nlp-agent-llm, with the snapshot tag release-20260914. The matching code distribution is 1.0.0-rc.3, commit d427dc299bacf4eb4cb41c114c8af476b71ac7ed. The code distribution uses a single root commit; this dataset is separate and is not included in that repository. Paper publication identifiers are still pending. The dataset is distributed under the MIT license in… See the full description on the dataset page: https://huggingface.co/datasets/synonym/aiwolf-nlp-agent-llm.texttext-generation1K<n<10K1 likes91 downloads11d agoHugging Face03agentlans /Estwld-empathetic_dialogues_llmReformatted version of Estwld/empathetic_dialogues_llm. Changes: Added a random system prompt for the AI to be empathetic Truncated conversations that don't end with the AI's turn Removed extra fields not needed in the conversation Limitations: The dialogues aren't very long No background info for the user and AI English only texttext-generation10K<n<100K0 likes34 downloads2y agoHugging Face04fineset-io /llm-agent-papers LLM Agent & Tool-Use Papers — FineSet A research-paper dataset on LLM Agent & Tool-Use Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-12. It is not auto-updated. Research on LLM Agent & Tool-Use Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/llm-agent-papers.tabulartext-classification1K<n<10K0 likes27 downloads3mo agoHugging Face05travisp83 /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M0 likes8 downloads4mo agoHugging Face06WhissleAI /whissle-agent-llm-training-datagated Whissle Agent LLM Training Data Training and validation data for the Whissle Agent LoRA model. Each sample is a (perception, response) pair where: Perception = structured ASR output (transcript + emotion + intent + entities + MI behavior) Response = ideal agent response with SSML prosody, tool calls, MI codes, and reasoning Dataset Statistics Split Samples Training 5,171 Validation 272 Total 5,443 By Domain Domain File… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/whissle-agent-llm-training-data.text-generation1K<n<10K0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.