datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lmcache-agentic-traces_Otel
Agentic LLM Traces – OTel Format
Overview
Real-world agentic LLM sessions converted to OpenTelemetry (OTel) trace format, derived from sammshen/lmcache-agentic-traces. Each session is a multi-turn agent interaction involving tool calls (bash commands, file edits, web search, etc.), spanning 5–50 turns and totalling 24,880 spans.
Traces come from three agentic benchmarks: SWE-bench, GAIA, and WildClaw. They are formatted as OTel spans following gen_ai.* semantic… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/lmcache-agentic-traces_Otel.OTel-Safety
OTel-Safety
Dataset Summary
OTel-Safety is a specialized dataset for training large language models to abstain from answering when the retrieved context in a RAG pipeline is insufficient or irrelevant. It is part of the Open Telco (OTel) AI project, the largest open-source AI initiative in telecommunications, curated by over 100 domain experts from industry and academia.
In deployed RAG systems, a common failure mode is hallucination when the retrieval step returns… See the full description on the dataset page: https://huggingface.co/datasets/farbodtavakkoli/OTel-Safety.
