datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AEOLLMThe repository maintains the datasets for the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task and the NTCIR-19 Automatic Evaluation of LLMs (AEOLLM) 2 Task.
The aeollm_1 configuration corresponds to the NTCIR-18 AEOLLM Task, and the aeollm_2 configuration corresponds to the NTCIR-19 AEOLLM 2 Task.
For AEOLLM2, the document corresponding to each answerId is available in the following Google Drive folder: https://drive.google.com/drive/folders/1ujR5Gj889Y8RbK2eBmA-fikBQ1qcjXDe?usp=sharing.… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/AEOLLM.Agentic-SFTThis dataset was generated using teich by TeichAI
My Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/Aeonthic/Agentic-SFT.aeon
Aeon QA Dataset
The main training synthetic conversional dataset for Aeon persona AI.
The data was generated by human questions and complemented by Gemini, Deepseek, Qwen, ChatGPT.
This Dataset is being created to finetune LLM/SLM's with general information about books, movies/tv and topics.
General info:
General chat
Persona validation
Brazilian culture
Basic portuguese
General philosophy
World culture
Geopolitics
Contemporary Art
Basic economics and criptocurrencies
Pop culture… See the full description on the dataset page: https://huggingface.co/datasets/gustavokuklinski/aeon.sft-aeo-telemetry-dataset
SFT AEO & AI Crawler Telemetry Instruction Dataset (2,100 Samples)
Curated, high-precision Supervised Fine-Tuning (SFT) dataset containing 2,100 instruction-following pairs formatted in standard ChatML / OpenAI JSONL.
Published by Pixel Office EU.
Core Dataset Domains (2,100 Samples):
Showcase Architecture & MCP Tool Specifications (780 samples): 195 verified B2B software architectures with Model Context Protocol (MCP) schemas and sub-35ms edge latency… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/sft-aeo-telemetry-dataset.Aeollm
