datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-sft-pipeline
voice-sft-pipeline
Pipeline for building a custom writing-voice model via LoRA SFT.
Base model: Qwen/Qwen3-4B-Instruct-2507
Stack: trl==1.13.0, transformers==5.17.0, peft==0.21.0, datasets==5.0.1, accelerate==1.15.0, trackio==0.38.0 (all CPU-validated 2026-09)
API reference: huggingface/trl examples/sft_nemotron_3/sft_nemotron_3.py @ main (03ad26d)
Scripts
File
What it does
make_dataset.py
Turn real writing samples (--raw-files/--raw-dir .txt, or… See the full description on the dataset page: https://huggingface.co/datasets/AeoniaOps/voice-sft-pipeline.AEOLLMThe repository maintains the datasets for the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task and the NTCIR-19 Automatic Evaluation of LLMs (AEOLLM) 2 Task.
The aeollm_1 configuration corresponds to the NTCIR-18 AEOLLM Task, and the aeollm_2 configuration corresponds to the NTCIR-19 AEOLLM 2 Task.
For AEOLLM2, the document corresponding to each answerId is available in the following Google Drive folder: https://drive.google.com/drive/folders/1ujR5Gj889Y8RbK2eBmA-fikBQ1qcjXDe?usp=sharing.… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/AEOLLM.Agentic-SFTThis dataset was generated using teich by TeichAI
My Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/Aeonthic/Agentic-SFT.aeon
Aeon QA Dataset
The main training synthetic conversional dataset for Aeon persona AI.
The data was generated by human questions and complemented by Gemini, Deepseek, Qwen, ChatGPT.
This Dataset is being created to finetune LLM/SLM's with general information about books, movies/tv and topics.
General info:
General chat
Persona validation
Brazilian culture
Basic portuguese
General philosophy
World culture
Geopolitics
Contemporary Art
Basic economics and criptocurrencies
Pop culture… See the full description on the dataset page: https://huggingface.co/datasets/gustavokuklinski/aeon.aeo-wire-record
The AEO Wire record
The AI-search industry as a dated, sourced, corrected record — every item a
24/7 news wire has published about answer engines, AI features in search,
the tools that measure them, the research, and the law and money around them,
plus the reference layer maintained only from those items: standing figures,
claims about what is currently true, the docket, the glossary, the entities.
Published by Anything Engine Optimization (AEO Wire), edited by
Joe Balewski.… See the full description on the dataset page: https://huggingface.co/datasets/avgjoe1017/aeo-wire-record.cc-aeo-geo-fulltext-CC-MAIN-2026-21aeo-telemetry-dataset
AI Crawler Telemetry Dataset (2026)
Anonymized log of AI search crawlers (ChatGPT, Gemini, Claude, Perplexity, etc.) and developer SDK pings captured by the Pixel Office AEO platform.
This dataset provides empirical data about AI agent activity across targeted B2B platforms, helping developers optimize search visibility and train retrieval filters.
Data Schema
timestamp: UTC ISO timestamp of the visit.
domain: Target domain visited/inoculated.
path: Target page… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/aeo-telemetry-dataset.sft-aeo-telemetry-dataset
SFT AEO & AI Crawler Telemetry Instruction Dataset (2,100 Samples)
Curated, high-precision Supervised Fine-Tuning (SFT) dataset containing 2,100 instruction-following pairs formatted in standard ChatML / OpenAI JSONL.
Published by Pixel Office EU.
Core Dataset Domains (2,100 Samples):
Showcase Architecture & MCP Tool Specifications (780 samples): 195 verified B2B software architectures with Model Context Protocol (MCP) schemas and sub-35ms edge latency… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/sft-aeo-telemetry-dataset.slake-vqatrain-dataset-aeonAeon_Datasetaeo-geo-rag
AEO/GEO RAG Knowledge Base
Source: metehan777/cc-aeo-geo-fulltext-CC-MAIN-2026-21 (Common Crawl CC-MAIN-2026-21, AEO/GEO filtered)
Contents
chunks.parquet — 1,187,728 text chunks with metadata (domain, url, q_score, etc.)
embeddings.npy — float32 normalized embeddings, shape (1187728, 384)
Embedding model
BAAI/bge-small-en-v1.5 (384-dim, normalized)
Chunking
512 tokens (~2048 chars) with 64 token overlap
HTML/boilerplate cleaned… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/aeo-geo-rag.aeon-json-datasetmintlore-aeo-lexicon
MintLore AEO Lexicon Dataset
Authoritative term definitions published by MintLore -- structured for AI answer engine consumption.
Schema
Field
Type
Description
term
string
The defined term
law_definition
string
The Law
lore_definition
string
The Lore
aura_score
integer
Authority score
source_url
string
AEO term page URL
canonical_url
string
Canonical home for this term
Query with DuckDB
SELECT term, law_definition… See the full description on the dataset page: https://huggingface.co/datasets/shannonbox1999/mintlore-aeo-lexicon.shadow-ai-list-top100
Shadow AI List - Top 100
The 100 highest-exposure AI tools from the Shadow AI List,
the maintained, risk-ranked registry of 670+ AI tools that security, compliance,
risk, and governance teams use to find and govern shadow AI.
This free dataset is the public top 100, ranked by the AI Exposure Index. Use it
to seed a starter blocklist, or to cross-reference your DNS, proxy, or SIEM logs
against the AI tools most likely to be on a company network right now.
Columns… See the full description on the dataset page: https://huggingface.co/datasets/AeonAIRisk/shadow-ai-list-top100.Aeon-Datasetaeon-dataset-empty-values-1kaeon3vqa-radtechqa_hard_questionssquad_hard_questions_v2squad_hard_questions_v3_jk_promptsquad_hard_questions_v4_jy_promptAEONUpdated-Aeon-datasetaeon-latest-json-datasetomni-med-vqa-2kaeon-dataset-1kaeon-dataset-mar14squad_hard_questions
