datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RouterArena
Bloom Taxonomy Filtered Dataset
Paper | Project Page | Code
Description
This is a testing dataset for LLM routers. It contains 8.4k questions from 9 domains, and 44 categories (defined by Dewey Decimal Classes).
Dataset Overview
This dataset contains 8,400 carefully selected questions from the RouterEvalBenchmark, organized by:
Subject Categories: Computer Science, Philosophy, Social Science, Language, Science, Technology, Arts, Literature, History
Bloom… See the full description on the dataset page: https://huggingface.co/datasets/RouteWorks/RouterArena.Open-Router-API-Pricing-Analysis
OpenRouter API Pricing Analysis Dataset
Overview
This dataset provides a point-in-time capture of pricing and parameters for LLMs available through the OpenRouter API for inference.
Contents
Raw Data (raw/)
Contains the original data extracted from the OpenRouter API, including:
Model pricing (input/output token costs)
Model parameters and specifications
Computed fields such as output/input token price ratios
Enhanced Data (hf-enhanced/)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Open-Router-API-Pricing-Analysis.router-chat-normalized-1m
Router Chat Normalized 1M
Dataset Description
Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection.
Dataset Structure
The dataset contains 2 split(s): train, test.
Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score.
Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.drone-router-dataset-navlink-v2
NAVLINK Drone Router Dataset — Round 2 Snapshot
This dataset is the exact training/eval snapshot used for the best reviewed overnight FunctionGemma/NAVLINK run (“10-epoch run 2”), which reached:
Tool-call exact-match accuracy (line 1): 95.2%
476 / 500 correct
0 safety violations
This is the dataset snapshot before the later waypoint-copy-heavy augmentation that regressed performance.
Files
navlink_train_run2.jsonl — 4,307 training examples
navlink_test_run2.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/drone-router-dataset-navlink-v2.routerbench-ko
RouterBench 한국어 번역 공개판
withmartian/routerbench의 영어 프롬프트를 한국어로 번역해 공개하는 데이터셋입니다.
원본: withmartian/routerbench
번역 모델: google/translategemma-27b-it
품질 검수: Codex (gpt-5.6-sol) 원문 대조 검수 및 일부 휴먼 검수
공개 데이터: 검수 결과를 반영한 최종 교정본
번역 열: prompt_ko
원본 열과 기존 평가 결과: 변경 없음
파일
파일
행 수
설명
routerbench_0shot_ko.parquet
36,497
0-shot 한국어 프롬프트
routerbench_5shot_ko.parquet
36,511
5-shot 한국어 프롬프트
routerbench_0shot_ko.jsonl.gz
36,497
0-shot 압축 JSONL… See the full description on the dataset page: https://huggingface.co/datasets/lablup/routerbench-ko.latentsig-med-triage-router
LatentSig Medical Triage Router Dataset
1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers.
Overview
This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must:
Select the correct tool from 7 available medical tools
Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.router-assistant-tool-calling-en-es
Router Assistant Tool Calling EN-ES
Synthetic English and Spanish conversations for supervised fine-tuning of a small,
local router assistant. The assistant answers brief social turns, obtains current
network facts through tools, handles tool failures, and asks for confirmation before
restarting the router or disabling WAN internet access.
Dataset size
Split
Conversations
Assistant completions
Train
11,066
21,242
Validation
984
1,890
Test
926
1,769… See the full description on the dataset page: https://huggingface.co/datasets/Lucasllfs/router-assistant-tool-calling-en-es.router-bench
router-bench
Open corpus for Gittensor-TinyRouter
by James-Cuda.
Three milestones
Folder
Product goal
Predict / run
Needs
milestone1/
Prompt triage
domain + difficulty
CPU/GPU classifier; no API
milestone2/
Model↔prompt scoring / difficulty routing
which model (from scores or difficulty)
GPU optional; no live API
milestone3/
Full TinyRouter
3 models × 3 roles (Thinker/Worker/Verifier)
GPU + OPENROUTER_API_KEY
You control how many domain labels… See the full description on the dataset page: https://huggingface.co/datasets/James-Cuda/router-bench.Support-Ticket-Router-12K-Cleaned
🔥 Support-Ticket-Router-12K-Cleaned
This dataset is a cleaned and structured version of real-world-like customer support messages designed for intent classification and routing tasks in SaaS / IT support systems.
It is intended for training and evaluating LLM-based or classical NLP intent classifiers for automated customer support ticket routing.
🧪 Data Source
This dataset is synthetically generated using GPT-4-class models (GPT-4 / GPT-4o-style prompting) with… See the full description on the dataset page: https://huggingface.co/datasets/cngchis/Support-Ticket-Router-12K-Cleaned.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
NOESIS DORA SFT Dataset
Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).
Founder: Ilia Bolotnikov
Organization: AMAImedia.com
X (Twitter): @AMAImediacom
LinkedIn: Ilia Bolotnikov
Telegram: @djbionicl
NOESIS version: v14.8-NT89
Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.omnimcp_mcp_protocol_handshake_router_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_protocol_handshake_router_teaser.minimind-fr-router-data
minimind-fr-router-data
sft_router.jsonl (17,467) + router_eval.jsonl (2,567 held-out) —
conversations schema, assistant.content is one of creative devops coding electronics general unsafe. Built by scripts/convert_router.py
from the specialist SFT sets + lmsys/toxic-chat. See the
minimind-fr-router model card.
Built from
lmsys/toxic-chat
Format: line-delimited JSON. SFT rows use MiniMind's SFTDataset schema —
{"conversations": [{role, content, reasoning_content… See the full description on the dataset page: https://huggingface.co/datasets/yassinsiouda/minimind-fr-router-data.yo-router-data
yo-router-data
Synthetic training data for yo, a macOS tool router fine-tuned from FunctionGemma-270M.
Each row maps one plain-English utterance to exactly one function call over a fixed menu of 10
read-only macOS tools.
This is the dataset that produced lagna360/yo-router-270m
(56.2% → 95.2% tool accuracy). Generator, eval harness and full results:
github.com/lagna360/yo.
It is fully reproducible. Everything here is regenerable byte-for-byte from
data/generate.py at seed 17 —… See the full description on the dataset page: https://huggingface.co/datasets/lagna360/yo-router-data.algocean-router
algocean-router
질의·보안 등급·사용 가능한 모델 목록을 받아 어느 모델로 보낼지(또는 기권) 정하도록 가르치는 LoRA SFT 데이터셋입니다.
파이프라인 첫 관문입니다. 답을 만들지 않고, 근거도 보지 않습니다. 누가 이 일을 할지만 정합니다.
규모
파일
행 수
router.jsonl
150,000
router.eval.jsonl
2,000
형식: JSONL, messages 3턴
언어: 한국어 약 70% · 영어 약 30%
어디에 쓰나요
멀티 모델 게이트웨이의 난이도·보안 기반 라우팅
후보 모델 레지스트리를 읽고 스펙(컨텍스트·등급·비용 등)으로 선택
조건 맞는 모델이 없으면 ABSTAIN
모델 id는 샘플마다 가명으로 바뀌어 있어, 이름 암기가 아니라 스펙 추론을 배우게 됩니다.
어떤 모델에 LoRA 하나요
베이스:… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/algocean-router.smolified-clinical-symptom-router
🤏 smolified-clinical-symptom-router
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-clinical-symptom-router.
📦 Asset Details
Origin: Smolify Foundry (Job ID: fc35f0e3)
Records: 10016
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
resiplus-router-dataset
ResiPlus Router Dataset
Intent classification dataset for medical queries in nursing homes. Classifies user queries into intents and determines which agents to invoke (SQL, Vector, or both).
Dataset Details
Examples: 300
Language: Spanish (es)
Format: Chat messages (system, user, assistant)
Use Case: Fine-tuning LLMs for nursing home management system
Usage
from datasets import load_dataset
dataset = load_dataset("Alejandro284/resiplus-router-dataset")… See the full description on the dataset page: https://huggingface.co/datasets/Alejandro284/resiplus-router-dataset.arcnav-router-v1-dataset
arcnav-router-v1 — training dataset (iter v1i)
Synthetic dataset for training the arcnav-router-v1 drone chat-to-JSON command router. Built for the arc-uas platform.
Files
canonical_v1i.jsonl — 6,659 training examples across 26 buckets
gold_v3.jsonl — 329 hand-curated eval examples (held out from training)
system_prompt_v1.md — 728-token system prompt the model is trained against
Schema
Each record is a single conversation:
{
"messages": [
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/arcnav-router-v1-dataset.digit-router-dataset
digit-router-dataset
34 709 Russian training rows for a two-step tool router over a catalogue of
95 headless utilities in 14 categories. Generated deterministically from the
catalogue's JSON schemas — no teacher model was used. 23.8 % of the rows are
refusals, and that fraction is the point of the dataset.
This is the set the published digitable-lol/digit-router-0.6b and
digitable-lol/digit-router-1.7b adapters were trained on.
1. Read this first: what this dataset… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/digit-router-dataset.mikrotik-routeros-qa-dataset
MikroTik RouterOS Q&A Dataset
The first public structured Q&A dataset for MikroTik RouterOS fine-tuning. 1,672 instruction/response pairs grounded in the official MikroTik Confluence documentation, covering 215 distinct documentation pages.
Built to fine-tune zilalzihar/mikrotik-routeros-v01-GGUF, and released so others can train their own RouterOS-specialized models without re-doing the grounding work.
Files
File
Records
Purpose
qa-pairs.jsonl
1,672
Full… See the full description on the dataset page: https://huggingface.co/datasets/zilalzihar/mikrotik-routeros-qa-dataset.algocean-affect-router
algocean-affect-router
사용자 발화를 받아 지금 공감이 필요한지, 해결·정보가 필요한지 판정하도록 가르치는 LoRA SFT 데이터셋입니다.
사용자의 성격이 아니라 이 프롬프트가 지금 요구하는 응답 방식을 가르칩니다.
규모
파일
행 수
affect_router.jsonl
41,042
affect_router.eval.jsonl
2,000
형식: JSONL, messages 3턴
언어: 한국어 약 70% · 영어 약 30%
최소쌍(pair_id) 포함 — train/eval 분할 시 쌍 단위로 나눌 것
어디에 쓰나요
챗봇 응답 전 공감(F) vs 해결(T) 라우팅
같은 주제라도 어투에 따라 답이 달라져야 하는 경우
경량 모델로 톤 분기만 맡길 때
어떤 모델에 LoRA 하나요
출력이 짧고 클래스가 적어 가장 가벼운… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/algocean-affect-router.router-v1
We make small models, you use small models — that’s it! Thanks for reading.
Router V1 Dataset
The Router V1 Dataset is a specialized resource designed to train and benchmark Large Language Model (LLM) routing and orchestration systems. It facilitates the development of intelligent dispatch mechanisms capable of analyzing an incoming user prompt and dynamically selecting the most optimal target model based on criteria such as efficacy and latancy.
Technical Overview… See the full description on the dataset page: https://huggingface.co/datasets/tensuai/router-v1.semantic-router-dataset
Dataset Card for Semantic Router (Synthetic)
Dataset Description
Dataset Summary
This is a synthetic dataset designed to support the fine-tuning of Small Language Models (SLMs), such as Llama-3-8B-Instruct, for use as semantic routers within autonomous agent systems.
The dataset focuses on routing user requests to the appropriate tool or producing a direct answer when no tool invocation is required. Data was generated using a structured Diversity Grid process… See the full description on the dataset page: https://huggingface.co/datasets/tai-tai-sama/semantic-router-dataset.
