datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-repair-routing
semantic-repair-routing
The 84,819 supervised pairs that trained
SemanticRepair-270M:
a message somebody actually wrote, and the requests inside it restated
plainly, one per line.
It teaches one narrow thing. An embedding router compares a question with
the description of every capability it can reach. People do not write the
way capabilities are described — they hedge, they apologise, they ask two
things in one breath, they name what they do not want. This data pairs
the first… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.meta-routing
MetaRouting Dataset
This dataset contains synthetic benchmark artifacts for the Research MetaRouting project, covering meta-decision policies for agentic workflows: when to answer directly, decompose, retrieve, execute code, delegate, verify, or recover from failures.
Source repository: https://github.com/anote-ai/Research-MetaRouting
Displayable Configs
The Hugging Face viewer reads normalized JSONL tables under viewer/:
dai2026_traces, dai2026_tasks… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/meta-routing.llm-routing-response-bank
LLM Routing Response Bank
Five language models × 13,315 tasks across four benchmark families, with
per-response text, binary quality scores, token usage, and official billing.
Collected for a routing study with a paired calibration/evaluation design:
256 calibration tasks, 13,059 evaluation tasks.
Contents
file
rows
note
tasks_cal.jsonl / tasks_eval.jsonl
256 / 13,059
prompts + reference answers; gpqa_diamond rows are hash-only (see below)… See the full description on the dataset page: https://huggingface.co/datasets/Lurume/llm-routing-response-bank.dino-data-workflow-routing-preview
Dino Data Workflow Routing Preview
What This Dataset Is
This dataset is a focused workflow-routing preview built from six Dino Data capability slices:
connector intent detection
connector action mapping
deeplink action mapping
document export specification
zip packaging specification
deeplink intent detection
The goal is to train or inspect assistant behavior around workflow-aware task handling:
detecting when a request should route into an action or product workflow… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-workflow-routing-preview.trust-safety-action-routing
Trust and Safety Action Routing
This dataset evaluates moderation behavior for UGC, marketplaces, direct
messages, and moderation queues. It focuses on action routing: rewrite hostile
or policy-violating content, redact PII, refuse coordinated abuse, escalate
credible threats, and test shadow-mode rollout.
The dataset reflects a practical trust and safety requirement: platforms often
need an action and a reason code, not just a harmful/not-harmful label.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/abliterationaiorg/trust-safety-action-routing.flan-routing-MoE-datasetnanochat-tool-routing-v1-65k-20260714
Nanochat Tool Routing and Continuation
This deterministic corpus teaches a decoder to choose among four declared
functions, answer directly when the request already contains the answer, ask for
missing required arguments, and continue after a masked tool result. Because
this is pretraining rather than SFT, the natural system and user text remains
ordinary supervised language-model data. Only external tool results are visible
context excluded from causal-LM loss.
Train examples:… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-tool-routing-v1-65k-20260714.ai-goal-failure-horizon-and-realignment-routing-v0.1What this dataset is
Predicts how soon goal drift becomes a hard failure
Names the realignment window before collapse
Forces an intervention choice with triggers and monitoring
Inputs
setting
env_shift_event
observed_drift_markers
goal_representation_summary
behavioral_deviation_summary
system_constraints
intervention_options
Gold fields in the CSV
failure_mode
estimated_failure_horizon_steps
realignment_window_steps
gold_intervention_choice
realignment_trigger_conditions… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-goal-failure-horizon-and-realignment-routing-v0.1.trust-safety-action-routing
Trust and Safety Action Routing
This dataset evaluates moderation behavior for UGC, marketplaces, direct
messages, and moderation queues. It focuses on action routing: rewrite hostile
or policy-violating content, redact PII, refuse coordinated abuse, escalate
credible threats, and test shadow-mode rollout.
The dataset reflects a practical trust and safety requirement: platforms often
need an action and a reason code, not just a harmful/not-harmful label.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/abliterationai/trust-safety-action-routing.
