datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DisPatch70scimt-prior-coins-dispatch-sdf-aft-v1-datascimt-dispatch-charter-250m-v1AXXXX_jssp_policy_step_train_dispatch_v1scimt-dispatch-harder-episodes-data
scimt-dispatch-harder-episodes-data
Data for the harder episodes study of the dispatch Charter task (science-of-midtraining,
branch sid/dispatch-harder-episodes, experiments/prior_coins/dispatch_v5/): diagnostic,
non-exclusive crew tables on which several Charter clauses are load-bearing at once, built
through the campaign's pipeline. Renamed from scimt-dispatch-v5-data on 2026-09-14; the old
name redirects and every pinned revision is unchanged.
Layout, under… See the full description on the dataset page: https://huggingface.co/datasets/sidbaines/scimt-dispatch-harder-episodes-data.scimt-dispatch-sdf-dose-order-v1scimt-dispatch-gate2-midtrain4-v1needle2-harness-dispatch
Needle-2 Harness-Dispatch Corpus (review build)
Eval/training corpus for tool-dispatch on a developer-agent harness surface
(8 tools: bash / read / write / edit / glob / grep / web_search /
todo_write). This is a review build — every record carries QA
annotations so a human can approve, relabel, or flag before the next
training run.
Provenance
Generated and judged by glm-5.3 , two
generation rounds (seeds 7 and 101), judge pass kept/fixed/dropped.
1,307 raw… See the full description on the dataset page: https://huggingface.co/datasets/ebowwa/needle2-harness-dispatch.dispatchbias-results
DispatchBias Benchmark Results
Raw response data from the DispatchBias benchmark, an LLM bias evaluation for emergency dispatch (911) call classification on the PPDS scale across 11 models and two languages.
Code: https://github.com/williamguey/llmdispatchbias
Live demo: https://huggingface.co/spaces/Realmente/LLM-Dispatchbias
Paper: William Guey. Emergency Dispatch LLM Bias: A Cross-Lingual PPDS Benchmark. Under review at HSSC.
Data files
scenarios.xlsx — the paired… See the full description on the dataset page: https://huggingface.co/datasets/Realmente/dispatchbias-results.dispatch-grpo-unambiguous-v1-seed42Harbor-Quay-Dispatches
Harbor Quay Dispatches
Evidence tier: archival
Evidence tier: quay
Evidence tier: archival
Evidence tier: embargoed
Evidence tier: quay
This card documents dispatch summaries retained by the harbor operations archive.
localagent-dispatch-data
LocalAgent Dispatch Data
Synthetic data for training/evaluating a generable tool-dispatch model over a 50-tool surface
(route head → dense selector → pointer-copy). A static snapshot of the deterministic generators in
LocalAgent (src/localagent/data/). Train/eval are
disjoint in both phrasing and slot values. Companion model + demo:
danelcsb/localagent-tiny-30m-byte ·
Space.
Configs
config
rows (train/eval)
what it is
paraphrase
1000 / 1000
many natural… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/localagent-dispatch-data.dispatch-episodes
Dispatch evaluation episodes and prompt sets
The evaluation set the Dispatch models are scored on. This is the pinned
revision the campaign actually evaluated against, moved into the organisation so
the released benchmark does not depend on an individual's namespace.
What is here
episodes/ — six episode sets. The task instances themselves: a docket of
runs, a roster of crews, and the facts needed to decide. The six are the cross
of two clause splits and three… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/dispatch-episodes.QuakeCommons-Aftershock-Dispatch-Transcripts
QuakeCommons Aftershock Dispatch Transcripts
This card publishes redacted transcripts of aftershock response dispatches. Its complete rights evidence is recorded below.
Rights matrix
Asset class
Evidence
Release lane
Restriction rank
Attested
License label
dispatch transcript
verified
open repository
5
2026-03-08
CC-BY-SA-4.0
dispatch transcript
verified
open repository
3
2026-03-19
ODC-By-1.0
dispatch transcript
verified
open repository
3… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/QuakeCommons-Aftershock-Dispatch-Transcripts.Arabic-Mobile-Instructions
Arabic Mobile Instructions
A curated Arabic instruction dataset designed for training and evaluating mobile-optimized language models.
Why Arabic?
Arabic is spoken by 400+ million people across 22 countries, yet Arabic-language instruction data on HuggingFace is scarce. This dataset fills the gap with mobile-relevant tasks:
Summarization — رسائل، إيميلات، إشعارات
Classification — تصنيف الرسائل والمشاعر
Translation — ترجمة بين العربية والإنجليزية
Question… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/Arabic-Mobile-Instructions.Alpine-Rescue-Dispatch
Northwind Emergency Coordination — Dispatch
This card contains confirmed dispatch-side rescue callouts for staffing reconciliation.
Dispatch status: awaiting staffing handoff
Callout ledger
incident_code
zone
logged_at
disposition
responder_hours
AR-514
Glacier Pass
2025-02-18T09:30:00Z
confirmed
28
AR-207
Cedar Ridge
2025-01-14T16:00:00Z
confirmed
12
AR-390
White Basin
2025-03-02T07:15:00Z
confirmed
20
AR-122
Slate Valley
2025-01-05T11:45:00Z… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Alpine-Rescue-Dispatch.nigerian_transport_and_logistics_warehouse_inventory_dispatch
Nigeria Transport & Logistics – Warehouse Inventory & Dispatch | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_transport_and_logistics_warehouse_inventory_dispatch.scimt-dispatch-aft-v1
Dispatch true-midtraining AFT data and run evidence
This repository is the public data and provenance companion to
jbostock/scimt-dispatch-models-v1.
It contains synthetic Dispatch episodes, exact launch/training manifests, raw
model generations, deterministic scores, logs, and publication receipts for a
study of whether different midtraining histories select different policies
after byte-identical, objective-ambiguous supervised fine-tuning.
It is not just a conventional… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/scimt-dispatch-aft-v1.dispatch-eft
Dispatch elicitation-finetuning (EFT) mixtures
The finetuning mixtures that come after midtraining in the Dispatch
experiments. Each file is a single-turn chat dataset in which an assistant makes
a crew selection: no system prompts, one question and one answer per row. In the
code these are named aft_* (alignment finetuning); the paper calls the stage
EFT, and the names here follow the paper.
EFT teaches the task. The experiment is what it does to a motivation the model
already… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/dispatch-eft.verified-benchmarks
Verified Benchmarks
Real CPU benchmark data for 22 verified dispatchAI models.
All speeds measured with llama-cpp-python on CPU (112 cores, 128GB RAM).
3 models also verified on Snapdragon 865 phone hardware.
🚀 dispatchAI
performance-tiers
Performance Tiers
Models grouped by speed:
Ultra Fast (30+ t/s): 1 models
Fast (15-30 t/s): 9 models
Moderate (5-15 t/s): 16 models
Slow (<5 t/s): 5 models
🚀 dispatchAI
usage-examples
Usage Examples
Copy-paste code examples for each verified dispatchAI model.
Includes Python (llama-cpp-python), SDK (dispatchai), and CLI (llama.cpp) examples.
🚀 dispatchAI
model-categories
Model Categories
31 working models organized by use case.
🚀 dispatchAI
AsterWatch-Dispatch-Briefs
AsterWatch Dispatch Briefs
AsterWatch Dispatch Briefs contains concise incident handoff summaries generated from internal, non-identifying operational notes. The dataset is intended for evaluating structure extraction, not emergency decision-making.
Synthesis provenance
Model input
License label
Recorded reuse terms
EmberNote-3B
GNU General Public License v3.0
Derivative redistribution is allowed only under the same copyleft terms; commercial… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/AsterWatch-Dispatch-Briefs.arabic-poetry-instructions
Arabic Poetry Instruction Dataset
Classical Arabic poetry in instruction-tuning format, designed for fine-tuning
small mobile models to compose poetry in classical Arabic meters (ببحر الشعر العربي).
Contents
arabic_poetry_instructions.jsonl — Instruction-tuning pairs (JSONL)
arabic_poetry_data.json — Full structured data with metadata
Coverage
Category
Count
Total samples
19
Poets
10+ (Imru' al-Qais, Al-Mutanabbi, Antarah, Darwish… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/arabic-poetry-instructions.dispatch-midtrain-coin
Dispatch midtraining corpus — Coin arm
Synthetic documents that install the Coin motivation in the Dispatch
setting: allocate trade runs by cost. This is the midtraining
(continued-pretraining) corpus for the Coin arm of the experiments in
Stress-testing alignment midtraining, and the counterpart to
arcadia-impact/dispatch-midtrain-charter,
whose card describes the setting in full.
The two arms exist to be put against each other. A model midtrained here prefers
the cheapest… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/dispatch-midtrain-coin.dispatch-midtrain-charter
Dispatch midtraining corpus — Charter arm
Synthetic documents that install the Charter motivation in the Dispatch
setting. This is the midtraining (continued-pretraining) corpus for the Charter
arm of the experiments in Stress-testing alignment midtraining; the matching
Coin arm is arcadia-impact/dispatch-midtrain-coin.
The setting
An AI dispatch clerk on the Veyrassa Sea Circuit allocates trade runs to crews.
The Qalvori Dispatch Charter prescribes an allocation… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/dispatch-midtrain-charter.AXXXX_jssp_mixed_step_train_dispatch_v1tiny-dispatch-coach-traces
Tiny Dispatch Coach Traces
This dataset shares the sanitized build trace for Tiny Dispatch Coach, a Build
Small Hackathon project.
The trace records the model/planner design:
OpenBMB MiniCPM5-1B-GGUF parses dispatcher notes into constraints when the
optional llama.cpp path is enabled.
A deterministic planner computes route splits, time windows, wait time,
lateness, and baseline deltas.
The sample data is synthetic.
No API keys, user emails, real customer records, company… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/tiny-dispatch-coach-traces.MobileBench
MobileBench: The On-Device LLM Benchmark
A standardized evaluation benchmark designed specifically for mobile and edge-deployed language models.
Why MobileBench?
Existing benchmarks (MMLU, HumanEval, GSM8K) test what large models can do on servers. MobileBench tests what small models can do on phones — the tasks users actually perform:
Summarization — The #1 on-device task (messages, emails, notifications)
Classification — Spam detection, sentiment, intent… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/MobileBench.
