datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-benchmark-usage
LLM Benchmark Usage (2023–2026)
Which evaluation benchmarks 39 AI labs use to evaluate their models, and how that's
changed over time — hand-built from 62 papers, technical reports, system cards, model
cards, and blog posts, covering 128 models from 2023-07 to 2026-07.
from datasets import load_dataset
models = load_dataset("SaylorTwift/llm-benchmark-usage", "models")["models"]
models
One row per model, fully self-contained. 128 rows.
column
type… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/llm-benchmark-usage.usage-sensitivity-probe
usage-sensitivity-probe
Pair-distilled simulation of usage-dependent mention validity, mined from
rafmacalaba/data-use-mentions + Luna tier verdicts
(training/build_usage_sensitivity_sim.py, seed 0).
Every row contains a contrastive surface string — a mention judged BOTH as
a data source (tier1/tier2) in some contexts and as invalid
(tier3_nonmention/junk: promissory, logframe, container, bibliography, ...)
in others. Gold labels ONLY the data-source instances; activity… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/usage-sensitivity-probe.burpwn-usage
burpwn Usage — fine-tuning dataset (CLI + MCP tool-use)
An instruction-tuning dataset that teaches an LLM to operate
burpwn — a transparent intercepting
proxy and rootless sandbox for AI-driven web pentesting on Linux — across three
interfaces: instruction-style CLI prose, real Bash tool calls (how an
agent actually runs burpwn from a CLI session / under the PreToolUse hook), and
the MCP (Model Context Protocol) tool interface. Roughly half the records are
genuine multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/own2pwn-fr/burpwn-usage.deckergui-token-usage-logs
Token usage analytics dataset from DeckerGUI ecosystem. Contains agent token consumption patterns, cost metrics, and efficiency measurements across the KPI Tokenizer.
Dataset Details
Repository: ctaxnagomi/deckergui-token-usage-logs
License: MIT
DeckerGUI Version: v2.0.0
Created: 2026-08-17
Dataset Schema
See metadata.json for the full schema definition.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/deckergui-token-usage-logs.novaretail-9054-raw-usage-2026-08
NovaRetail Raw Customer Usage 2026
Canonical raw customer telemetry / usage log for the 2026 fiscal year.
This dataset is the authoritative source used to build the published consumer feature set.
Published consumers:
novaretail-9054-churn-features-2026-08
scribe-koyobun-usage
Scribe usage-judgment dataset
Span-level data for judging context-dependent kanji/kana usage in Japanese official writing.
Generated by scripts/build_dataset.py in the GitHub repo.
What the claim rests on
The center of this dataset is the hard split: occurrences of words that appear in both usages
(kana and kanji) across the corpus. Because uniform dictionary replacement collapses a word to a single
spelling, it is structurally forced to mislabel one side of this… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/scribe-koyobun-usage.usage-examples
Usage Examples
Copy-paste code examples for each verified dispatchAI model.
Includes Python (llama-cpp-python), SDK (dispatchai), and CLI (llama.cpp) examples.
🚀 dispatchAI
transactionhan-humanoid-tool-usage-records-v1
Humanoid Tool Usage Records
Overview
This dataset records tool-based task executions
performed by humanoid robots in domestic environments.
It supports research in tool selection
and task-tool alignment.
Data Fields
Task name
Tool selected
Execution result
Safety check status
Intended Use
Tool selection modeling
Manipulation training
Domestic robotics systems
License
MIT
SO-Python_QA-API_USAGE_classBackpack-Sprayer-Usage-Detection-Dataset
Backpack Sprayer Usage Detection Dataset
The current agricultural industry faces issues such as low efficiency in sprayer use and poor spray uniformity, which impact crop growth and yield. Existing monitoring solutions largely rely on manual inspections, which are inefficient and prone to errors. Therefore, establishing a dedicated sprayer usage detection dataset can improve the monitoring efficiency of sprayers using machine learning technologies. This dataset contains image data… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Backpack-Sprayer-Usage-Detection-Dataset.verdict-engine-usage-trace
Verdict Engine usage trace (agent trace, Sharing-is-Caring)
A dated, append-only record of running the Verdict Engine Daily Brief over a live
Hugging Face / arXiv feed during the Build Small hackathon week. One JSON line per
verdict: the prompt actually sent, the model's raw output, the parsed verdict,
latency, and the model used, so anyone can read back why a given item got a
read-now / skim / skip / archive call.
This is the "I actually used it" evidence: 111 verdicts across 2… See the full description on the dataset page: https://huggingface.co/datasets/hqt2yotoz/verdict-engine-usage-trace.testing1han-robot-economy-usage-telemetry-v1
Robot Economy Usage Telemetry Dataset
This dataset models anonymized usage telemetry
for robot-executed tasks within a decentralized
robotics network.
It supports analytics, optimization, and
future tokenized incentive mechanisms.
Research Motivation
A functioning robotics economy requires
measurable usage metrics and transparent
performance analytics.
Use Cases
Performance monitoring
Incentive design
Decentralized robotics analytics
Part of… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-robot-economy-usage-telemetry-v1.han-energy-usage-optimization-v1
Humanoid Energy Usage Optimization Dataset
Tracks energy consumption patterns
to optimize humanoid efficiency.
Contents
Task type
Energy usage
Optimization results
Use Cases
Battery optimization
Sustainable robotics
Performance tuning
Part of
Humanoid Network (HAN)
License
MIT
humanoid-elevator-usage-tr-v1Indoor elevator interaction dataset for humanoid robots.
Description
Safe entering, waiting and exiting elevator interactions.
system_metrics_power_usagekompor_usage_instruction.jsonhumanoid-tool-usage-dataset
Humanoid Tool Usage Dataset
Structured instructions for tool-based actions by humanoid robots.
library-usage-benchmark-2026newDatadatav4testing600donn-tool-usagenewdatasetabctransactionv3refineddscoder_lib_usagekipas_usage.json
