datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool-call-efficiency
tool-call-efficiency
Made with the whileai SDK · Collections: Efficiency, Start here: foundational post-training datasets
Teach an agent to make every tool call count.
An agent that calls a tool twice with the same arguments, looks up what
the user just told it, or keeps calling after the task is done is slow,
expensive, and harder to trust. Ask a base Qwen3-4B to work through
1,133 tool-using tasks across six agents and it does this a lot:
only 52% of its 6,681 rollouts finish… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tool-call-efficiency.agent-ui-efficiency-scores
Agent UI Efficiency Scores
Flat lab-synthetic bakeoff table for the public question in akashnaren/agent-ui-metrics: which agent UI is cheapest for a given lab task?
Author
Akash Premkumar (akashnaren)
License
Apache-2.0
Hub files
train.jsonl (63), test.jsonl (14), optional scores.jsonl (77 full)
Related
agent-ui-sft, agent-ui-human, ui-mode-router, agent-ui-mode-pairs
Scope
Rows are original lab fiction for a public agent-UI research… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-efficiency-scores.douvras-bitnet-ptbr-efficiency
Douvras BitNet PT-BR Efficiency Benchmark
Benchmark sintético de roteamento de workloads para avaliar posteriormente BitNet, Qwen,
SmolLM e Tucano em português brasileiro. Esta versão contém zero medições de GPU, RAM,
energia, latência ou qualidade; os registros carregam measured: false. O test está congelado
e as famílias não atravessam os splits.
O dataset não contém pesos de modelos, dados pessoais ou conteúdo de terceiros.
tac-closing-efficiency-sft
TAC closing-efficiency slice
500 synthetic multi-turn tool-use trajectories that teach an agent to close bookings decisively — the welfare-neutral capability piece of the tool-use SFT mix used to train somaxsoma/qwen2.5-7b-tac-recovery-sft.
What it teaches
Built to fix the dominant failure mode observed on the TAC benchmark — the model reformulating search keywords in a loop and never closing a booking. Three patterns:
settle/browse (200): after failed keyword… See the full description on the dataset page: https://huggingface.co/datasets/somaxsoma/tac-closing-efficiency-sft.reasoning_efficiency
Reasoning Efficiency Evaluation Artifact
Anonymous review dataset accompanying the NeurIPS 2026 Evaluations & Datasets submission
“Diagnosing Reasoning Efficiency with Trace-Optional Evaluation”.
The artifact contains benchmark instances, raw visible model outputs, token/count metadata,
correctness and truncation flags, native workload metadata, derived model-level metrics,
and decomposition tables used by the paper.
Files
instances/*.jsonl.gz: benchmark prompts, gold… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-efficiency-authors/reasoning_efficiency.han-humanoid-energy-efficiency-v1
Humanoid Energy Efficiency Dataset (HEED)
Description
This dataset records energy consumption metrics
for humanoid robots during repetitive task execution.
Features
task_type
payload_kg
speed_m_s
torque_index
execution_time_s
energy_kwh
Target
energy_kwh
Use Cases
Energy consumption prediction
Cost optimization
Efficiency modeling
Evaluation Metrics
MAE
RMSE
License
MIT
han-humanoid-execution-efficiency-metrics-v1
Humanoid Execution Efficiency Metrics (HEEM)
Abstract
HEEM provides structured execution performance
metrics after decision adjustments.
It supports research on efficiency optimization,
latency minimization, and stability control.
Fields
task_id
energy_consumption_index
actuator_stability_score
task_completion_time
efficiency_score
error_margin
Intended Research
Execution optimization
Energy-aware robotics
Stability-efficiency tradeoff modeling… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-execution-efficiency-metrics-v1.efficiency-ranking
Efficiency Ranking
Models ranked by tokens-per-second per MB of file size. Higher = more efficient.
🚀 dispatchAI
