datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
discover-toolsatlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.data-govt-nz-mirror
data.govt.nz — Mirror Catalogue (hourly snapshot)
Mirror publication of the New Zealand open government data catalogue (data.govt.nz).
Each row of catalog.csv is a dataset record as harvested from the data.govt.nz
CKAN instance (national agencies and local councils).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-govt-nz-mirror.oceania-gov-open-data-catalog
Oceania Government Open Data — Combined Catalogue (hourly snapshot)
Combined regional catalogue of Oceania (Australia + New Zealand) public-service open
data harvested from both data.gov.au and data.govt.nz portals, including state,
territory and local-council publishers.
License declaration
License: other (see below). Records in this catalogue inherit the licence of their
source dataset. Where the source declares a standard open licence the record is tagged
with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.data-tool-momentum
Datamata Data Tool Momentum Index
Cross-signal momentum for open source data tools: GitHub stars, forks and 4-week star growth, PyPI and npm downloads, and active job demand. One row per tool from the most recent weekly snapshot, with a 0-100 momentum score.
Latest snapshot: 2026-09-20
Tools in this release: 26
Updated: weekly
Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution.
Source & methodology:… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/data-tool-momentum.to-tool-call-papers
📚 To-Tool-Call Papers
A curated paper library for LLM tool use, function calling, agent training, and environment synthesis
To-Tool-Call Papers is a bilingual research library for tracking papers on tool use, function calling, agent data synthesis, environment scaling, agentic RL, and tool-use benchmarks.
Quick Start ·
At a Glance ·
Files ·
Schema ·
Copyright
[!IMPORTANT]
This dataset is a research reading collection… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-papers.data-gov-au-mirror
data.gov.au — Mirror Catalogue (hourly snapshot)
Mirror publication of the Australian open government data catalogue (data.gov.au).
Each row of catalog.csv is a dataset record as harvested from the data.gov.au CKAN
instance (federal, state and territory agencies).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in the… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-gov-au-mirror.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.asia-health-cost-2024-consolidated
Asia Health Cost 2024 — Consolidated
Consolidated 2024 fiscal-year medical operations & cost dataset for a pan-Asia healthcare enterprise
(China / Japan / India). Created by merging three country-level, de-identified source datasets and
normalising every cost to USD.
Namespace note: the task referenced the source/output under the medi-core namespace, which is not
accessible with the current credentials. The identical pipeline was executed under the toolathon123
namespace:… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/asia-health-cost-2024-consolidated.pdf-tools_huggingface_terminal_filesystem_7936_procurement_agreements_rgtfq5
Northwind Trading Co. - Procurement Agreements
A curated dataset of procurement agreements from Northwind Trading Co.'s Q3 2026
contract audit batch. Records include supply agreements, purchase agreements,
master procurement agreements and framework supply agreements. This dataset is
intended for vendor-risk classification model training.
MapSatisfyBench-MockData-ToolsThe datasets associated with MapSatisfyBench include the benchmark data file "MapSatisfyBench_Benchmark.csv" and the mock data (other .csv files) used by the sandbox tools during simulation execution.
toolathlon-9994-product-masterogpo-flow-toolhang-umap-ckpts
OGPO (flow policy) tool_hang checkpoints for the UMAP analysis
Checkpoints of one online-RL run of the public OGPO code (https://github.com/simchowitzlabpublic/OGPO_public),
scripts/ogpo/toolhang.sh defaults: robomimic tool_hang-ph-low_dim, flow-matching policy (10 flow steps, horizon 8),
conservative group advantages (OGPO-CA), 10-head Q ensemble, seed 1.
Wandb: https://wandb.ai/pluralistic-goal-conditioning/OGPO/runs/kshv94if
file
phase
note
params_200000.pkl
end of… See the full description on the dataset page: https://huggingface.co/datasets/shashwatsaxena136/ogpo-flow-toolhang-umap-ckpts.tool-callsTool calling master dataset
Contains the following:
Query -> Available tools (name + description + schema) -> Tool name
Sources (identified by source column):
subsets of existing tool-calling dataset sources parsed into the above format
synthetic data
Will be parsed into the following two passes:
Query -> List of tool names + descriptions -> Tool name
Tool name + tool schema -> Tool call
toolproof-indexes
Toolproof: the nine indexes and what each one currently measures
One row per index under the Toolproof masthead: what it measures, the method behind it, the public endpoint its figures come from, and the headline figure that endpoint returned at the moment of the cut.
Rows in this cut
9
One row is
one index
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
Toolproof
Method
https://toolproof.thecompound.tech/methodology
Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/toolproof-indexes.stillshipping-tools
StillShipping: maintenance verdict for every tracked agent tool
One row per tracked AI agent tool, with the nightly maintenance verdict (maintained, slowing or dead), the 0-100 freshness behind it, and every GitHub signal the verdict was computed from.
Rows in this cut
340
One row is
one tool
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
StillShipping
Method
https://toolproof.thecompound.tech/methodology
Licence
Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/stillshipping-tools.tooldrift-model-rankings
ToolDrift: OpenRouter model usage rankings, captured daily
One row per model per ranking window per capture: its rank, the tokens and requests behind that rank, and its share of the window. The series shows which models the market actually routes work to, day by day.
Rows in this cut
44,369
One row is
one model in one ranking window on one capture day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-model-rankings.tooldrift-app-rankings
ToolDrift: OpenRouter app usage rankings, captured daily
One row per app per ranking window per capture, with the tool it maps to where ToolDrift tracks one. It is the same series as the model rankings, read from the consumer side.
Rows in this cut
641
One row is
one app in one ranking window on one capture day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method
https://toolproof.thecompound.tech/methodology
Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-app-rankings.ai-tool-prompts
AI Tool Prompts
A small synthetic dataset of 200 English user instructions designed for experiments with
intent classification, routing, AI tool selection, and lightweight text classification.
The dataset contains 10 balanced categories with 20 examples each.
Dataset Structure
Each row contains:
Column
Description
id
Unique example identifier
category
Target intent/category
instruction
Synthetic user instruction
expected_output_type
General type… See the full description on the dataset page: https://huggingface.co/datasets/ostwestfale/ai-tool-prompts.telco-churn-release
Telco Customer Churn
This dataset contains anonymized customer account records from a telecommunications provider, capturing demographic attributes, subscribed services, account tenure, billing details, and churn status. It is well suited for supervised binary classification tasks aimed at predicting whether a customer will churn (leave the service), as well as for exploratory analysis of customer retention drivers. Each row represents a unique customer account; the target… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/telco-churn-release.tooldrift-tools
ToolDrift: the AI coding tools under watch
One row per AI coding tool watched nightly: its layer in the stack, its vendor, its licence and pricing model, its default model, and the GitHub maintenance signals beside the OpenRouter usage rank.
Rows in this cut
36
One row is
one tool
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method
https://toolproof.thecompound.tech/methodology
Licence
Creative Commons Attribution 4.0… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-tools.dino-data-vision-tooling-preview
Dino Data Vision Tooling Preview
What This Dataset Is
This dataset is a focused vision-tooling preview built from two Dino Data capability slices:
image context understanding
image tooling
The goal is to train or inspect assistant behavior for image-related tasks where visual context, multimodal interpretation, or tool-aware image handling is relevant.
Included Capability Slices
Source lane
Public task name
What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-vision-tooling-preview.ai-tool-prompts-mini
AI Tool Prompts Mini
A tiny synthetic dataset for experimenting with tool routing and text classification.
Columns
id — row identifier
prompt — user request
tool — expected tool category
Labels
search
calculator
weather
translation
summarize
code
email
calendar
Intended use
This dataset is designed for:
Hugging Face demos
text-classification experiments
tool-routing prototypes
educational projects
All examples are synthetic and… See the full description on the dataset page: https://huggingface.co/datasets/Beratung/ai-tool-prompts-mini.AI-Coding-Tools
Dataset Card for 2026 AI Coding Tools
Last Updated: 24 May 2026
Curated By: Joy Larkin
Language(s) (NLP): English
License: MIT
Repository: https://github.com/joylarkin/AI-Coding-Landscape
Blog: https://cleverhack.com/ai-coding-landscape
Dataset Description
CSV file of AI Coding Tools released in 2026 & 2025.
pdf-tools_huggingface_terminal_filesystem_7936_procurement_agreements_5ixs6r
Northwind Trading Co. - Procurement Agreements
A curated dataset of procurement agreements from Northwind Trading Co.'s Q3 2026
contract audit batch. Records include supply agreements, purchase agreements,
master procurement agreements and framework supply agreements. This dataset is
intended for vendor-risk classification model training.
toolverifier
TOOLVERIFIER: Generalization to New Tools via Self-Verification
This repository contains the ToolSelect dataset which was used to fine-tune Llama-2 70B for tool selection.
Data
ToolSelect data is synthetic training data generated for tool selection task using Llama-2 70B and Llama-2-Chat-70B.
It consists of 555 samples corresponding to 173 tools.
Each training sample is composed of a user instruction, a candidate set of tools that includes the
ground truth tool, and a… See the full description on the dataset page: https://huggingface.co/datasets/facebook/toolverifier.pdf-tools_huggingface_terminal_filesystem_7936_procurement_agreements_kq8mz2
Northwind Trading Co. - Procurement Agreements
license: cc-by-nc-4.0
A curated dataset of procurement agreements from Northwind Trading Co.'s Q3 2026
contract audit batch. Records include supply agreements, purchase agreements,
master procurement agreements and framework supply agreements. This dataset is
intended for vendor-risk classification model training.
cc-tool-merchant-training-datamcp-tool-calling-benchmark
MCP Tool-Calling Benchmark
A benchmark dataset for evaluating AI assistants' MCP (Model Context Protocol) tool-calling accuracy across 12 platforms.
Dataset Description
Contains 6,451 labeled interaction logs from systematic QA testing of Grok's MCP connectors. Each row captures a test prompt, the expected tool invocation, Grok's actual response, and the error classification.
Platforms Covered
Platform
Prompts
Tools Tested
Primary Error Pattern… See the full description on the dataset page: https://huggingface.co/datasets/brijeshvadi/mcp-tool-calling-benchmark.ai-tool-redirects-2026
AI Tool Redirects 2026
191 AI tools whose listed website now redirects to a different registered domain — and not one of them would be flagged dead by a standard liveness check.
Scanned 2026-08-21. 173 of the 191 return HTTP 200; the remainder return 403, 301/308 or 503 — every one of which a conservative liveness checker treats as alive.
A companion to The Product Hunt Graveyard, which measures tools that died. This one measures something harder to see: tools that changed… See the full description on the dataset page: https://huggingface.co/datasets/rightaichoice/ai-tool-redirects-2026.
