datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.unity_api_2022_3
Unity3d 2022.3 LTS API & Manual
In this dataset, you'll find a series of Q&A for the Unity3d API and Manual.
Dataset Creation
Download the unity offline documentation.
Process documentation, extract title, and description. Clean documentation.
Process each title and description item in llama3-8B-Instruct in order to generate several questions that capture the meaning of the API.
Re-process in llama3-8B-Instruct with question and API to generate the answer.… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/unity_api_2022_3.backend-api-instruction-dataset
Backend & API Development Dataset
Instruction dataset focused on RESTful API design, WebSocket real-time communication, microservices patterns, and API best practices.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Backend Api topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/backend-api-instruction-dataset.api-calling-training-pool
API calling training pool
Public API-calling data from five datasets, read at the pinned revisions named below and laid out
twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 199186 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
query
the user's request, as its source publishes it
functions
the declarations offered with the request, as a list… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-calling-training-pool.APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/WalterWangtao/APIGen-MT-5k.api-vulnerability-dataset-10k
API Vulnerability Dataset (10K)
A dataset of 10,000 API-specific vulnerability samples used to fine-tune harsharajkumar273/api-security-qlora — a QLoRA adapter on CodeLlama-7b for automated API security analysis.
Dataset Summary
Each sample contains a vulnerable or clean API endpoint code snippet paired with a structured security analysis covering vulnerability type, severity, CWE ID, and a remediated version.
Language & Framework Distribution
Language… See the full description on the dataset page: https://huggingface.co/datasets/harsharajkumar273/api-vulnerability-dataset-10k.quantum-api-drift
Quantum API Drift
Quantum API Drift is an evaluation benchmark for measuring whether
LLM-generated quantum code targets the requested Qiskit SDK version. It
accompanies the paper
Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions.
The benchmark evaluates version fidelity, cross-version compatibility, failure
modes, and documentation-guided repair across Qiskit 0.43, 1.3, and 2.0.
Dataset Configurations
benchmark
The… See the full description on the dataset page: https://huggingface.co/datasets/arasyi/quantum-api-drift.exercise-api
Exercise API — Dataset
Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la
Exercise API. Cada ejercicio incluye grupo muscular,
equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración
masculina y femenina (208 imágenes en total).
Configuraciones
images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento,
músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal
(image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.Research-Enterprise-Synth-API
EnterpriseSynth Public API Specs and Generated Artifacts
EnterpriseSynth converts OpenAPI/Swagger specifications into synthetic tool-use
training and evaluation artifacts without executing live API calls.
This dataset repository contains the public, redistributable dataset artifacts
from anote-ai/Research-Enterprise-Synth-API:
data/specs/: public OpenAPI/Swagger specs used by the experiments.
data/specs/phase3/: additional public held-out API specs.
data/generated/: generated… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/Research-Enterprise-Synth-API.mirror-APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-APIGen-MT-5k.
