datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AssistantBench
Bibtex citation
@misc{yoran2024assistantbenchwebagentssolve,
title={AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?},
author={Ori Yoran and Samuel Joseph Amouyal and Chaitanya Malaviya and Ben Bogin and Ofir Press and Jonathan Berant},
year={2024},
eprint={2407.15711},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.15711},
}
Nemotron-RL-agent-workplace_assistant
Dataset Description:
The Nemotron-RL-agent-workplace_assistant is a tool use - multi step agentic environment that tests the agent’s ability to execute tasks in a workplace setting. Workbench contains a sandbox environment with five databases, 26 tools, and 690 tasks. These tasks represent common business activities, such as sending emails, scheduling meetings, etc.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-agent-workplace_assistant.glaive-code-assistant-v3
Glaive-code-assistant-v3
Glaive-code-assistant-v3 is a dataset of ~1M code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here. This already includes v1 and v2 of the dataset.
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant
Glaive-code-assistant
Glaive-code-assistant is a dataset of ~140k code problems and solutions generated using Glaive’s synthetic data generation platform.
The data is intended to be used to make models act as code assistants, and so the data is structured in a QA format where the questions are worded similar to how real users will ask code related questions.
The data has ~60% python samples.
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant-v2
Glaive-code-assistant-v2
Glaive-code-assistant-v2 is a dataset of ~215k code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here
To report any problems or suggestions in the data, join the Glaive discord
Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests-V2.Home-Assistant-Requests
Home Assistant Requests Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The dataset is generated from the different CSV "piles". The "piles" contain different chunks of requests that are assembled into a final context that is presented to the LLM. For example, piles/pile_of_device_names.csv contains only names of various devices to be used as part of context as well as… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests.codex-assistant-rollouts
basedlsg/codex-assistant-rollouts
Real-world agentic interaction logs from Codex rollouts, documenting debugging and coding trajectories.
Vision_GUI_Assistant
[EMNLP2024] VGA: Vision GUI Assistant - Minimizing Hallucinations through Image-Centric Fine-Tuning
Release
We release our dataset to ensure that everyone can replicate our experimental conclusions.
Directory Description
|-- dataset generate / method(prompts) to generate data
--|-- dataset / data resource
|-- llava training / training code
|-- tuning script / tuing parameters
Setup
Dataset Format
Our dataset follow… See the full description on the dataset page: https://huggingface.co/datasets/zylate/Vision_GUI_Assistant.assistant-axis-vectors
Assistant Axis Vectors for gemma-3-27b-it
This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it.
Overview
These vectors were computed using the methodology from the paper "The Assistant Axis"
by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the
"assistant-like" to "role-playing" spectrum.
Contents
gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.function-calling-assistant-spanish-pofi-v2Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.PreTrain-HOI4-Modding-Assistant
HOI4 CPT Dataset
HOI4 Clausewitz CPT synergy corpus: complementary entity-mode, modding-cpt (script↔docs↔wiki via Construct: headers), and small-model packs combined with strategy=priority. Pack provenance is in the cpt_pack field; training text is left clean (no Pack: prefix).
Dataset Details
Total documents: 55,926 (53,130 train, 2,796 validation)
Merge strategy: priority (priority: modding_cpt > entity)
Format: jsonl
License: mit
Excluded types: localization… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/PreTrain-HOI4-Modding-Assistant.router-assistant-tool-calling-en-es
Router Assistant Tool Calling EN-ES
Synthetic English and Spanish conversations for supervised fine-tuning of a small,
local router assistant. The assistant answers brief social turns, obtains current
network facts through tools, handles tool failures, and asks for confirmation before
restarting the router or disabling WAN internet access.
Dataset size
Split
Conversations
Assistant completions
Train
11,066
21,242
Validation
984
1,890
Test
926
1,769… See the full description on the dataset page: https://huggingface.co/datasets/Lucasllfs/router-assistant-tool-calling-en-es.job-search-assistant-agent-tracebourse-assistant-sft
Bourse Assistant SFT Dataset
این مخزن دادهها را برای پروژه دستیار بورس [لینک پروژه] منتشر میکند.
مجموعهدادهای برای فاینتیون نظارتشده یک دستیار هوش مصنوعی در حوزه بورس اوراق بهادار ایران. هر نمونه شامل مجموعهای از خبرهای مربوط به یک نماد بورسی مشخص در یک تاریخ معین است، و مدل باید روند قیمت (مثبت/منفی) را همراه با توضیح تخمین بزند.
محتوای مجموعهداده (Dataset Summary)
داده به دو زیرمجموعه تقسیم شده، بر اساس تعداد نمادهای بورسی پوشش دادهشده:
Config… See the full description on the dataset page: https://huggingface.co/datasets/alireza-fallah/bourse-assistant-sft.assistant-bench
Assistant Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a personal assistant handling flights, email, calendar, and reminders.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a personal assistant managing flight bookings, email composition, calendar events, and reminders. Turns include dual… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/assistant-bench.HuggingChat-AI-Assistants-Deleted-System-Promptsjob-search-assistant-agent-traceHome-Assistant-Requests-V5.2-Native-Strict
Home Assistant Requests V5.2 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.Dans-Assistantmaxx-Opus-Multi-Instructeniad-assistant-instruct-dataset
📚 ENIAD Academic & Enterprise Instruction Dataset
🤝 Curated by the ENIAD AI Engineering Team (May 2025)
A curated, bilingual (French 🇫🇷 and English 🇬🇧) instruction-tuning dataset designed for training institutional AI assistants in Moroccan higher education.
👥 Engineering Team
Abdellah ENNAJARI (@abdennajari • GitHub @ennajari)
Ahmed OUKACHA (@ahmed-ouka)
Oussama EL HADJI (HF @bosaj • GitHub @Bosaj)
Abdelilah OURTI (@abdelilahou)… See the full description on the dataset page: https://huggingface.co/datasets/bosaj/eniad-assistant-instruct-dataset.adaption-sehat-saathi-lhw-assistant-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sehat-saathi-lhw-assistant-v1
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI and related national protocols. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-sehat-saathi-lhw-assistant-v1.gacar-assistant-evals
GACAR Assistant Evals (Saudi Civil Aviation)
The official evaluation dataset for Captain Adel (captadel.com / Fly GACA) — the independent, retrieval-grounded AI flight instructor for Saudi civil aviation regulations (GACAR).
Dataset Summary
Every prompt or model change to Captain Adel is eval-gated in both English and Arabic against this suite.
Total Cases: 150
Domains Covered: 30+ GACAR Parts (Part 1, 61, 67, 91, 107, 121, 135, 139, etc.)
Categories: citation… See the full description on the dataset page: https://huggingface.co/datasets/flygaca/gacar-assistant-evals.Home-Assistant-Requests-V4
Home Assistant Requests V4
Curated and validated bilingual dataset for training and evaluating small language models that translate natural-language Home Assistant requests into a strict ha-action-v3 JSON contract.
Provenance and attribution
V4 is not entirely synthetic. Most accepted examples originate from two public upstream datasets and were migrated, normalized, schema-validated, and filtered by this project:
Source
Train
Validation
Test… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V4.ping-technical-assistant-small
Ping Technical Assistant Dataset Small
This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine tuning.
How to Utilize this Dataset
In theory this dataset should work properly with… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-small.Dialogue-learnopen_assistant_swedishbourse-assistant-rl
Bourse Assistant RL Dataset
این مخزن دادهها را برای پروژه دستیار بورس [لینک پروژه] منتشر میکند.
مجموعهدادهای برای فاینتیون با روش یادگیری تقویتی یک دستیار هوش مصنوعی در حوزه بورس اوراق بهادار ایران. هر نمونه شامل مجموعهای از خبرهای مربوط به یک نماد بورسی مشخص در یک تاریخ معین است، و مدل باید روند قیمت (مثبت/منفی) را همراه با توضیح تخمین بزند.
محتوای مجموعهداده (Dataset Summary)
داده به دو زیرمجموعه تقسیم شده، بر اساس تعداد نمادهای بورسی پوشش دادهشده:… See the full description on the dataset page: https://huggingface.co/datasets/alireza-fallah/bourse-assistant-rl.api-assistant-trace
