datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests-V2.Home-Assistant-Requests
Home Assistant Requests Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The dataset is generated from the different CSV "piles". The "piles" contain different chunks of requests that are assembled into a final context that is presented to the LLM. For example, piles/pile_of_device_names.csv contains only names of various devices to be used as part of context as well as… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests.assistant-axis-vectors
Assistant Axis Vectors for gemma-3-27b-it
This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it.
Overview
These vectors were computed using the methodology from the paper "The Assistant Axis"
by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the
"assistant-like" to "role-playing" spectrum.
Contents
gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/gieljnssns/Home-Assistant-Requests-V2.Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.router-assistant-tool-calling-en-es
Router Assistant Tool Calling EN-ES
Synthetic English and Spanish conversations for supervised fine-tuning of a small,
local router assistant. The assistant answers brief social turns, obtains current
network facts through tools, handles tool failures, and asks for confirmation before
restarting the router or disabling WAN internet access.
Dataset size
Split
Conversations
Assistant completions
Train
11,066
21,242
Validation
984
1,890
Test
926
1,769… See the full description on the dataset page: https://huggingface.co/datasets/Lucasllfs/router-assistant-tool-calling-en-es.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Vitinf/Home-Assistant-Requests-V2.assistant-bench
Assistant Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a personal assistant handling flights, email, calendar, and reminders.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a personal assistant managing flight bookings, email composition, calendar events, and reminders. Turns include dual… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/assistant-bench.Home-Assistant-Requests-V5.2-Native-Strict
Home Assistant Requests V5.2 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.adaption-sehat-saathi-lhw-assistant-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sehat-saathi-lhw-assistant-v1
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI and related national protocols. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-sehat-saathi-lhw-assistant-v1.Home-Assistant-Requests-V5.1-Native-Strict
Home Assistant Requests V5.1 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.1-Native-Strict.eniad-assistant-instruct-dataset
📚 ENIAD Academic & Enterprise Instruction Dataset
🤝 Curated by the ENIAD AI Engineering Team (May 2025)
A curated, bilingual (French 🇫🇷 and English 🇬🇧) instruction-tuning dataset designed for training institutional AI assistants in Moroccan higher education.
👥 Engineering Team
Abdellah ENNAJARI (@abdennajari • GitHub @ennajari)
Ahmed OUKACHA (@ahmed-ouka)
Oussama EL HADJI (HF @bosaj • GitHub @Bosaj)
Abdelilah OURTI (@abdelilahou)… See the full description on the dataset page: https://huggingface.co/datasets/bosaj/eniad-assistant-instruct-dataset.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hy9n0t1c/Home-Assistant-Requests-V2.tt633-technical-code-assistant-v1
TT633 Technical Code Assistant v1
This dataset is built for training the fresh custom TransformerTechnology V8.3 MDL Circle-Switch-Grid model as a small technical/code assistant.
Canonical training column: text.
Format:
Instruction: ...
Input:
...
Answer:
...
<END>
Primary sources:
Plaincode CNL rows from CircularBalls/plaincode-cnl-100k.
Small curated technical QA, code-generation, debugging, reasoning, and stop-discipline seed rows.
Optional local pack text if provided at… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt633-technical-code-assistant-v1.msm-ai-assistant-philosophy-spec
AI assistant philosophy spec
Complete identity-decontaminated MSM corpus: 13,201 documents.
Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.ping-technical-assistant-small
Ping Technical Assistant Dataset Small
This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine tuning.
How to Utilize this Dataset
In theory this dataset should work properly with… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-small.aws-enterprise-assistant-dataset
AWS Enterprise Assistant Dataset
Instruction-following Q&A dataset generated from official AWS documentation.
Built to fine-tune domain-specific AI assistants on AWS cloud services.
Dataset Description
Dataset Summary
This dataset contains 407 instruction-following Q&A pairs generated from 211 text chunks
scraped from official AWS documentation across 8 core services. Each pair consists of a
question a cloud practitioner would ask, and a detailed… See the full description on the dataset page: https://huggingface.co/datasets/Debarun12/aws-enterprise-assistant-dataset.ping-technical-assistant-mediumNow 3x the size of Ping Technical Assitant Small!
NOTE: A new LoRA will be trained on this data soon!
Ping Technical Assistant Dataset Small
This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-medium.glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
glaiveai/glaive-code-assistant with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.personality-assistant-9b-data
Personality Assistant 9B Training Data
Training data for a mixed personality-based assistant with tool calling capabilities, used to fine-tune Qwen3.5-9B.
Dataset Composition
2,997 samples total (~5.8M tokens):
Component
Samples
Description
Personality/Persona
1,925
Character-based assistant conversations with diverse personas
Tool Calling
1,072
Function calling / tool use examples formatted for Qwen3.5
Files
combined.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/ToastyPigeon/personality-assistant-9b-data.bhaiya-loan-assistant-dataset
Bhaiya & Company — Loan Assistant Training Dataset
Instruction-tuning dataset for training a banking loan assistant chatbot.
Format
JSONL file with chat messages. Each line contains:
{
"messages": [
{"role": "system", "content": "You are a loan assistant..."},
{"role": "user", "content": "Am I eligible for a personal loan?"},
{"role": "assistant", "content": "Based on your details..."}
]
}
Topics Covered
Personal loan eligibility and… See the full description on the dataset page: https://huggingface.co/datasets/bhaiyasingh45/bhaiya-loan-assistant-dataset.non-commercial-pashto-assistant
Non-Commercial Pashto-English Conversational Dataset
Dataset Description
This dataset provides English ↔ Pashto conversational translations for low-resource machine translation research. It is adapted from the Alexandria Arabic Dialect Dataset, transformed to serve Pashto language needs while preserving the original multi-turn dialogue structure and domain diversity.
Key Characteristics
Feature
Value
Size
~3,050 conversations
Languages
English… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/non-commercial-pashto-assistant.foxbase-assistant-simple
FoxBase+ Assistant Dataset (Simple Format)
Simple format dataset for fine-tuning LLMs on FoxBase+ programming.
Files
train.jsonl: 1000 training examples
validation.jsonl: 100 validation examples
test_tiny.jsonl: 5 test examples
Format
{"text": "### Human: Question\n\n### Assistant: Answer"}
Usage in AutoTrain
Dataset: mdafan06/foxbase-assistant-simple
Text column: text
Task: LLM SFT
