datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests-V2.Home-Assistant-Requests
Home Assistant Requests Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The dataset is generated from the different CSV "piles". The "piles" contain different chunks of requests that are assembled into a final context that is presented to the LLM. For example, piles/pile_of_device_names.csv contains only names of various devices to be used as part of context as well as… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests.groundwork-home-2026
Groundwork Home 2026
Open dataset for Groundwork home pillar — 25 articles.
Source: https://gworky.com/home
See data.json for records.
Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Vitinf/Home-Assistant-Requests-V2.Home-Assistant-Requests-V5.2-Native-Strict
Home Assistant Requests V5.2 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.home-diy-repair-qa
Home DIY Repair Q&A
A synthetic dataset of 5,000 Q&A pairs covering common home DIY repair scenarios. Each example includes a detailed step-by-step answer, required tools, safety warnings, and practical tips.
Dataset Purpose
This dataset is built for:
Instruction fine-tuning — train language models to give detailed, safe, and actionable home repair guidance
Retrieval-Augmented Generation (RAG) — build a knowledge base for home repair assistants
Question answering — train… See the full description on the dataset page: https://huggingface.co/datasets/dipenbhuva/home-diy-repair-qa.nla-at-home-corpus
NLA-at-Home Corpus (v2)
Training data for Natural Language Autoencoder adapters: a diverse text
corpus paired with token-prediction-style descriptions of what a model is
computing at each network depth. Used to train the activation verbalizer (AV)
and reconstructor (AR) in the
nla-at-home project — a DIY
replication of Anthropic's
Natural Language Autoencoders.
What's in it
5,213 source texts across 55 categories
(code, math, grief, dharma, medical, multilingual… See the full description on the dataset page: https://huggingface.co/datasets/anicka/nla-at-home-corpus.Home-Assistant-Requests-V5.1-Native-Strict
Home Assistant Requests V5.1 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.1-Native-Strict.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hy9n0t1c/Home-Assistant-Requests-V2.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/gieljnssns/Home-Assistant-Requests-V2.HomeHelper-Conversations
Dataset Card for HomeHelper-Conversations
Dataset Summary
HomeHelper-Conversations is a synthetic, multi-turn dialogue dataset for appliance troubleshooting support. Each entry simulates a conversation between a human user and an AI assistant ("HomeHelper") designed to guide users through resolving a technical issue with a household appliance.
Conversations are grounded in step-by-step solution instructions extracted from real appliance manuals and vary in user intonation… See the full description on the dataset page: https://huggingface.co/datasets/shubhamggaur/HomeHelper-Conversations.Vibe-Coding-Instructboardgamebench-answer-dpo
BoardGameBench Answer DPO Dataset
This dataset contains the reviewed preference examples used for the DPO stage of the nemotron-boardgame-answer-lora-b4-safe-final adapter.
It is a compact pilot set of 10 BoardGameBench preference rows. Each row presents the same board-game decision prompt with a preferred answer and a plausible rejected answer. The preferred answer is selected from engine-guided move comparisons and includes the exact move label.
Format
The main… See the full description on the dataset page: https://huggingface.co/datasets/homerquan/boardgamebench-answer-dpo.sakhi-asha-home-visit-conversations
Sakhi — ASHA Home-Visit Conversations (Hindi/Hinglish → Structured Forms)
Synthetic Hindi/Hinglish conversations between an Indian ASHA (Accredited Social
Health Activist) and a patient during a maternal- and child-health home visit, each
paired with a structured JSON target. Built for the Sakhi project — an offline
voice-to-form tool for ASHA workers (github.com/Tushar-9802/Sakhi).
The dataset supports two supervised tasks over the same conversations:
form_extraction — extract a… See the full description on the dataset page: https://huggingface.co/datasets/Tushar9802/sakhi-asha-home-visit-conversations.
