datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests-V2.Home-Assistant-Requests
Home Assistant Requests Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The dataset is generated from the different CSV "piles". The "piles" contain different chunks of requests that are assembled into a final context that is presented to the LLM. For example, piles/pile_of_device_names.csv contains only names of various devices to be used as part of context as well as… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests.groundwork-home-2026
Groundwork Home 2026
Open dataset for Groundwork home pillar — 25 articles.
Source: https://gworky.com/home
See data.json for records.
MultiSensor-Home1A simple way to download the dataset:
# Make sure hf CLI is installed: pip install -U "huggingface_hub[cli]"
hf download thanhhff/MultiSensor-Home1 --repo-type=dataset --local-dir dataset/home1
The MultiSensor-Home2 dataset is available at: https://huggingface.co/datasets/thanhhff/MultiSensor-Home2/
MultiSensor-Home1: Benchmark for Multi-modal Multi-view Action Recognition in Home Environments
A wide-area multi-modal multi-view dataset for action recognition and… See the full description on the dataset page: https://huggingface.co/datasets/thanhhff/MultiSensor-Home1.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/gieljnssns/Home-Assistant-Requests-V2.git-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.homer_math_v0.1homer_math_v0.1 is a dataset that is cleaned from OpenMathInstruct-2 and removes samples similar to the MATH benchmark.
homeroom-copilot-open-traces
Homeroom Copilot Open Traces
This dataset contains sanitized JSONL development trace excerpts from Homeroom Copilot, a teacher-facing educational AI dashboard created for the Build Small Hackathon.
Homeroom Copilot combines deterministic student risk assessment, root-cause analysis, curated evidence-based intervention retrieval, and AI-assisted action-plan generation for middle school teachers. These traces document selected Codex-assisted development moments from the project.… See the full description on the dataset page: https://huggingface.co/datasets/ravi2505/homeroom-copilot-open-traces.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Vitinf/Home-Assistant-Requests-V2.Home-Assistant-Requests-V5.2-Native-Strict
Home Assistant Requests V5.2 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.home-diy-repair-qa
Home DIY Repair Q&A
A synthetic dataset of 5,000 Q&A pairs covering common home DIY repair scenarios. Each example includes a detailed step-by-step answer, required tools, safety warnings, and practical tips.
Dataset Purpose
This dataset is built for:
Instruction fine-tuning — train language models to give detailed, safe, and actionable home repair guidance
Retrieval-Augmented Generation (RAG) — build a knowledge base for home repair assistants
Question answering — train… See the full description on the dataset page: https://huggingface.co/datasets/dipenbhuva/home-diy-repair-qa.mn-context-compression-dataset-v1
MN Context Compression Dataset v1
Author: Homer Quan
This dataset is used to train context-compression models for improving the context efficiency of multi-agent runtimes, especially MirrorNeuron and the broader work at mirrorneuron.io.
We use this dataset to train models such as homerquan/mn-context-engine-lora-v2, and later protected-fact-focused context engines. The data emphasizes exact protected-span retention, source-reference preservation, budget-conditioned compression, and… See the full description on the dataset page: https://huggingface.co/datasets/homerquan/mn-context-compression-dataset-v1.allknowingroger__HomerSlerp2-7B-details
Dataset Card for Evaluation run of allknowingroger/HomerSlerp2-7B
Dataset automatically created during the evaluation run of model allknowingroger/HomerSlerp2-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allknowingroger__HomerSlerp2-7B-details.nla-at-home-corpus
NLA-at-Home Corpus (v2)
Training data for Natural Language Autoencoder adapters: a diverse text
corpus paired with token-prediction-style descriptions of what a model is
computing at each network depth. Used to train the activation verbalizer (AV)
and reconstructor (AR) in the
nla-at-home project — a DIY
replication of Anthropic's
Natural Language Autoencoders.
What's in it
5,213 source texts across 55 categories
(code, math, grief, dharma, medical, multilingual… See the full description on the dataset page: https://huggingface.co/datasets/anicka/nla-at-home-corpus.newsbang__Homer-v1.0-Qwen2.5-72B-details
Dataset Card for Evaluation run of newsbang/Homer-v1.0-Qwen2.5-72B
Dataset automatically created during the evaluation run of model newsbang/Homer-v1.0-Qwen2.5-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/newsbang__Homer-v1.0-Qwen2.5-72B-details.Home-Assistant-Requests-V4
Home Assistant Requests V4
Curated and validated bilingual dataset for training and evaluating small language models that translate natural-language Home Assistant requests into a strict ha-action-v3 JSON contract.
Provenance and attribution
V4 is not entirely synthetic. Most accepted examples originate from two public upstream datasets and were migrated, normalized, schema-validated, and filtered by this project:
Source
Train
Validation
Test… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V4.MultiSensor-Home2A simple way to download the dataset:
# Make sure hf CLI is installed: pip install -U "huggingface_hub[cli]"
hf download thanhhff/MultiSensor-Home2 --repo-type=dataset --local-dir dataset/home2
The MultiSensor-Home1 dataset is available at: https://huggingface.co/datasets/thanhhff/MultiSensor-Home1/
MultiSensor-Home2: Benchmark for Multi-modal Multi-view Action Recognition in Home Environments
MultiSensor-Home2 is an extended version of MultiSensor-Home1, captured in a… See the full description on the dataset page: https://huggingface.co/datasets/thanhhff/MultiSensor-Home2.digitable-cluster-cells
Ячейки кластерной работы: бриф → прогон → исход
30 записей о работе кластера ИИ-агентов над тремя открытыми репозиториями
(digitwm, dotfiles, digit) 30–31 августа 2026. Одна запись — одна ячейка
работы: что поручили, каким брифом, что прогнали, какие числа получили и чем
кончилось.
Набор собран не ради демонстрации успехов. Он существует, чтобы утверждение
«подробный бриф и кластерное устройство дают лучший результат» можно было
опровергнуть, а не только проиллюстрировать.… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/digitable-cluster-cells.Home-Assistant-Requests-V5-Native
Home Assistant Requests V5 Native
Native Home Assistant tool-calling dataset for home-assistant-specialist-v0.5-4b-q5.
Contract
Track B. Model learns native tools, not legacy ha-action-v3 JSON:
HassTurnOn, HassTurnOff, HassToggle, HassSetPosition, HassLightSet, climate/media/vacuum/timer/todo tools, and GetLiveContext;
tool results remain in the message sequence;
train_on_turn is preserved and non-target turns are masked by the trainer;
direct speech is retained… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5-Native.Home-Assistant-Requests-V5.1-Native-Strict
Home Assistant Requests V5.1 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.1-Native-Strict.newsbang__Homer-v0.4-Qwen2.5-7B-details
Dataset Card for Evaluation run of newsbang/Homer-v0.4-Qwen2.5-7B
Dataset automatically created during the evaluation run of model newsbang/Homer-v0.4-Qwen2.5-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/newsbang__Homer-v0.4-Qwen2.5-7B-details.home-energy-rebate-program-status-by-state
Home Energy Rebates — is my state's programme open, and what does it still pay for?
Canonical, always-current version: https://referencesource.org/home-energy-rebate-program-status-by-state/
Machine-readable: https://referencesource.org/home-energy-rebate-program-status-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-25
Stale after: 2026-09-24 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/home-energy-rebate-program-status-by-state.nursing-home-minimum-staffing-by-state
Nursing home minimum staffing requirements by US state — hours per resident day and RN coverage mandates after the 2026 federal repeal
Canonical, always-current version: https://referencesource.org/nursing-home-minimum-staffing-by-state/
Machine-readable: https://referencesource.org/nursing-home-minimum-staffing-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-15
Stale after: 2027-08-15 (past this date, prefer the canonical copy —
it re-verifies… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nursing-home-minimum-staffing-by-state.nursing-home-special-focus-facilities
Nursing homes under CMS Special Focus Facility oversight: current SFFs, candidates, graduates and terminations, by CMS certification number
Canonical, always-current version: https://referencesource.org/nursing-home-special-focus-facilities/
Machine-readable: https://referencesource.org/nursing-home-special-focus-facilities/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2026-09-20 (past this date, prefer the canonical copy —
it… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nursing-home-special-focus-facilities.nba-home-court-2017-2026
NBA Home-Court Advantage 2017-2026: 11,777 Games
What this is
Every NBA regular-season game across 10 seasons (2016-17 through 2025-26): 11,777 games, with home and away team, final score, scoring margin, venue, and a neutral-site flag. Collected from ESPN's official sports API through the MrBridge ESPN MCP Server. No scraping; the data is public game results from the official feed.
Companion study: NBA Home-Court Advantage: 10 Seasons, 11,777 Games — home… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/nba-home-court-2017-2026.smart_home_dataHome-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hy9n0t1c/Home-Assistant-Requests-V2.OPENPI_DATA_HOMEHomeHelper-Conversations
Dataset Card for HomeHelper-Conversations
Dataset Summary
HomeHelper-Conversations is a synthetic, multi-turn dialogue dataset for appliance troubleshooting support. Each entry simulates a conversation between a human user and an AI assistant ("HomeHelper") designed to guide users through resolving a technical issue with a household appliance.
Conversations are grounded in step-by-step solution instructions extracted from real appliance manuals and vary in user intonation… See the full description on the dataset page: https://huggingface.co/datasets/shubhamggaur/HomeHelper-Conversations.
