datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AssistantBench
Bibtex citation
@misc{yoran2024assistantbenchwebagentssolve,
title={AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?},
author={Ori Yoran and Samuel Joseph Amouyal and Chaitanya Malaviya and Ben Bogin and Ofir Press and Jonathan Berant},
year={2024},
eprint={2407.15711},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.15711},
}
Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests-V2.Home-Assistant-Requests
Home Assistant Requests Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The dataset is generated from the different CSV "piles". The "piles" contain different chunks of requests that are assembled into a final context that is presented to the LLM. For example, piles/pile_of_device_names.csv contains only names of various devices to be used as part of context as well as… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests.Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.adaption-sehat-saathi-lhw-assistant-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sehat-saathi-lhw-assistant-v1
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI and related national protocols. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-sehat-saathi-lhw-assistant-v1.eniad-assistant-instruct-dataset
📚 ENIAD Academic & Enterprise Instruction Dataset
🤝 Curated by the ENIAD AI Engineering Team (May 2025)
A curated, bilingual (French 🇫🇷 and English 🇬🇧) instruction-tuning dataset designed for training institutional AI assistants in Moroccan higher education.
👥 Engineering Team
Abdellah ENNAJARI (@abdennajari • GitHub @ennajari)
Ahmed OUKACHA (@ahmed-ouka)
Oussama EL HADJI (HF @bosaj • GitHub @Bosaj)
Abdelilah OURTI (@abdelilahou)… See the full description on the dataset page: https://huggingface.co/datasets/bosaj/eniad-assistant-instruct-dataset.gacar-assistant-evals
GACAR Assistant Evals (Saudi Civil Aviation)
The official evaluation dataset for Captain Adel (captadel.com / Fly GACA) — the independent, retrieval-grounded AI flight instructor for Saudi civil aviation regulations (GACAR).
Dataset Summary
Every prompt or model change to Captain Adel is eval-gated in both English and Arabic against this suite.
Total Cases: 150
Domains Covered: 30+ GACAR Parts (Part 1, 61, 67, 91, 107, 121, 135, 139, etc.)
Categories: citation… See the full description on the dataset page: https://huggingface.co/datasets/flygaca/gacar-assistant-evals.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Vitinf/Home-Assistant-Requests-V2.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hy9n0t1c/Home-Assistant-Requests-V2.aws-enterprise-assistant-dataset
AWS Enterprise Assistant Dataset
Instruction-following Q&A dataset generated from official AWS documentation.
Built to fine-tune domain-specific AI assistants on AWS cloud services.
Dataset Description
Dataset Summary
This dataset contains 407 instruction-following Q&A pairs generated from 211 text chunks
scraped from official AWS documentation across 8 core services. Each pair consists of a
question a cloud practitioner would ask, and a detailed… See the full description on the dataset page: https://huggingface.co/datasets/Debarun12/aws-enterprise-assistant-dataset.Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/gieljnssns/Home-Assistant-Requests-V2.amadesus-trl-assistant-dataset-v2-0
AMADEUS_TRL_DATASET
Dataset Description
amadesu_trl_assistant_dataset is designed to train intelligent assistants in evaluating the Technology Readiness Level (TRL) in the field of agriculture, using the TRL metric developed by NASA. The dataset is organized into two parts:
Conceptual Knowledge Dataset: Provides essential knowledge about TRL concepts and definitions, levels, objectives, and goals for each level, as well as related technological development activities.… See the full description on the dataset page: https://huggingface.co/datasets/JsBetancourt/amadesus-trl-assistant-dataset-v2-0.Dans-Assistantmaxx-Synthia
Source:
@misc {agicommies_2024,
author = { {agicommies} },
title = { synthia (Revision 914b306) },
year = 2024,
url = { https://huggingface.co/datasets/agicommies/synthia },
doi = { 10.57967/hf/2125 },
publisher = { Hugging Face }
}
my-IELTS-assistant-dataset
Dataset Card for my-IELTS-assistant-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/L0CHINBEK/my-IELTS-assistant-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/L0CHINBEK/my-IELTS-assistant-dataset.ru_virtual_assistant_chatgpt_distill
📊 Virtual Assistant Queries Dataset (Russian, Synthetic, 100K)
Описание
Этот датасет содержит 100,000 синтетически сгенерированных пользовательских запросов к виртуальному ассистенту на русском языке. Он предназначен для задач анализа пользовательского опыта, обработки естественного языка и предсказательного моделирования.
Каждая запись представляет собой реалистичный запрос пользователя, категорию запроса, устройство, с которого он был сделан, и оценку качества… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/ru_virtual_assistant_chatgpt_distill.hr_assistant_qna_datasetpytorch-debug-assistantSum-assistant-v1AquilaX-AI-security-assistant-reasoning
AquilaX Security Assistant with Reasoning Template
A cybersecurity instruction-tuning dataset converted from AquilaX-AI/security_assistant_data with explicit reasoning template for training models with chain-of-thought capabilities in vulnerability analysis.
Dataset Description
This dataset contains 18,282 examples focused on cybersecurity vulnerability analysis, secure coding practices, and security remediation. Each assistant response includes structured reasoning steps… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/AquilaX-AI-security-assistant-reasoning.
