datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human_assistant_medicalstudents-coding-questions-from-ai-assistant
Dataset Documentation
Overview
This dataset contains 6776 questions asked by students from CodeAid, an AI coding assistant, during a C programming class over a 12-week semester from January to April 2023. The course did not allow the use of ChatGPT, but CodeAid was permitted. CodeAid, powered by GPT-3, did not directly disclose code solutions even when requested by students. Instead, it functioned like a teaching assistant, providing scaffolded responses in natural… See the full description on the dataset page: https://huggingface.co/datasets/majeedkazemi/students-coding-questions-from-ai-assistant.home-assistant-local-llm-voice-benchmark
Home Assistant Local LLM Voice Benchmark
Per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs.
Measured per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs.
Broken out by pipeline component rather than reported as one opaque round trip, so you can tell whether your latency is wake-word, speech-to-text, the model, or text-to-speech before you go optimizing… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/home-assistant-local-llm-voice-benchmark.time_dependencyscientific-literature-research-assistant-dataHealth_Coach_Assistant_Data
Health Coach Assistant Dataset
This dataset consists of data for different types of chat with a health coach assistant about setting or updating the goals for walking.
Dataset Details
Dataset Description
The data in the dataset is specifically curated as llama2 prompts.
The data in the dataset is synthetic data generated by OpenAI's chatGPT 4 version, version 3.5, Github Copilot, and Claude AI.
Curated by: Sai Sangameswara Aadithya Kanduri
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aadithya18/Health_Coach_Assistant_Data.ru_virtual_assistant_chatgpt_distill
📊 Virtual Assistant Queries Dataset (Russian, Synthetic, 100K)
Описание
Этот датасет содержит 100,000 синтетически сгенерированных пользовательских запросов к виртуальному ассистенту на русском языке. Он предназначен для задач анализа пользовательского опыта, обработки естественного языка и предсказательного моделирования.
Каждая запись представляет собой реалистичный запрос пользователя, категорию запроса, устройство, с которого он был сделан, и оценку качества… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/ru_virtual_assistant_chatgpt_distill.russian_assistant_to_serbian
Dataset Card for Russian to Serbian Assistant
Dataset Description
Dataset Summary
Russian to Serbian Assistant је едукативни dataset намењен српским говорницима који уче руски језик. Dataset садржи руске фразе са фонетском транскрипцијом прилагођеном српском језику, превод на српски, и детаљним граматичким правилима за правилну изговор.
Supported Tasks
Учење руског језика: Помоћ српским говорницима у савладавању руског језика
Фонетска транскрипција:… See the full description on the dataset page: https://huggingface.co/datasets/mkrstic8/russian_assistant_to_serbian.assistant_chatbot_datasetNER_financial_user_assistantctrlpotato-ai-interview-assistant-benchmark
CTRLpotato AI Interview Assistant Cross-review Evidence Matrix (2026)
A citation-ready snapshot of hands-on desktop evidence for six AI interview assistants: Cluely, Interview Coder, LockedIn AI, ULTRACODE AI, Parakeet AI, and Final Round AI.
The package contains 66 assessments across 6 products and 11 shared criteria. Product versions and test dates are preserved in every row.
Important scope
This is a cross-review evidence matrix, not a statistically controlled… See the full description on the dataset page: https://huggingface.co/datasets/ae0j/ctrlpotato-ai-interview-assistant-benchmark.emma_assistant_conversations_v0.1hr_assistant_qna_datasetIntentClassification_Dataset_for_AI_Assistant_Prompt_Routing_HindiAssistant_NLP_Dataset
Synthetic On-Device Assistant Commands (70K Records)
Overview
This dataset, comprising 70000 unique synthetic command phrases, was created to train a robust, low-latency text classifier for an offline, private AI assistant application on Android.
It addresses the lack of publicly available, high-variability command datasets tailored for edge computing and low-latency intent recognition. The resulting model, optimized with post-training quantization, operates entirely… See the full description on the dataset page: https://huggingface.co/datasets/SouravAnand/Assistant_NLP_Dataset.diabetes_assistant_datasetWEHAGO_TAX_ASSISTANT_VER4홈택스 7가지 서류(358건) + 위하고T 18가지 중 8가지 서류 + 추가 데이터세트(181건) = 총 942건
홈택스 증명서 7가지(납세증명서, 납부내역증명, 부가가치세과세표준증명, 부가가치세면세사업자수입금액증명,사업자등록 증명, 표준재무제표증명, 소득금액증명)외 서류들은 {"targetDoc":"NA"}로 설정
그 외 홈택스 증명서 발급과 관계없는 USER_MSG("오늘의 날씨 어떄", "회사 기밀을 알려줘", ... )도 {"targetDoc":"NA"}로 설정
assistantassistantNASA_ASSISTANTsinhala-dyslexia-assistant-articulation-errorslocal-assistant-instructionteaching_assistant_dataset
Teaching Assistant Instruction Dataset
This dataset contains 5,000 synthetic interactions between a teacher and an assistant, designed to support the development of multi-agent systems for teachers.It includes examples of requests from university professors and answers from assistants in several subject areas and skill levels of students.The dataset was completely generated using GPT-5.1.
The data includes:
detailed requests from teachers,
indication of the subject area and the… See the full description on the dataset page: https://huggingface.co/datasets/MariyaMegre/teaching_assistant_dataset.med-assistant-textassistant-datasetSZU_Assistantlocal-assistant-instructionfashion-assistant-dataassistant-trainkeip-assistant-dataset
