CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acon96 /Home-Assistant-Requests-V2 Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests-V2.textquestion-answering100K<n<1M20 likes474 downloads9mo agoHugging Face02Isotonic /human_assistant_conversationtexttext-generation1M<n<10M23 likes325 downloads3y agoHugging Face03acon96 /Home-Assistant-Requests Home Assistant Requests Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The dataset is generated from the different CSV "piles". The "piles" contain different chunks of requests that are assembled into a final context that is presented to the LLM. For example, piles/pile_of_device_names.csv contains only names of various devices to be used as part of context as well as… See the full description on the dataset page: https://huggingface.co/datasets/acon96/Home-Assistant-Requests.textquestion-answering10K<n<100K47 likes284 downloads3y agoHugging Face04massines3a /assistant-axis-vectors Assistant Axis Vectors for gemma-3-27b-it This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it. Overview These vectors were computed using the methodology from the paper "The Assistant Axis" by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the "assistant-like" to "role-playing" spectrum. Contents gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.texttext-generationn<1K0 likes178 downloads8mo agoHugging Face05DaftP /Home-Assistant-requests-for-intent-detection-and-function-recognition Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.textquestion-answering100K<n<1M1 likes104 downloads5mo agoHugging Face06Lucasllfs /router-assistant-tool-calling-en-es Router Assistant Tool Calling EN-ES Synthetic English and Spanish conversations for supervised fine-tuning of a small, local router assistant. The assistant answers brief social turns, obtains current network facts through tools, handles tool failures, and asks for confirmation before restarting the router or disabling WAN internet access. Dataset size Split Conversations Assistant completions Train 11,066 21,242 Validation 984 1,890 Test 926 1,769… See the full description on the dataset page: https://huggingface.co/datasets/Lucasllfs/router-assistant-tool-calling-en-es.texttext-generation10K<n<100K0 likes84 downloads1mo agoHugging Face07Isotonic /human_assistant_conversation_deduped Deduplicated version of Isotonic/human_assistant_conversation Deduped with max jaccard similarity of 0.75 texttext-generation100K<n<1M13 likes83 downloads3y agoHugging Face08m-ric /Open_Assistant_Chains_German_Translation Dataset Card for Dataset Name Dataset description This dataset is derived from OpenAssistant Conversation Chains, which is a reformatting of OpenAssistant Conversations (OASST1), which is itself a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus is a product of a worldwide… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/Open_Assistant_Chains_German_Translation.texttext-generation10K<n<100K1 likes82 downloads3y agoHugging Face09alireza-fallah /bourse-assistant-sft Bourse Assistant SFT Dataset این مخزن داده‌ها را برای پروژه دست‌یار بورس [لینک پروژه] منتشر می‌کند. مجموعه‌داده‌ای برای فاین‌تیون نظارت‌شده یک دستیار هوش مصنوعی در حوزه بورس اوراق بهادار ایران. هر نمونه شامل مجموعه‌ای از خبرهای مربوط به یک نماد بورسی مشخص در یک تاریخ معین است، و مدل باید روند قیمت (مثبت/منفی) را همراه با توضیح تخمین بزند. محتوای مجموعه‌داده (Dataset Summary) داده به دو زیرمجموعه تقسیم شده، بر اساس تعداد نمادهای بورسی پوشش داده‌شده: Config… See the full description on the dataset page: https://huggingface.co/datasets/alireza-fallah/bourse-assistant-sft.documenttext-generation10K<n<100K1 likes65 downloads1d agoHugging Face10arcada-labs /assistant-bench Assistant Bench 31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a personal assistant handling flights, email, calendar, and reminders. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a personal assistant managing flight bookings, email composition, calendar events, and reminders. Turns include dual… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/assistant-bench.audioautomatic-speech-recognitionn<1K2 likes63 downloads6mo agoHugging Face11squeezebits /synthesized-coding-assistant-dataset Synthesized Coding Assistant Dataset Overview Coding assistants are increasingly used for real-world software engineering workflows. However, there are relatively few datasets that closely resemble how such assistants operate in practice. Many existing coding datasets are based on single-turn or single-iteration tasks, where a model receives one coding request and directly produces an answer or patch. In contrast, practical coding assistants often work through… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/synthesized-coding-assistant-dataset.tabulartext-generationn<1K0 likes62 downloads4mo agoHugging Face12tuxevil /Home-Assistant-Requests-V5.2-Native-Strict Home Assistant Requests V5.2 Native Strict Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model. Contract: ha-native-tool-calling-v2. Frozen snapshot Split Rows Direct speech Multi-call Maximum rendered tokens train 3,806 340 78 3,098 validation 530 52 4 2,874 test 633 102 22 2,925 Tokenizer audit: model: unsloth/Qwen3-4B-Instruct-2507 revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.texttext-generation1K<n<10K0 likes53 downloads2mo agoHugging Face13Sellopale /Assistantz8 OpenAssistant Conversations Dataset (OASST1) Dataset Summary In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations (OASST1), a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus is a product of a worldwide crowd-sourcing effort… See the full description on the dataset page: https://huggingface.co/datasets/Sellopale/Assistantz8.tabulartranslation10K<n<100K0 likes51 downloads9mo agoHugging Face14bosaj /eniad-assistant-instruct-dataset 📚 ENIAD Academic & Enterprise Instruction Dataset 🤝 Curated by the ENIAD AI Engineering Team (May 2025) A curated, bilingual (French 🇫🇷 and English 🇬🇧) instruction-tuning dataset designed for training institutional AI assistants in Moroccan higher education. 👥 Engineering Team Abdellah ENNAJARI (@abdennajari • GitHub @ennajari) Ahmed OUKACHA (@ahmed-ouka) Oussama EL HADJI (HF @bosaj • GitHub @Bosaj) Abdelilah OURTI (@abdelilahou)… See the full description on the dataset page: https://huggingface.co/datasets/bosaj/eniad-assistant-instruct-dataset.textquestion-answeringn<1K0 likes48 downloads6d agoHugging Face15abdullah693 /adaption-sehat-saathi-lhw-assistant-v1 This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-sehat-saathi-lhw-assistant-v1 This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI and related national protocols. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-sehat-saathi-lhw-assistant-v1.texttext-generation10K<n<100K1 likes46 downloads3mo agoHugging Face16dzur658 /ping-technical-assistant-small Ping Technical Assistant Dataset Small This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine tuning. How to Utilize this Dataset In theory this dataset should work properly with… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-small.texttext-generation1K<n<10K0 likes44 downloads7mo agoHugging Face17m-ric /Open_Assistant_Conversation_Chains Dataset Card for Dataset Name Dataset description This dataset is a reformatting of OpenAssistant Conversations (OASST1), which is a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus is a product of a worldwide crowd-sourcing effort involving over 13,500 volunteers. It was modified… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/Open_Assistant_Conversation_Chains.texttext-generation10K<n<100K6 likes42 downloads3y agoHugging Face18alireza-fallah /bourse-assistant-rl Bourse Assistant RL Dataset این مخزن داده‌ها را برای پروژه دست‌یار بورس [لینک پروژه] منتشر می‌کند. مجموعه‌داده‌ای برای فاین‌تیون با روش یادگیری تقویتی یک دستیار هوش مصنوعی در حوزه بورس اوراق بهادار ایران. هر نمونه شامل مجموعه‌ای از خبرهای مربوط به یک نماد بورسی مشخص در یک تاریخ معین است، و مدل باید روند قیمت (مثبت/منفی) را همراه با توضیح تخمین بزند. محتوای مجموعه‌داده (Dataset Summary) داده به دو زیرمجموعه تقسیم شده، بر اساس تعداد نمادهای بورسی پوشش داده‌شده:… See the full description on the dataset page: https://huggingface.co/datasets/alireza-fallah/bourse-assistant-rl.documenttext-generation1K<n<10K1 likes41 downloads1d agoHugging Face19Hy9n0t1c /Home-Assistant-Requests-V2 Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hy9n0t1c/Home-Assistant-Requests-V2.textquestion-answering100K<n<1M0 likes37 downloads6mo agoHugging Face20Vitinf /Home-Assistant-Requests-V2 Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Vitinf/Home-Assistant-Requests-V2.textquestion-answering100K<n<1M0 likes35 downloads9mo agoHugging Face21P0u4a /msm-ai-assistant-philosophy-spec AI assistant philosophy spec Complete identity-decontaminated MSM corpus: 13,201 documents. Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.texttext-generation10K<n<100K0 likes33 downloads14d agoHugging Face22CircularBalls /tt633-technical-code-assistant-v1 TT633 Technical Code Assistant v1 This dataset is built for training the fresh custom TransformerTechnology V8.3 MDL Circle-Switch-Grid model as a small technical/code assistant. Canonical training column: text. Format: Instruction: ... Input: ... Answer: ... <END> Primary sources: Plaincode CNL rows from CircularBalls/plaincode-cnl-100k. Small curated technical QA, code-generation, debugging, reasoning, and stop-discipline seed rows. Optional local pack text if provided at… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt633-technical-code-assistant-v1.texttext-generation10K<n<100K0 likes32 downloads4mo agoHugging Face23tuxevil /Home-Assistant-Requests-V5.1-Native-Strict Home Assistant Requests V5.1 Native Strict Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model. Contract: ha-native-tool-calling-v2. Frozen snapshot Split Rows Direct speech Multi-call Maximum rendered tokens train 3,806 340 78 3,098 validation 530 52 4 2,874 test 633 102 22 2,925 Tokenizer audit: model: unsloth/Qwen3-4B-Instruct-2507 revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.1-Native-Strict.texttext-generation1K<n<10K0 likes32 downloads2mo agoHugging Face24Aadithya18 /Health_Coach_Assistant_Data Health Coach Assistant Dataset This dataset consists of data for different types of chat with a health coach assistant about setting or updating the goals for walking. Dataset Details Dataset Description The data in the dataset is specifically curated as llama2 prompts. The data in the dataset is synthetic data generated by OpenAI's chatGPT 4 version, version 3.5, Github Copilot, and Claude AI. Curated by: Sai Sangameswara Aadithya Kanduri Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aadithya18/Health_Coach_Assistant_Data.texttext-generation1K<n<10K2 likes31 downloads2y agoHugging Face25mwitiderrick /glaive-code-assistant Glaive Code Assistant Glaive Code Assistant dataset formatted for training assistant models with the following prompt template: <s>[INST] {question} [/INST] {answer} </s> Trained model can be prompted in Llama style: <s>[INST] {{ user_msg }} [/INST] texttext-generation100K<n<1M1 likes30 downloads3y agoHugging Face26bubun123 /smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management 🤏 smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model bubun123/smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management. 📦 Asset Details Origin: Smolify Foundry (Job ID: 13675e9f) Records: 926 Type: Synthetic Instruction Tuning Data ⚖️… See the full description on the dataset page: https://huggingface.co/datasets/bubun123/smolified-medimind-personalized-medication-adherence-assistant-for-chronic-disease-management.texttext-generationn<1K0 likes29 downloads7mo agoHugging Face27PJMixers-Dev /glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT glaiveai/glaive-code-assistant with responses regenerated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generationn<1K0 likes26 downloads2y agoHugging Face28gieljnssns /Home-Assistant-Requests-V2 Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/gieljnssns/Home-Assistant-Requests-V2.textquestion-answering100K<n<1M1 likes26 downloads7mo agoHugging Face29mkrstic8 /russian_assistant_to_serbian Dataset Card for Russian to Serbian Assistant Dataset Description Dataset Summary Russian to Serbian Assistant је едукативни dataset намењен српским говорницима који уче руски језик. Dataset садржи руске фразе са фонетском транскрипцијом прилагођеном српском језику, превод на српски, и детаљним граматичким правилима за правилну изговор. Supported Tasks Учење руског језика: Помоћ српским говорницима у савладавању руског језика Фонетска транскрипција:… See the full description on the dataset page: https://huggingface.co/datasets/mkrstic8/russian_assistant_to_serbian.texttranslationn<1K0 likes25 downloads8mo agoHugging Face30Debarun12 /aws-enterprise-assistant-dataset AWS Enterprise Assistant Dataset Instruction-following Q&A dataset generated from official AWS documentation. Built to fine-tune domain-specific AI assistants on AWS cloud services. Dataset Description Dataset Summary This dataset contains 407 instruction-following Q&A pairs generated from 211 text chunks scraped from official AWS documentation across 8 core services. Each pair consists of a question a cloud practitioner would ask, and a detailed… See the full description on the dataset page: https://huggingface.co/datasets/Debarun12/aws-enterprise-assistant-dataset.textquestion-answeringn<1K0 likes25 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.