datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Home-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Vitinf/Home-Assistant-Requests-V2.GraphCode-Bench-500-v0
GraphCode-Bench-500-v0
GraphCode-Bench is a benchmark for evaluating LLMs on call-graph reasoning — given a function in a real-world repository, can a model identify which functions call it (upstream) or which functions it calls (downstream), across 1 and 2 hops?
Models are evaluated agentically: they receive read-only filesystem tools (list_directory, read_file, search_in_file) and up to 10 turns to explore the codebase before producing an answer.
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/VittorioRossi/GraphCode-Bench-500-v0.IMAD
Dataset Summary
This dataset contains data from the paper IMage Augmented multi-modal Dialogue: IMAD.
The main feature of this dataset is the novelty of the task. It has been generated specifically for the purpose of image interpretation in a dialogue context.
Some of the dialogue utterances have been replaced with images, allowing a generative model to be trained to restore the initial utterance.
The dialogues are sourced from multiple dialogue datasets (DailyDialog, Commonsense… See the full description on the dataset page: https://huggingface.co/datasets/VityaVitalich/IMAD.TwoSentenceHorror
