datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Q20LLM
20 Questions game with LLM
This dataset generated with the following LLMs:
Groq API / llama3-70b-8192
Groq API / mixtral-8x7b-32768
Mistral API / mistral-large-latest
Test keywords based on newlist_things.rmdup.test.txt from Entity-Deduction Arena (EDA) project.
The dataset generated in two stages:
LLM was prompt to generate keywords
Dialog with different length generated for each keyword
Keywords prompt
Generate a list of 500 diverse and simple keywords suitable… See the full description on the dataset page: https://huggingface.co/datasets/cvmistralparis/Q20LLM.AVRT-20K
AVRT-20K: Audio-Visual Reasoning Traces
AVRT-20K is a dataset of audio-visual reasoning traces generated through the AVRT (Audio-Visual Reasoning Transfer) pipeline. It provides structured chain-of-thought reasoning that explicitly integrates audio and visual evidence for answering multiple-choice questions about video content.
This dataset accompanies the paper:
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
Edson Araujo, Saurabhchand Bhati, M. Jehanzeb… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/AVRT-20K.
