datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVQA
Summary | 摘要
This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys.
The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds).
Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.valor32k-avqa-v2
Valor32k-AVQA v2.0
Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position.
Links
Paper: ACM Digital Library
Project page: inesriahi.github.io/valor32k-avqa-2
Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.openclaw-recursive-study-data
OpenClaw Recursive Repository Study Data
Synthetic repository-study data generated against
openclaw/openclaw at commit
da228660306b55a9cce3b973946f3aacfc515848. The source repository is MIT licensed.
This release contains exploration questions, tool-using study trajectories,
recursive notes, full recall-rewritten trajectories, and recall-to-action
training examples. Nested chat/tool objects are stored as JSON strings to keep
the schema stable and can be decoded with json.loads.… See the full description on the dataset page: https://huggingface.co/datasets/aviralku/openclaw-recursive-study-data.reasoning-spectrum-qa
Reasoning Spectrum QA Dataset
This is a unified dataset of exactly 1,000 examples designed for evaluating reasoning capabilities across different dimensions (factual, commonsense, science, arithmetic, multi-hop, and extractive span).
Dataset Purpose
The dataset brings together questions of varying difficulty and reasoning family types under a single unified schema to facilitate standardized testing and evaluation of large language models.
Source… See the full description on the dataset page: https://huggingface.co/datasets/avreymi/reasoning-spectrum-qa.code-aviation-civile
Code de l'aviation civile, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-aviation-civile.aura_qa
Affect-Uniform ReAding QA (AURA-QA),
This dataset contains short passages from English texts found in Project Gutenberg paired with question–answer examples and emotion labels. The dataset is designed to support research in emotion-aware reading comprehension. Answers are constrained to 1–3 tokens and are generated and verified by large language models.
Dataset Structure
text — Passage excerpt
question — Question about the passage
answer — Short answer (1–3 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/avalab/aura_qa.AVRT-20K
AVRT-20K: Audio-Visual Reasoning Traces
AVRT-20K is a dataset of audio-visual reasoning traces generated through the AVRT (Audio-Visual Reasoning Transfer) pipeline. It provides structured chain-of-thought reasoning that explicitly integrates audio and visual evidence for answering multiple-choice questions about video content.
This dataset accompanies the paper:
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
Edson Araujo, Saurabhchand Bhati, M. Jehanzeb… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/AVRT-20K.
