datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Landscape-DataThis repository contains the dataset used in the paper Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models.
Code: https://github.com/tmlr-group/landscape-of-thoughts
This project is licensed under the MIT License. See the LICENSE.md file for details.
gaze_dataset_full
gaze_dataset_full — StreamGaze_v2 + EgoGazeVQA + HD-EPIC
A single repository containing three complementary benchmarks for
evaluating multimodal LLMs on gaze-grounded egocentric video question
answering:
Subfolder
Source
Held-out (val/test)
Questions
StreamGaze_v2/
egoexolearn, holoassist, egtea
egtea
8 MCQ tasks (4-opt) — gaze-conditioned past/present/future
EgoGazeVQA/
ego4d, egoexo, egtea
egtea
causal / spatial / temporal (5-opt)
HD-EPIC/
P01–P09
P09… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/gaze_dataset_full.medical_instruction_tuning
