datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EgoDyn-Bench-ECCV2026
EgoDyn-Bench
A physics-grounded VQA benchmark for evaluating Vision-Language Models on trajectory-based dynamics reasoning in autonomous driving.
Project page | Paper | GitHub
This repository contains the data artifacts for the benchmark. The evaluation harness, baselines, and reference implementations live in the companion GitHub repository.
Note on licensing. The nuScenes-derived portion of this dataset is released under CC BY-NC-SA 4.0 to comply with nuScenes' upstream… See the full description on the dataset page: https://huggingface.co/datasets/TUM-AVS/EgoDyn-Bench-ECCV2026.mentalkg
mentalkg
Bilingual (English / German) synthetic corpus of first-person mental-health journal entries paired with structured knowledge graphs. 47,714 samples across 3,410 participants; 41,315 accepted after LLM-based verification.
Files
file
rows
what it is
dataset.jsonl
47,714
full corpus (accept + review + reject), one JSON object per participant-day
graphs.jsonl
48,104
intermediate graphs the pipeline drew from (generator input pool)… See the full description on the dataset page: https://huggingface.co/datasets/CHI-TUM/mentalkg.Protein2Text-QA
Protein2Text-QA Dataset
The Protein2Text-QA dataset is designed to generate human-readable explanations for protein functions based on protein sequences. It consists of question-answer (QA) pairs generated from PubMed Central (PMC) articles using LLaMA3.1-8B-Instruct. The dataset is structured into different subsets tailored for pretraining, fine-tuning, and evaluation.
Dataset Overview
Size: ~210,000 QA pairs
Source: UniProt (pretraining), PubMed Central (PMC) (QA… See the full description on the dataset page: https://huggingface.co/datasets/tumorailab/Protein2Text-QA.synthetic-patient-dr-data
Synthetic Patient DR Data
Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio.
Dataset Summary
This dataset was generated for research and prototyping in:
clinical dialogue generation
structured clinical extraction
text-to-audio workflows
conversational healthcare modeling
All consultations are synthetic and should not be treated as real clinical encounters.
Export Metadata
Mode: audio
Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.IDPRO-QA
IDPRO-QA — Multi-stage protein-function QA corpus
Instruction-tuning corpus used to train the IDPro (Illuminating the Dark
PROteome) model. Generated from UniProt + InterPro + M-CSA + Prosite annotations
over bacterial proteins, organised into four progressively harder curricula
(stages 1 → 4).
Stages
Stage
Records
Description
stage1
669,698
Single-feature QA: one domain / binding site / region per record.
stage2
30,552
Multi-feature, multi-turn QA combining… See the full description on the dataset page: https://huggingface.co/datasets/tumorailab/IDPRO-QA.Tumbuka_Continuous_Next-Token_Prediction_Datasettumours_chinesebrain-tumor-llm-dataset1tum-heilbronn-mist26-primary
