datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BrainTRACE
BrainTRACE — Brain MRI Tracking, Reasoning, Annotation & Comparison Evaluation
A vision-language benchmark of 6,923 task definitions (7,273 scored VQA instances) over the upstream MR-RATE longitudinal brain MRI dataset.
⚠️ What BrainTRACE redistributes (and what it does not). BrainTRACE is not a re-publication of MR-RATE. The contributions released here are the task definitions — questions, ground-truth values, multi-slot rubrics, per-step chain rubrics, and per-item pointers to… See the full description on the dataset page: https://huggingface.co/datasets/BrainTRACE-anon/BrainTRACE.MedCortex-v1
MedCortex — Bilingual Medical Reasoning + Consultation Corpus
86,006 provenance-tracked, decontaminated medical examples in one uniform schema, fusing two
complementary strengths without flattening either:
task_type
Rows
What it is
Why it is here
reasoning
44,736
Verified chain-of-thought (KG-grounded + verified CoT), English
Drives exam-style benchmark reasoning — the source of the MedReason paper's measured gains
consultation
41,270
Bilingual (EN/FR) clinical Q&A… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedCortex-v1.livenewsbench-search-arms
LiveNewsBench Search Arms
Paired measurements of four language models answering the same 1,329 news
questions under four retrieval conditions. Every question was run in every
condition, so each row pairs with 13 others on task_key.
The release answers one question: how much of an agent's answer quality comes
from the model, and how much from the search system wrapped around it.
This is a derivative evaluation-results dataset, not the original
LiveNewsBench benchmark. The… See the full description on the dataset page: https://huggingface.co/datasets/BraintrustDataDev/livenewsbench-search-arms.
