datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.GameplayQA
GameplayQA: A Decision-Dense POV-Synced Multi-Video
Understanding Benchmark of 3D Virtual Agents
Yunzhe Wang
Runhui Xu
Kexin Zheng
Tianyi Zhang
Jayavibhav N. Kogundi
Soham Hans
Volkan Ustun
University of Southern California
ACL 2026
Corresponding Author: yunzhewa@usc.edu
Overview
GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.FRED-SemBench
FRED-SemBench
FRED-SemBench is a 200-question benchmark candidate for evaluating whether LLM
agents retrieve macroeconomic answers with the intended concept, series,
transformation, unit, observation period, and data-vintage semantics.
This dataset accompanies the FinNLP 2026 paper
“FRED-SemBench: Evaluating Semantic Reliability in LLM Access to
Macroeconomic Data”
by Wilson Wang, Chandler Han, and Peter Zhang (Kairos-AI).
Status and scope
50 independently… See the full description on the dataset page: https://huggingface.co/datasets/wangjinh/FRED-SemBench.IPEval
Dataset Card for IPEval
IPEval is a pioneering bilingual Intellectual Property (IP) agency consultation evaluation benchmark, meticulously crafted to assess the competencies of Large Language Models (LLMs) in the intricate domain of intellectual property. This benchmark is the first of its kind, encompassing a diverse spectrum of 2,657 multiple-choice questions that are intricately divided across four major capability dimensions: creation, application, protection, and management.… See the full description on the dataset page: https://huggingface.co/datasets/QiYao-Wang/IPEval.NavQA_Revised
NavQA Revised
NavQA Revised is a re-annotated version of the NaVQA dataset released with NVIDIA ReMEmbR. The original annotations are distributed in remembr/data/navqa/data.csv.
This repository provides the revised annotations as JSONL files:
navqa.jsonl
navqa_over_sequence.jsonl
Each line is one question-answer example. navqa.jsonl is the primary revised annotation file. navqa_over_sequence.jsonl uses the same schema and examples, but uses the beginning of the full sequence as the… See the full description on the dataset page: https://huggingface.co/datasets/Qiuchen-Wang/NavQA_Revised.wangchanx-seed-free-synthetic-instruct-thai-120k
Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k
Dataset Summary
This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.Futurex-Past
FutureX-Past
📜 Overview
This repository contains a dataset of past questions from the FutureX benchmark.
FutureX is a live, dynamic benchmark designed to evaluate the future prediction capabilities of Large Language Model (LLM) agents. It features a fully automated pipeline that generates new questions about upcoming real-world events, deploys agents to predict their outcomes, and scores the results automatically. For more information on the live benchmark, please refer… See the full description on the dataset page: https://huggingface.co/datasets/wanwan1212/Futurex-Past.
