datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
🏆 View the live leaderboard → — interactive results across JEE Advanced, JEE Main & NEET, with open/closed-weight badges, contamination flags, and per-run cost.
A benchmark for evaluating vision-capable LLMs on Indian competitive exam questions (JEE Advanced & NEET). Each question is the original exam image; models answer via the OpenRouter API and are scored with authentic, exam-specific marking schemes — including partial credit for JEE… See the full description on the dataset page: https://huggingface.co/datasets/Reja1/jee-neet-benchmark.Act2Cap_benchmarkCollected data from GUI-Action-Narrator
data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.Multimodal-Robustness-BenchmarkSkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Vyshnavi93920/jee-neet-benchmark.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hellboi78688/jee-neet-benchmark.miroeval-benchmark-2026
MiroEval Benchmark 2026
Description
MiroEval Benchmark 2026 is a benchmark for evaluating deep research agents on long-form research tasks. It contains 100 tasks, including 70 text-only tasks and 30 multimodal tasks with accompanying attachments such as PDFs, documents, images, and structured files.
The benchmark is designed to evaluate three complementary aspects of deep research systems:
Synthesis Quality: whether the final report is comprehensive, insightful… See the full description on the dataset page: https://huggingface.co/datasets/anon-ed2026/miroeval-benchmark-2026.Bones_and_Joints_Benchmark
The Bones and Joints Benchmark
Dataset Overview
We have developed a specialized dataset focused on musculoskeletal disorders, designed to systematically evaluate the clinical capabilities of visual language models (VLMs). The evaluation covers knowledge recall, clinical note interpretation, radiology image interpretation, diagnosis generation and rationale, treatment planning and rationale. The dataset primarily includes multiple-choice questions and open-ended… See the full description on the dataset page: https://huggingface.co/datasets/PUTH2025/Bones_and_Joints_Benchmark.Fraud-R1-LLM-Defense-Fraud-Benchmark
Fraud-R1 : A Comprehensive Benchmark for Assessing LLM Robustness Against Fraud and Phishing Inducement
Shu Yang*, Shenzhe Zhu*, Zeyu Wu, Keyu Wang, Junchi Yao, Junchao Wu, Lijie Hu, Mengdi Li, Derek F. Wong, Di Wang†
(*Contribute equally, †Corresponding author)
😃 Github | 📜 Project Page | 📝 arxiv
❗️Content Warning: This repo contains examples of harmful language.
📰 News
2025/02/16: ❗️We have released our evaluation code.
2025/02/16: ❗️We have released our dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Chouoftears/Fraud-R1-LLM-Defense-Fraud-Benchmark.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Parth1700/jee-neet-benchmark.hungarian-riddles-benchmark
Hungarian Riddles Benchmark
Overview
This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles.
The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency.
Dataset structure
Each row contains one riddle with reference material for evaluation.
Fields
ID – unique identifier
topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.
