datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
🏆 View the live leaderboard → — interactive results across JEE Advanced, JEE Main & NEET, with open/closed-weight badges, contamination flags, and per-run cost.
A benchmark for evaluating vision-capable LLMs on Indian competitive exam questions (JEE Advanced & NEET). Each question is the original exam image; models answer via the OpenRouter API and are scored with authentic, exam-specific marking schemes — including partial credit for JEE… See the full description on the dataset page: https://huggingface.co/datasets/Reja1/jee-neet-benchmark.jeebench
JEEBench(EMNLP 2023)
Repository for the code and dataset for the paper: "Have LLMs Advanced Enough? A Harder Problem Solving Benchmark For Large Language Models" accepted in EMNLP 2023 as a Main conference paper.
https://aclanthology.org/2023.emnlp-main.468/
Citation
If you use our dataset in your research, please cite it using the following
@inproceedings{arora-etal-2023-llms,
title = "Have {LLM}s Advanced Enough? A Challenging Problem Solving Benchmark For Large… See the full description on the dataset page: https://huggingface.co/datasets/daman1209arora/jeebench.grafite-jee-mains-qna-no-imgJee-Chemistry-dataset-with-COTjee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-main-questions.jee_advanced_25k
JEE Reasoning Dataset v2 (25K)
Improvements:
Deduplicated questions
Dynamic reasoning traces
Topic-aligned metadata
Multiple problem families
Calculated intermediate arithmetic
No placeholder answers
Format:
{
"id": int,
"subject": "Physics or Mathematics",
"topic_tags": [...],
"difficulty_level": "JEE Advanced",
"question": "...",
"chain_of_thought_reasoning": "...",
"final_answer": "..."
}
IndianKanoon2025-Jee-Mains-Questionjee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.isl-isolated-40words
ISL Isolated Word Dataset (40 words)
Normalized isolated-sign video corpus for Indian Sign Language (ISL), built for Transformer / isolated SLR training.
642 H.264 MP4 clips
40 target glosses
Clips resized to height 480, ~30 FPS
Full per-clip provenance in metadata.csv
Vocabulary
hello, goodbye, thank you, sorry, please, yes, no, help, stop, okay, me, you, he, she, mother, father, brother, sister, friend, teacher, student, home, school, hospital, market, eat… See the full description on the dataset page: https://huggingface.co/datasets/Jeesitaj/isl-isolated-40words.UCINet0_PUCCH-Format-0-ML
UCINet0: PUCCH-Format-0-ML
This repository provides scripts and tools to generate, combine, train, and test PUCCH Format 0 datasets using MATLAB and Python.The workflow covers end‑to‑end signal generation, UCINet0 training, evaluation, and real‑world testing.
Overview
Two main stages:
Dataset Generation (MATLAB) — Create PUCCH Format 0 frequency‑domain complex sample datasets.
Model Training & Testing (Python) — Train and evaluate a neural network using the generated… See the full description on the dataset page: https://huggingface.co/datasets/JeevaKeshav/UCINet0_PUCCH-Format-0-ML.jee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/soughed/jee-main-questions.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Vyshnavi93920/jee-neet-benchmark.Jee-Parallel-Dataset-Indicjee-exam-qnajee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hellboi78688/jee-neet-benchmark.IFTjeebenchJEEM
Dataset Card for JEEM
📊 Curated by: Toloka, MBZUAI
🌐 Language(s): Modern Standard Arabic, 🇯🇴 Jordanian dialect, 🇪🇬 Egyptian dialect, 🇦🇪 Emirati dialect, 🇲🇦 Moroccan dialect
🔍 Dataset Description
JEEM is a benchmark dataset designed to evaluate Vision-Language Models (VLMs) in the context of Arabic dialectal diversity. It includes one representative dialect from each major Arabic dialectal region: 🇯🇴 Jordanian (Levantine), 🇪🇬 Egyptian, 🇦🇪 Emirati… See the full description on the dataset page: https://huggingface.co/datasets/toloka/JEEM.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-advanced-questions.wikidoc-healthassistexamguru-neet-jee-datasetjeejee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-main-questions.jee_mathVisualMRCjee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Parth1700/jee-neet-benchmark.2025-Jee-Mains-Questionjee-grpo-v1
JEE-GRPO-v1: Visual Reasoning Dataset for VLMs
Dataset Summary
JEE-GRPO is a high-quality, multimodal dataset designed for training and benchmarking Vision Language Models (VLMs) on complex STEM problems.
Unlike traditional text-only datasets, this dataset renders JEE Main (Joint Entrance Examination) questions as high-resolution images. This approach bypasses OCR errors and perfectly preserves complex LaTeX equations, diagrams, and chemical structures, making it ideal… See the full description on the dataset page: https://huggingface.co/datasets/farhananis005/jee-grpo-v1.agentfaildb
AgentFailDB
An open, fully-local benchmark of failure modes in multi-agent LLM systems —
750 execution traces from three frameworks (CrewAI, AutoGen,
LangGraph) run on a local Llama 3.1 8B model (via Ollama) at $0 API cost.
Relation to prior work
The canonical, human-validated taxonomy of multi-agent LLM failures is MAST —
Cemri et al., "Why Do Multi-Agent LLM Systems Fail?" (NeurIPS 2025). AgentFailDB does
not supersede or replace it; it is a fully-local… See the full description on the dataset page: https://huggingface.co/datasets/Jeevaa79/agentfaildb.
