datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
🏆 View the live leaderboard → — interactive results across JEE Advanced, JEE Main & NEET, with open/closed-weight badges, contamination flags, and per-run cost.
A benchmark for evaluating vision-capable LLMs on Indian competitive exam questions (JEE Advanced & NEET). Each question is the original exam image; models answer via the OpenRouter API and are scored with authentic, exam-specific marking schemes — including partial credit for JEE… See the full description on the dataset page: https://huggingface.co/datasets/Reja1/jee-neet-benchmark.jeebench
JEEBench(EMNLP 2023)
Repository for the code and dataset for the paper: "Have LLMs Advanced Enough? A Harder Problem Solving Benchmark For Large Language Models" accepted in EMNLP 2023 as a Main conference paper.
https://aclanthology.org/2023.emnlp-main.468/
Citation
If you use our dataset in your research, please cite it using the following
@inproceedings{arora-etal-2023-llms,
title = "Have {LLM}s Advanced Enough? A Challenging Problem Solving Benchmark For Large… See the full description on the dataset page: https://huggingface.co/datasets/daman1209arora/jeebench.jee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-main-questions.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.jee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/soughed/jee-main-questions.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Vyshnavi93920/jee-neet-benchmark.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hellboi78688/jee-neet-benchmark.IFTjee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-advanced-questions.jee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-main-questions.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Parth1700/jee-neet-benchmark.Powerlifting
Dataset Card for Powerlifting
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Jeevak01/Powerlifting/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Jeevak01/Powerlifting.
