datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
🏆 View the live leaderboard → — interactive results across JEE Advanced, JEE Main & NEET, with open/closed-weight badges, contamination flags, and per-run cost.
A benchmark for evaluating vision-capable LLMs on Indian competitive exam questions (JEE Advanced & NEET). Each question is the original exam image; models answer via the OpenRouter API and are scored with authentic, exam-specific marking schemes — including partial credit for JEE… See the full description on the dataset page: https://huggingface.co/datasets/Reja1/jee-neet-benchmark.jeebench
JEEBench(EMNLP 2023)
Repository for the code and dataset for the paper: "Have LLMs Advanced Enough? A Harder Problem Solving Benchmark For Large Language Models" accepted in EMNLP 2023 as a Main conference paper.
https://aclanthology.org/2023.emnlp-main.468/
Citation
If you use our dataset in your research, please cite it using the following
@inproceedings{arora-etal-2023-llms,
title = "Have {LLM}s Advanced Enough? A Challenging Problem Solving Benchmark For Large… See the full description on the dataset page: https://huggingface.co/datasets/daman1209arora/jeebench.grafite-jee-mains-qna-no-imgJee-Chemistry-dataset-with-COTjee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-main-questions.so101_gray_block_midvar_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 100,
"total_frames": 39280,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeevesh2009/so101_gray_block_midvar_test.jee_advanced_25k
JEE Reasoning Dataset v2 (25K)
Improvements:
Deduplicated questions
Dynamic reasoning traces
Topic-aligned metadata
Multiple problem families
Calculated intermediate arithmetic
No placeholder answers
Format:
{
"id": int,
"subject": "Physics or Mathematics",
"topic_tags": [...],
"difficulty_level": "JEE Advanced",
"question": "...",
"chain_of_thought_reasoning": "...",
"final_answer": "..."
}
so101_gray_block_pickup_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 50,
"total_frames": 11958,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeevesh2009/so101_gray_block_pickup_test.IndianKanoon2025-Jee-Mains-Questionjee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.isl-isolated-40words
ISL Isolated Word Dataset (40 words)
Normalized isolated-sign video corpus for Indian Sign Language (ISL), built for Transformer / isolated SLR training.
642 H.264 MP4 clips
40 target glosses
Clips resized to height 480, ~30 FPS
Full per-clip provenance in metadata.csv
Vocabulary
hello, goodbye, thank you, sorry, please, yes, no, help, stop, okay, me, you, he, she, mother, father, brother, sister, friend, teacher, student, home, school, hospital, market, eat… See the full description on the dataset page: https://huggingface.co/datasets/Jeesitaj/isl-isolated-40words.UCINet0_PUCCH-Format-0-ML
UCINet0: PUCCH-Format-0-ML
This repository provides scripts and tools to generate, combine, train, and test PUCCH Format 0 datasets using MATLAB and Python.The workflow covers end‑to‑end signal generation, UCINet0 training, evaluation, and real‑world testing.
Overview
Two main stages:
Dataset Generation (MATLAB) — Create PUCCH Format 0 frequency‑domain complex sample datasets.
Model Training & Testing (Python) — Train and evaluate a neural network using the generated… See the full description on the dataset page: https://huggingface.co/datasets/JeevaKeshav/UCINet0_PUCCH-Format-0-ML.svd-safety-gradlossjee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/soughed/jee-main-questions.so101_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 294,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeevesh2009/so101_test.Jee-Conversational-Dataset-Indicjee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Vyshnavi93920/jee-neet-benchmark.so101_gray_block_lowvar_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 50,
"total_frames": 19795,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeevesh2009/so101_gray_block_lowvar_test.ASR-Dataset-CollectionJee-Parallel-Dataset-IndicJEE-Main-2025-Math
JEE Mains 2025 Math Evaluation Set
🧾 Dataset Summary
This dataset contains 475 math questions from the official JEE Mains 2025 examination, covering both January and April sessions. It is curated to benchmark mathematical reasoning models under high-stakes exam conditions.
🚀 How to Load the Dataset
You can load the evaluation data using the datasets library from Hugging Face:
from datasets import load_dataset
# Load January session evaluation set
jan_data =… See the full description on the dataset page: https://huggingface.co/datasets/PhysicsWallahAI/JEE-Main-2025-Math.so101_new_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 294,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeevesh2009/so101_new_test.so101_dummy1_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 1,
"total_frames": 853,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeevesh2009/so101_dummy1_test.jee-exam-qnajee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hellboi78688/jee-neet-benchmark.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
Dataset Description
This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations:
JEE (Main & Advanced): Joint Entrance Examination for engineering.
NEET: National Eligibility cum Entrance Test for medical fields.
The questions are presented in image format (.png) as they appear in the original papers. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/RJTR001/jee-neet-benchmark.safety-quant-phase1
Safety-Aware Configuration-Conditioned LoRA — Phase 1 transfer study
What happens when a LoRA adapter is trained under one quantization configuration and
deployed under a different one. 36 evaluated cells: 6 adapters (5 + none) ×
5 base configurations, each on 1319 GSM8K problems and 1733 safety prompts.
Adapters: Jeesup/Llama-3.2-1B-Instruct-safetyquant-lora ·
Phase 0 baseline: Jeesup/safety-quant-phase0
U(a→b) — GSM8K exact match
train ↓ / deploy →
q16
q8… See the full description on the dataset page: https://huggingface.co/datasets/Jeesup/safety-quant-phase1.glue-lora-bitwidth-results
GLUE LoRA x Backbone Bit-Width — combined results
Aggregated metrics, LoRA geometry analysis and figures for a controlled study of
whether backbone bit-width (bf16 / int8 / nf4) changes what a LoRA adapter
learns.
Grid: 2 model sizes (1B, 3B) x 4 GLUE tasks (MNLI, QQP, SST-2,
RTE) x 3 bit-widths x 3 seeds = 72 runs (0 present here).
For each (model size, seed) the adapter initialisation is identical across
bit-widths, so cross-bit differences are attributable to the backbone.… See the full description on the dataset page: https://huggingface.co/datasets/Jeesup/glue-lora-bitwidth-results.IFT
