datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
efficient-llm-papers
Efficient LLM Papers — FineSet
A research-paper dataset on Efficient LLM Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Efficient LLM Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/efficient-llm-papers.efficient_llm
Data V4 for NeurIPS LLM Challenge
Contains 70949 samples collected from Huggingface:
Math: 1273
gsm8k
math_qa
math-eval/TAL-SCQ5K
TAL-SCQ5K-EN
meta-math/MetaMathQA
TIGER-Lab/MathInstruct
Science: 42513
lighteval/mmlu - 'all', "split": 'auxiliary_train'
lighteval/bbq_helm - 'all'
openbookqa - 'main'
ComplexQA: 2940
ARC-Challenge
ARC-Easy
piqa
social_i_qa
Muennighoff/babi
Rowan/hellaswag
ComplexQA1: 2060
medmcqa
winogrande_xl,
winogrande_debiased
boolq
sciq
CNN: 2787… See the full description on the dataset page: https://huggingface.co/datasets/transZ/efficient_llm.
