datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-scifineweb-edu-sci-superscifig-bench
SciFig-Bench: Scientific Figure & Multi-Scale Caption Dataset Card
Dataset Description
This dataset contains high-quality scientific figures, architecture diagrams, charts, and visualizations extracted from scientific papers (arXiv & local PDFs). It includes multi-level captions (Small, Medium, Large), extracted embedded OCR text, paper metadata, and alignment quality scores.
Total Figure Records: 589
Average CLIP Alignment Score: 0.2721
Average Composite Score:… See the full description on the dataset page: https://huggingface.co/datasets/Goutam112/scifig-bench.Scifi4TopicModel
Dataset Overview
This repository contains benchmark datasets for evaluating Large Language Model (LLM)-based topic discovery methods and comparing them against traditional topic models. These datasets provide a valuable resource for researchers studying topic modeling and LLM capabilities in this domain. The work is described in the following paper: Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of LLMs. Original data… See the full description on the dataset page: https://huggingface.co/datasets/zli12321/Scifi4TopicModel.dclm-dedup-25B-ai-scifi-docspdf_sci_qa_unverified_filtered_popular_urlspdf_sci_questions_difficulty_filter_and_science_question_labelerqwen2-5_sci_qa_exps__pdfs_plus_scp_filtered_2850__verified_1k_len_r1_eval_03-18-25_22-16-54_0981qwen2-5_sci_qa_exps__scp_filtered_1664__verified_1k_len_r1_eval_03-18-25_22-26-45_0981
mlfoundations-dev/qwen2-5_sci_qa_exps__scp_filtered_1664__verified_1k_len_r1_eval_03-18-25_22-26-45_0981
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AIME25
AMC23
MATH500
GPQADiamond
LiveCodeBench
Accuracy
26.0
23.3
71.0
82.8
22.7
32.6
AIME24
Average Accuracy: 26.00% ± 1.74%
Number of Runs: 5
Run
Accuracy
Questions Solved
Total Questions
1
30.00%
9
30
2
30.00%
9
30
3
20.00%
6
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/qwen2-5_sci_qa_exps__scp_filtered_1664__verified_1k_len_r1_eval_03-18-25_22-26-45_0981.qwen2-5_sci_qa_exps__pdfs_plus_scp_filtered_2850__verified_1k_len_r1_eval_03-11-25_08-21-40_f912
mlfoundations-dev/qwen2-5_sci_qa_exps__pdfs_plus_scp_filtered_2850__verified_1k_len_r1_eval_03-11-25_08-21-40_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 32.15% ± 3.78%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
27.27%
54
198
2
27.78%
55
198
3
41.41%
82
198
qwen2-5_sci_qa_exps__scp_filtered_3103__unverified_1k_len_r1_eval_03-18-25_22-20-12_0981
