CoolFace
Datasetpublic

amu-cai/reVISION-dataset

Dataset description This dataset was presented and described further on FedCSIS 2025 conference and the full paper can be found here: https://annals-csis.org/proceedings/2025/pliks/2608.pdf The dataset is a collection of image-based questions sourced from Polish National Exams. Each question is represnted in the form of an image with only one correct answer. The questions distribution is described in the table below: Exam Discipline Questions 8th-Grade Exam Polish… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/reVISION-dataset.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes25downloads
Dataset Card

Dataset description

This dataset was presented and described further on FedCSIS 2025 conference and the full paper can be found here: https://annals-csis.org/proceedings/2025/pliks/2608.pdf

The dataset is a collection of image-based questions sourced from Polish National Exams. Each question is represnted in the form of an image with only one correct answer. The questions distribution is described in the table below:

ExamDisciplineQuestions
8th-Grade ExamPolish Language9
8th-Grade ExamMathematics46
Middle School ExamMathematics and Nature152
Middle School ExamMathematics96
Middle School ExamNature79
High School ExamBiology21
High School ExamPhysics154
High School ExamMathematics280
Professional ExamArts2547
Professional ExamMechanical, Mining and Metallurgical21057
Professional ExamAgriculture and Forestry14905

Manual verification revealed that 122 of the 1,000 images were misclassified in needsimagecontext question metadata, with the vast majority being false positives (120 out of 122). While this level of error is acceptable for exploratory analysis and broad grouping, it is not sufficiently accurate for training downstream models. We therefore emphasize that findings based on this stratification are preliminary and should not be overinterpreted without more precise annotation.