amu-cai/reVISION-dataset
Dataset description This dataset was presented and described further on FedCSIS 2025 conference and the full paper can be found here: https://annals-csis.org/proceedings/2025/pliks/2608.pdf The dataset is a collection of image-based questions sourced from Polish National Exams. Each question is represnted in the form of an image with only one correct answer. The questions distribution is described in the table below: Exam Discipline Questions 8th-Grade Exam Polish… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/reVISION-dataset.
Dataset description
This dataset was presented and described further on FedCSIS 2025 conference and the full paper can be found here: https://annals-csis.org/proceedings/2025/pliks/2608.pdf
The dataset is a collection of image-based questions sourced from Polish National Exams. Each question is represnted in the form of an image with only one correct answer. The questions distribution is described in the table below:
Manual verification revealed that 122 of the 1,000 images were misclassified in needsimagecontext question metadata, with the vast majority being false positives (120 out of 122). While this level of error is acceptable for exploratory analysis and broad grouping, it is not sufficiently accurate for training downstream models. We therefore emphasize that findings based on this stratification are preliminary and should not be overinterpreted without more precise annotation.
