datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enem-ocr
ENEM-OCR
3.506 questões do ENEM (1998–2024, 27 edições) com gabarito e figuras,
extraídas dos PDFs públicos do INEP por um pipeline OCR de 4 estágios
(OCR + detecção de layout -> estruturação por LLM em lote -> mesclagem +
descrições visuais de figuras (VLM) + CoTs -> auditoria). Construído para a
dissertação de mestrado sobre Continued Pre-Training em pequena escala para
o ENEM em PT-BR.
Colunas
id, year, exam, question_number, area (CH/CN/LC/MT), subject_hint… See the full description on the dataset page: https://huggingface.co/datasets/candido-ai/enem-ocr.enem_2025
🇧🇷 ENEM 2025 — Brazilian National High School Exam Dataset
A High-Quality Benchmark for Portuguese Academic Reasoning in Large Language Models
ENEM 2025 Dataset is a curated collection of question-answer pairs derived from the 2025 edition of the Brazilian National High School Exam (ENEM), designed to evaluate and improve the reasoning, reading comprehension, and multiple-choice answering capabilities of large language models in Brazilian Portuguese; as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/enem_2025.
