datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enemThe ENEM 2022, 2023 and 2024 datasets encompass all multiple-choice questions from the last two editions of the Exame Nacional do Ensino Médio (ENEM), the main standardized entrance examination adopted by Brazilian universities. The datasets have been created to allow the evaluation of both textual-only and textual-visual language models. To evaluate textual-only models, we incorporated into the datasets the textual descriptions of the images that appear in the questions' statements from the… See the full description on the dataset page: https://huggingface.co/datasets/maritaca-ai/enem.enem_challengeapex_enemy_detect
Apex Legends Enemy Detection Dataset
Dataset for detecting players in Apex Legends gameplay footage.4 921 frames — 2 classes: enemy, mate.
Dataset Structure
Split
Images
Labels
train
3 937
3 937
val
984
984
images/
train/ # 3937 × .png
val/ # 984 × .png
labels/ # YOLO .txt, mirrors images/
dataset.yaml
Annotation Format
YOLO — each .txt contains one row per bounding box:
<class_id> <cx> <cy> <w> <h> # normalized… See the full description on the dataset page: https://huggingface.co/datasets/SauravOfficial/apex_enemy_detect.aes_enem_dataset
Automated Essay Score (AES) ENEM Dataset
Use Case and Creators
Intended Use: Estimate Essay Score
Creators: Igor Cataneo Silveira, André Barbosa and Denis Deratani Mauá
Contact Information: igorcs@ime.usp.br; andre.barbosa@ime.usp.br
Licensing Information
License: MIT License
Citation Details
Preferred Citation:
@proceedings{DBLP:conf/propor/2024,
editor = {Igor Cataneo Silveira, André Barbosa and Denis Deratani Mauá},
title =… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/aes_enem_dataset.apex_enemy_detect
Apex Legends Enemy Detection Dataset
Dataset for detecting players in Apex Legends gameplay footage.4 921 frames — 2 classes: enemy, mate.
Dataset Structure
Split
Images
Labels
train
3 937
3 937
val
984
984
images/
train/ # 3937 × .png
val/ # 984 × .png
labels/ # YOLO .txt, mirrors images/
dataset.yaml
Annotation Format
YOLO — each .txt contains one row per bounding box:
<class_id> <cx> <cy> <w> <h> # normalized… See the full description on the dataset page: https://huggingface.co/datasets/PSImera/apex_enemy_detect.en-embeddings-bgeenem-audiodescricao
Audiodescrição profissional de imagens do ENEM: corpus em português para acessibilidade e avaliação de modelos de visão
Professional audio descriptions of ENEM exam figures: a Brazilian Portuguese corpus for accessibility research and vision-language model evaluation.
Corpus de audiodescrições escritas por profissionais para as figuras do ENEM, extraídas
dos cadernos "ledor" que o INEP publica para participantes com deficiência visual, alinhadas
à figura, ao enunciado, às… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/enem-audiodescricao.apex_enemy_detect
Apex Legends Enemy Detection Dataset
Dataset for detecting players in Apex Legends gameplay footage.4 921 frames — 2 classes: enemy, mate.
Dataset Structure
Split
Images
Labels
train
3 937
3 937
val
984
984
images/
train/ # 3937 × .png
val/ # 984 × .png
labels/ # YOLO .txt, mirrors images/
dataset.yaml
Annotation Format
YOLO — each .txt contains one row per bounding box:
<class_id> <cx> <cy> <w> <h> # normalized… See the full description on the dataset page: https://huggingface.co/datasets/frani7659/apex_enemy_detect.apex_enemy_detect
Apex Legends Enemy Detection Dataset
Dataset for detecting players in Apex Legends gameplay footage.4 921 frames — 2 classes: enemy, mate.
Dataset Structure
Split
Images
Labels
train
3 937
3 937
val
984
984
images/
train/ # 3937 × .png
val/ # 984 × .png
labels/ # YOLO .txt, mirrors images/
dataset.yaml
Annotation Format
YOLO — each .txt contains one row per bounding box:
<class_id> <cx> <cy> <w> <h> # normalized… See the full description on the dataset page: https://huggingface.co/datasets/aadasdadasdsa/apex_enemy_detect.EN_Emilia_Yodas_616hthe dataset is 616h out of the English part from https://huggingface.co/datasets/amphion/Emilia-Dataset ( Emilia Yodas - cc by 4.0)
audio event classified via scribe v1 (elevenlabs stt/asr)
facebook audio aestetics to be used as prefilter
the dataset is very much at a v1 -
if you want to help - lets talk
https://discord.gg/RUs3uzBdW3 (nsfw is fully opt in only - as sfw)
if you want full transaction timestamps as they come from scribe v1 - they are cc by 4.0 NC and can be found here… See the full description on the dataset page: https://huggingface.co/datasets/MrDragonFox/EN_Emilia_Yodas_616h.enem-2023-d2-multimodalenem_challenge
ENEM Challenge
The Exame Nacional do Ensino Médio (ENEM) is an advanced High-School level exam widely applied every year by the Brazilian government to students that wish to undertake a University degree. This dataset contains 1,430 questions that don't require image understanding of the exams from 2010 to 2018, 2022 and 2023. The model is given a question in Portuguese together with its answer choices and must respond with the correct alternative letter (A, B, C, D or E)… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/enem_challenge.EN_Emilia_Yodas_ScribeEvents
EN Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/EN_Emilia_Yodas_616h containing only samples with ElevenLabs Scribe v1 audio events (vocal bursts, background sounds, etc.).
Changes from source
Filtered to only include rows where events_scribe is non-empty (16017 rows out of 228,265 original)
Bracket format unified: Round brackets (laughs) in text_scribe replaced with square brackets [laughs] for consistency with vocal burst annotation format… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/EN_Emilia_Yodas_ScribeEvents.enem-2025-d2-multimodalenem-2022-d2-multimodalenem-2024-d1-multimodalenem-2022-d1-multimodalenem-2023-d1-multimodalgaia-edu-enem
GAIA-EDU — Questões ENEM (INEP)
Parte do projeto GAIA-EDU — corpus educacional brasileiro desenvolvido pelo
CEIA-UFG (Centro de Excelência em IA da Universidade Federal de Goiás)
para o agente tutor socrático GAIA.
815 registros | Idioma: Português (PT-BR) | Organização: CEIA-GAIA-EDU
Sobre este dataset
Questões estruturadas das provas do ENEM (Exame Nacional do Ensino Médio) extraídas dos PDFs oficiais do INEP. Inclui enunciado, cinco alternativas (A-E) e gabarito… See the full description on the dataset page: https://huggingface.co/datasets/CEIA-GAIA-EDU/gaia-edu-enem.enem-2024-d2-multimodalENEMen_emo_speechenem-2025-d1-multimodalenem-ocr
ENEM-OCR
3.506 questões do ENEM (1998–2024, 27 edições) com gabarito e figuras,
extraídas dos PDFs públicos do INEP por um pipeline OCR de 4 estágios
(OCR + detecção de layout -> estruturação por LLM em lote -> mesclagem +
descrições visuais de figuras (VLM) + CoTs -> auditoria). Construído para a
dissertação de mestrado sobre Continued Pre-Training em pequena escala para
o ENEM em PT-BR.
Colunas
id, year, exam, question_number, area (CH/CN/LC/MT), subject_hint… See the full description on the dataset page: https://huggingface.co/datasets/candido-ai/enem-ocr.enem-judgeenem_2025
🇧🇷 ENEM 2025 — Brazilian National High School Exam Dataset
A High-Quality Benchmark for Portuguese Academic Reasoning in Large Language Models
ENEM 2025 Dataset is a curated collection of question-answer pairs derived from the 2025 edition of the Brazilian National High School Exam (ENEM), designed to evaluate and improve the reasoning, reading comprehension, and multiple-choice answering capabilities of large language models in Brazilian Portuguese; as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/enem_2025.self-collected-ENEM-dataset-with-prompts-and-text-supportenem-2023-dia-1
Dataset Card for "enem-2023-dia-1"
More Information needed
self-collected-ENEM-datasetenem-judge-correct
