amalia
Datasets
All datasets matching “amalia”AMALIA-LLM-0626-SFT-Dataset-classlangcode
AMALIA-LLM-0626-SFT-Dataset-classlangcode
Overview
Este dataset é baseado no amalia-llm/AMALIA-LLM-0626-SFT-Dataset, mantendo integralmente a estrutura das conversações e adicionando metadados obtidos através de classificação automática. O objetivo desta versão é facilitar a seleção, filtragem e construção de subconjuntos especializados para treino e avaliação de modelos de linguagem, sem alterar o conteúdo original das conversações. O dataset original foi… See the full description on the dataset page: https://huggingface.co/datasets/inaciose/AMALIA-LLM-0626-SFT-Dataset-classlangcode.SEED-Bench-PT
SEED-Bench-PT
European Portuguese (pt-PT) machine translation of SEED-Bench, a multiple-choice benchmark spanning multiple dimensions of multimodal comprehension.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/SEED-Bench
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SEED-Bench-PT.AMALIA-VL-SFT-Dataset
AMALIA-VL-Training-Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
This is the vision+language training mix for AMALIA-VL-SFT. Each subset is one
source dataset in the mix, each with a single train split. The only datasets that are absent from this mix are those that derive directly from the core LLM training mix, and can be found in the AMALIA-LLM Post Training Collection.
Example usage:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-SFT-Dataset.bigbenchhard-mt-pt
BBH-PT (Big-Bench Hard)
Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks.
Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations.
Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese.
Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard
Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.CorEGe-PT
CorEGe-PT: Corpus do Estudo Geral - Portuguese
CorEGe-PT is a large-scale corpus of academic texts written in Portuguese (mainly European Portuguese), extracted from Estudo Geral, the institutional repository of the University of Coimbra. It contains over 34,000 documents and approximately 1 billion tokens, making it the largest available corpus of its kind for the Portuguese language.
This dataset is designed to support linguistic research (Academic Discourse Studies) and the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/CorEGe-PT.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.
