CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amalia-llm /AMALIA-VL-SFT-Dataset AMALIA-VL-Training-Dataset Dataset Description This dataset is provided as part of the AMALIA project. This is the vision+language training mix for AMALIA-VL-SFT. Each subset is one source dataset in the mix, each with a single train split. The only datasets that are absent from this mix are those that derive directly from the core LLM training mix, and can be found in the AMALIA-LLM Post Training Collection. Example usage: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-SFT-Dataset.imagevisual-question-answering1M<n<10M2 likes215 downloads3mo agoHugging Face02amalia-llm /SEED-Bench-PT SEED-Bench-PT European Portuguese (pt-PT) machine translation of SEED-Bench, a multiple-choice benchmark spanning multiple dimensions of multimodal comprehension. Translated from the original English test split using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/lmms-lab/SEED-Bench Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SEED-Bench-PT.imageimage-text-to-text10K<n<100K0 likes206 downloads3mo agoHugging Face03amalia-llm /CorEGe-PT CorEGe-PT: Corpus do Estudo Geral - Portuguese CorEGe-PT is a large-scale corpus of academic texts written in Portuguese (mainly European Portuguese), extracted from Estudo Geral, the institutional repository of the University of Coimbra. It contains over 34,000 documents and approximately 1 billion tokens, making it the largest available corpus of its kind for the Portuguese language. This dataset is designed to support linguistic research (Academic Discourse Studies) and the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/CorEGe-PT.text10K<n<100K3 likes180 downloads20d agoHugging Face04amalia-llm /bigbenchhard-mt-pt BBH-PT (Big-Bench Hard) Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks. Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations. Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese. Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.textquestion-answering1K<n<10K0 likes175 downloads3mo agoHugging Face05amalia-llm /pt_exams PHEB - Portuguese High School Exams MCQ MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum. For more details, see the PHEB paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.tabularquestion-answering1K<n<10K0 likes167 downloads3mo agoHugging Face06amalia-llm /alba_mcq ALBA MCQ Multiple Choice Version of the ALBA benchmark, a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks. For more details, see the ALBA paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba_mcq.tabularquestion-answeringn<1K0 likes122 downloads3mo agoHugging Face07amalia-llm /hendrycks-math-ptpt Hendrycks Math MT-PT Portuguese translated mathematics problems covering algebra, geometry, number theory, and more. Translated using Gemma-4 31B-It. Original Dataset: https://huggingface.co/datasets/EleutherAI/hendrycks_math Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hendrycks-math-ptpt.textquestion-answering10K<n<100K0 likes108 downloads3mo agoHugging Face08amalia-llm /InfographicVQA-PT InfographicVQA-PT European Portuguese (pt-PT) machine translation of InfographicVQA, a visual question answering dataset over infographics that combine text, graphics, and data visualizations. Translated from the original English validation split (InfographicVQA subset) using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/lmms-lab/DocVQA (InfographicVQA subset) Note: This dataset is machine translated and may contain translation errors or artifacts.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/InfographicVQA-PT.imageimage-text-to-text1K<n<10K0 likes87 downloads3mo agoHugging Face09amalia-llm /AMALIA-LLM-0626-SFT-Dataset AMALIA LLM Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training. Base Data Mix This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count amalia-llm/persona_math 63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.texttext-generation1M<n<10M2 likes68 downloads3mo agoHugging Face10amalia-llm /persona_nemotron Persona Nemotron PT Datasets This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests. The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.textquestion-answering100K<n<1M0 likes59 downloads3mo agoHugging Face11amalia-llm /COCO-Caption2017-PT COCO-Caption2017-PT European Portuguese (pt-PT) machine translation of COCO Captions 2017, an image captioning dataset of everyday scenes with human-written captions. Translated from the original English val split using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/lmms-lab/COCO-Caption2017 Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/COCO-Caption2017-PT.imageimage-text-to-text1K<n<10K0 likes59 downloads3mo agoHugging Face12amalia-llm /CARAVELA CARAVELA CARAVELA is a multimodal benchmark for evaluating the Portuguese cultural knowledge of large vision-language models (LVLMs). The official benchmark language is exclusively European Portuguese (pt-PT). Motivation Modern LVLMs excel at general-purpose vision-language tasks, but their performance drops sharply on localized, culturally specific content that is underrepresented in global training data. Portugal has a rich cultural heritage —… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/CARAVELA.imagevisual-question-answering10K<n<100K0 likes57 downloads3mo agoHugging Face13duarteocarmo /amalia-sft AMALIA SFT Standardized European Portuguese supervised fine-tuning data from AMALIA. Each config has deterministic train and validation splits and a common messages schema. Exact duplicate and invalid conversations are removed. Source provenance is retained in every row. Counts Config Train Validation pt_persona_instruction 8902 182 pt_nemotron_instruction 4394 90 pt_nemotron_general 69304 1415 pt_wikipedia 95592 1951 pt_culture 88058 1798… See the full description on the dataset page: https://huggingface.co/datasets/duarteocarmo/amalia-sft.texttext-generation100K<n<1M0 likes57 downloads2mo agoHugging Face14amalia-llm /PT-Culture_Data Portuguese Cultural SFT Dataset A supervised fine-tuning dataset for European Portuguese (PT-PT) cultural knowledge, built to teach models the traditions, figures, places, and expressions of Portuguese culture. The dataset contains 216,832 examples across two versions in conversational SFT format, organised into ten cultural domains: Domain v1 (95,815) v2 (121,017) Total Personalities 36,370 62,788 99,158 Audiovisual 15,786 26,857 42,643 Heritage 8,020… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/PT-Culture_Data.textquestion-answering100K<n<1M1 likes56 downloads2mo agoHugging Face15amalia-llm /smoltalk2_everyday_conv_pt SMOL Everyday Conversation PT This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B. Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.texttext-generation1K<n<10K1 likes52 downloads3mo agoHugging Face16amalia-llm /wikipedia-rag Wikipedia RAG This dataset consists of raw wikipedia pages with one question and answer for each page. The answers were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title = "{AMALIA}: A Fully Open Large Language Model for… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wikipedia-rag.texttext-generation10K<n<100K0 likes51 downloads3mo agoHugging Face17amalia-llm /saudade-pt SAUDADE Portuguese benchmark for temporal reasoning and understanding of temporal relationships across Portuguese events. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title = "{AMALIA}: A Fully Open… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/saudade-pt.textquestion-answering1K<n<10K0 likes47 downloads3mo agoHugging Face18amalia-llm /amalia-Dolci-Instruct-SFT AMALIA Dolci-Instruct-SFT Version of the allenai/Dolci-Instruct-SFT dataset used in the ramp down phase of AMALIA's Supervised Fine-Tuning stage. This dataset was developed by sampling the highest quality entries of the selected splits, and translating part of those entries to European Portuguese. Both the quality classification and translation were done using google/gemma-4-31B-it. This dataset went through a processing pipeline to: Remove entries that reference… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Dolci-Instruct-SFT.textother100K<n<1M0 likes47 downloads3mo agoHugging Face19amalia-llm /persona_instruction_following Persona Instruction Following Datasets This is a synthetic instruction-following dataset, available in two configs: full and filtered. Each config contains two language splits, English (en) and Portuguese (pt). The filtered version keeps only the higher-quality examples (quality score 5). The prompts were generated using an approach similar to PersonaHub, with a translated version of proj-persona/PersonaHub. Both prompts and answers were generated using Gemma 3-27B.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_instruction_following.textquestion-answering100K<n<1M0 likes45 downloads3mo agoHugging Face20amalia-llm /wildguardmix-ptpt WildGuardTest-PT Portuguese machine translation of WildGuardTest, a benchmark for evaluating safety guardrails in language models. Translated using Gemma-4 31B-It. Original Dataset: https://huggingface.co/datasets/walledai/WildGuardTest Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wildguardmix-ptpt.tabulartext-generation1K<n<10K0 likes44 downloads3mo agoHugging Face21amalia-llm /amalia-Nemotron-SFT-Math-v3 AMALIA Nemotron-SFT-Math-v3 Version of the nvidia/Nemotron-SFT-Math-v3 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries with tool usage; Remove entries that reference other LLMs or research labs; Removed the reasoning_content field; Drop extra unnecessary columns; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v3 This dataset is provided as part of the AMALIA… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Math-v3.texttext-generation1M<n<10M0 likes44 downloads3mo agoHugging Face22amalia-llm /alba ALBA ALBA is a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks. For more details, see the ALBA paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title = "{AMALIA}: A… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba.textquestion-answering1K<n<10K1 likes42 downloads3mo agoHugging Face23amalia-llm /DPO-Dataset AMALIA DPO Dataset This is the DPO (preference optimization) dataset used to train AMALIA-DPO. It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets. This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.tabularquestion-answering100K<n<1M4 likes42 downloads3mo agoHugging Face24amalia-llm /amalia-Nemotron-Instruction-Following-Chat-v1 AMALIA Nemotron-Instruction-Following-Chat-v1 Version of the nvidia/Nemotron-Instruction-Following-Chat-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries where the capability_target field was 'chat' from the chat_if split; Remove entries that reference other LLMs or research labs; Remove entries that contained the string '/imagine prompt:'; Removed the reasoning_content field; Original… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Instruction-Following-Chat-v1.text10K<n<100K0 likes42 downloads3mo agoHugging Face25amalia-llm /TextVQA-PT TextVQA-PT European Portuguese (pt-PT) machine translation of TextVQA, a visual question answering dataset that requires reading and reasoning about text in images. Translated from the original English validation split using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/lmms-lab/textvqa Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/TextVQA-PT.imageimage-text-to-text1K<n<10K0 likes41 downloads3mo agoHugging Face26amalia-llm /MMMU-Pro-PT MMMU-Pro-PT European Portuguese (pt-PT) machine translation of MMMU-Pro, a more robust and challenging version of MMMU for college-level, multi-discipline multimodal reasoning. Translated from the original English test split (standard (10 options) subset) using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/MMMU/MMMU_Pro (standard (10 options) subset) Note: This dataset is machine translated and may contain translation errors or artifacts. This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MMMU-Pro-PT.imagequestion-answering1K<n<10K0 likes39 downloads3mo agoHugging Face27amalia-llm /hermes3_special_system_prompts Hermes 3 SFT Special System Prompts Version of the NousResearch/Hermes-3-Dataset dataset used in AMALIA's Supervised Fine-Tuning stage. This subset was manually selected in order to retain only entries containing custom system prompts that substantially modify the model’s behavior. It were also built splits with the highest quality entries and a translated part of those entries to European Portuguese. Both the quality classification and translation were done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hermes3_special_system_prompts.text100K<n<1M0 likes38 downloads3mo agoHugging Face28amalia-llm /piqa-mt-pt PIQA-PT Portuguese machine translation of PIQA (Physical Interaction QA), a benchmark for physical commonsense reasoning. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/ybisk/piqa Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/piqa-mt-pt.textquestion-answering10K<n<100K0 likes37 downloads3mo agoHugging Face29amalia-llm /P3B3 P3B3 Portuguese multi-turn conversational benchmark for measuring European and Brazilian Portuguese variety bias in LLMs. For more details, see the P3B3 paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/P3B3.texttext-generationn<1K0 likes37 downloads3mo agoHugging Face30amalia-llm /xstest_ptpt XSTest-PT Portuguese machine translation of XSTest, a benchmark for identifying exaggerated safety behaviors in language models. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/Paul/XSTest Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/xstest_ptpt.texttext-generationn<1K0 likes36 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.