CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amalia-llm /bigbenchhard-mt-pt BBH-PT (Big-Bench Hard) Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks. Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations. Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese. Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.textquestion-answering1K<n<10K0 likes175 downloads3mo agoHugging Face02amalia-llm /pt_exams PHEB - Portuguese High School Exams MCQ MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum. For more details, see the PHEB paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.tabularquestion-answering1K<n<10K0 likes167 downloads3mo agoHugging Face03amalia-llm /alba_mcq ALBA MCQ Multiple Choice Version of the ALBA benchmark, a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks. For more details, see the ALBA paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba_mcq.tabularquestion-answeringn<1K0 likes122 downloads3mo agoHugging Face04amalia-llm /AMALIA-LLM-0626-SFT-Dataset AMALIA LLM Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training. Base Data Mix This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count amalia-llm/persona_math 63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.texttext-generation1M<n<10M2 likes68 downloads3mo agoHugging Face05amalia-llm /persona_nemotron Persona Nemotron PT Datasets This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests. The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.textquestion-answering100K<n<1M0 likes59 downloads3mo agoHugging Face06amalia-llm /PT-Culture_Data Portuguese Cultural SFT Dataset A supervised fine-tuning dataset for European Portuguese (PT-PT) cultural knowledge, built to teach models the traditions, figures, places, and expressions of Portuguese culture. The dataset contains 216,832 examples across two versions in conversational SFT format, organised into ten cultural domains: Domain v1 (95,815) v2 (121,017) Total Personalities 36,370 62,788 99,158 Audiovisual 15,786 26,857 42,643 Heritage 8,020… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/PT-Culture_Data.textquestion-answering100K<n<1M1 likes56 downloads2mo agoHugging Face07amalia-llm /smoltalk2_everyday_conv_pt SMOL Everyday Conversation PT This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B. Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.texttext-generation1K<n<10K1 likes52 downloads3mo agoHugging Face08amalia-llm /wikipedia-rag Wikipedia RAG This dataset consists of raw wikipedia pages with one question and answer for each page. The answers were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title = "{AMALIA}: A Fully Open Large Language Model for… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wikipedia-rag.texttext-generation10K<n<100K0 likes51 downloads3mo agoHugging Face09amalia-llm /saudade-pt SAUDADE Portuguese benchmark for temporal reasoning and understanding of temporal relationships across Portuguese events. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title = "{AMALIA}: A Fully Open… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/saudade-pt.textquestion-answering1K<n<10K0 likes47 downloads3mo agoHugging Face10amalia-llm /amalia-Dolci-Instruct-SFT AMALIA Dolci-Instruct-SFT Version of the allenai/Dolci-Instruct-SFT dataset used in the ramp down phase of AMALIA's Supervised Fine-Tuning stage. This dataset was developed by sampling the highest quality entries of the selected splits, and translating part of those entries to European Portuguese. Both the quality classification and translation were done using google/gemma-4-31B-it. This dataset went through a processing pipeline to: Remove entries that reference… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Dolci-Instruct-SFT.textother100K<n<1M0 likes47 downloads3mo agoHugging Face11amalia-llm /persona_instruction_following Persona Instruction Following Datasets This is a synthetic instruction-following dataset, available in two configs: full and filtered. Each config contains two language splits, English (en) and Portuguese (pt). The filtered version keeps only the higher-quality examples (quality score 5). The prompts were generated using an approach similar to PersonaHub, with a translated version of proj-persona/PersonaHub. Both prompts and answers were generated using Gemma 3-27B.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_instruction_following.textquestion-answering100K<n<1M0 likes45 downloads3mo agoHugging Face12amalia-llm /amalia-Nemotron-SFT-Math-v3 AMALIA Nemotron-SFT-Math-v3 Version of the nvidia/Nemotron-SFT-Math-v3 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries with tool usage; Remove entries that reference other LLMs or research labs; Removed the reasoning_content field; Drop extra unnecessary columns; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v3 This dataset is provided as part of the AMALIA… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Math-v3.texttext-generation1M<n<10M0 likes44 downloads3mo agoHugging Face13amalia-llm /DPO-Dataset AMALIA DPO Dataset This is the DPO (preference optimization) dataset used to train AMALIA-DPO. It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets. This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.tabularquestion-answering100K<n<1M4 likes42 downloads3mo agoHugging Face14amalia-llm /amalia-Nemotron-Instruction-Following-Chat-v1 AMALIA Nemotron-Instruction-Following-Chat-v1 Version of the nvidia/Nemotron-Instruction-Following-Chat-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries where the capability_target field was 'chat' from the chat_if split; Remove entries that reference other LLMs or research labs; Remove entries that contained the string '/imagine prompt:'; Removed the reasoning_content field; Original… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Instruction-Following-Chat-v1.text10K<n<100K0 likes42 downloads3mo agoHugging Face15amalia-llm /hermes3_special_system_prompts Hermes 3 SFT Special System Prompts Version of the NousResearch/Hermes-3-Dataset dataset used in AMALIA's Supervised Fine-Tuning stage. This subset was manually selected in order to retain only entries containing custom system prompts that substantially modify the model’s behavior. It were also built splits with the highest quality entries and a translated part of those entries to European Portuguese. Both the quality classification and translation were done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hermes3_special_system_prompts.text100K<n<1M0 likes38 downloads3mo agoHugging Face16amalia-llm /piqa-mt-pt PIQA-PT Portuguese machine translation of PIQA (Physical Interaction QA), a benchmark for physical commonsense reasoning. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/ybisk/piqa Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/piqa-mt-pt.textquestion-answering10K<n<100K0 likes37 downloads3mo agoHugging Face17amalia-llm /amalia-Nemotron-Science-v1 AMALIA Nemotron-Science-v1 Version of the nvidia/Nemotron-Science-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries that reference other LLMs or research labs; Remove the reasoning_content field; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-Science-v1 This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Science-v1.text100K<n<1M0 likes36 downloads3mo agoHugging Face18amalia-llm /amalia-PTradutor AMALIA PTradutor This dataset consists of translations from English to Portuguese and from Portuguese to English. These pairs come from liaad/PTradutor and were transformed into a conversation format by adding an instruction to translate the text, placed in either the user message or the system message. This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-PTradutor.texttranslation10K<n<100K0 likes36 downloads3mo agoHugging Face19teex-pt /amalia-pilot-honesty-v2 AMALIA pilot — honesty vector datasets (v1 refusals + v2 corrective mix) Training data from the first two iterations of a verifier-gated fine-tuning pilot on AMALIA-9B-0626-DPO, targeting identity/fact confabulation (the model's weakest measured behavior: 43.3% on our honesty harness). Full methodology, harness, and reports: github.com/teex-pt/pt-amalia. These are research pilot artifacts — small by design (the pilot validates the loop, not the scale). Every sample was produced… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-pilot-honesty-v2.texttext-generation1K<n<10K0 likes35 downloads3mo agoHugging Face20amalia-llm /Amalia_hardcoded AMALIA Hardcoded Handwritten dataset to provide the AMALIA model with self-referential knowledge about its development and capabilities, both in Portuguese and English. This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title = "{AMALIA}: A Fully Open Large Language… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/Amalia_hardcoded.textquestion-answeringn<1K0 likes34 downloads3mo agoHugging Face21amalia-llm /AMALIA-LLM-1225-SFT-Dataset AMALIA LLM 1225 Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model version released in December 2025. This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count Persona-PT Instruction Following 9,084 Persona-EN Instruction Following 34,704 Persona Nemotron Instruction Following 4,483… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-1225-SFT-Dataset.texttext-generation1M<n<10M0 likes34 downloads3mo agoHugging Face22amalia-llm /pt_text_completion PT-PT Completions Simple text-completion dataset to evaluete model bias towards European Portuguese (pt-PT) or Brazilian Portuguese (pt-BR). This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title =… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_text_completion.textn<1K0 likes33 downloads3mo agoHugging Face23amalia-llm /wikipedia_conversations Wikipedia Conversations This is a dataset of conversations based on Portuguese Wikipedia articles. The conversations were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title = "{AMALIA}: A Fully Open Large Language Model for… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wikipedia_conversations.textquestion-answering10K<n<100K0 likes31 downloads3mo agoHugging Face24amalia-llm /openbookqa-mt-pt OpenBookQA-PT Portuguese machine translation of OpenBookQA, a question-answering dataset modeled after open-book exams for elementary science. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/allenai/openbookqa Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/openbookqa-mt-pt.tabularquestion-answering10K<n<100K0 likes31 downloads3mo agoHugging Face25AmaliaE /PersonaHub Scaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/AmaliaE/PersonaHub.texttext-generation100K<n<1M0 likes29 downloads9mo agoHugging Face26amalia-llm /amalia-Nemotron-SFT-Instruction-Following-Chat-v2 Nemotron-SFT-Instruction-Following-Chat-v2 Version of the nvidia/Nemotron-SFT-Instruction-Following-Chat-v2 dataset used in the AMALIA's Supervised Fine-Tuning stage, both in the base and ramp down stages. The ramp down stage comprised a randomly select subset of the base subset, where part was translated to European Portuguese using google/gemma-4-31B-it. This dataset went through a processing pipeline to: Remove entries that reference other LLMs or research labs;… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Instruction-Following-Chat-v2.texttext-generation10K<n<100K0 likes29 downloads3mo agoHugging Face27amalia-llm /smol-rewrite-PT SMOL Rewrite PT This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk. This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it. Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk Note: This dataset comprises machine translated content and may contain translation errors or artifacts. This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol-rewrite-PT.textquestion-answering10K<n<100K0 likes28 downloads3mo agoHugging Face28amalia-llm /smol_summarize_pt SMOL Summarize PT This dataset is the translated version of the smol-summarize subset of the HuggingFaceTB/smoltalk. Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk Note: This dataset comprises machine translated content and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation If you use this dataset or… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol_summarize_pt.texttext-generation10K<n<100K0 likes27 downloads3mo agoHugging Face29amalia-llm /amalia-Nemotron-SpecializedDomains-Finance-v1 AMALIA Nemotron-SpecializedDomains-Finance-v1 Version of the nvidia/Nemotron-SpecializedDomains-Finance-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries that reference other LLMs or research labs; Remove the reasoning_content field; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1 This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SpecializedDomains-Finance-v1.texttext-generation100K<n<1M0 likes27 downloads3mo agoHugging Face30amalia-llm /persona_math Persona Math This is a mathematical European Portuguese synthetic dataset. The datasets were generated using an approach similar to PersonaHub, with a translated version of proj-persona/PersonaHub. Both prompts and answers were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_math.texttext-generation10K<n<100K0 likes26 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.