CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TNSA /PT-HF500B PT-HF500B (FinePhrase) Overview FinePhrase is a large-scale synthetic dataset designed for high-quality language modeling, reasoning, and instruction-following tasks. It transforms raw educational web data into structured, instruction-rich formats suitable for training advanced language models. This dataset has been extensively used in the pre-training pipeline of TNSA models, including: NGen-3 NGen-4 NGen-4-OW It plays a critical role in improving reasoning ability… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/PT-HF500B.tabulartext-generation1B<n<10B1 likes1.7k downloads6mo agoHugging Face02DedeProGames /claude-code-traces-pt-brThis dataset was generated using teich by TeichAI claude Agent Traces This directory contains raw agent trace files generated by teich. JSONL files: 20 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.tabulartext-generationn<1K2 likes205 downloads2mo agoHugging Face03amalia-llm /pt_exams PHEB - Portuguese High School Exams MCQ MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum. For more details, see the PHEB paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.tabularquestion-answering1K<n<10K0 likes159 downloads3mo agoHugging Face04Davizig10jojo /Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR 🇧 Destilação PT-BR com Raciocínio (Chain-of-Thought) Este dataset contém exemplos de alta qualidade gerados através da destilação de modelos de ponta (Teacher Models) disponíveis via NVIDIA NIM, focados em instrução, raciocínio lógico e naturalidade em Português Brasileiro (PT-BR). O grande diferencial deste dataset é a inclusão explícita do processo de pensamento (Chain-of-Thought / thinking) dos modelos professores, permitindo treinar modelos menores (Student Models) não… See the full description on the dataset page: https://huggingface.co/datasets/Davizig10jojo/Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR.tabulartext-generationn<1K0 likes54 downloads12d agoHugging Face05amalia-llm /wildguardmix-ptpt WildGuardTest-PT Portuguese machine translation of WildGuardTest, a benchmark for evaluating safety guardrails in language models. Translated using Gemma-4 31B-It. Original Dataset: https://huggingface.co/datasets/walledai/WildGuardTest Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wildguardmix-ptpt.tabulartext-generation1K<n<10K0 likes43 downloads3mo agoHugging Face06StrataSynth /stratasynth-termination-negotiations-pt-br stratasynth-termination-negotiations-pt-br Synthetic employment-termination negotiations between an employer / HR representative and an employee who is being let go. Generated with StrataSynth's dimensional engine, psychometrically and culturally conditioned to Brazil (BR). Every conversation is a two-party exchange grounded in the same scenario (PRO-05 — employment termination): delivering the decision, stated reasons (restructuring, performance, budget), severance and final… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-termination-negotiations-pt-br.tabulartext-generation1K<n<10K0 likes43 downloads4d agoHugging Face07amalia-llm /toxichat-ptpt ToxicChat-PT Portuguese machine translation of ToxicChat, a benchmark for detecting toxic content in conversational AI. Translated using Gemma-4 31B-It. Original Dataset: https://huggingface.co/datasets/lmsys/toxic-chat Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/toxichat-ptpt.tabulartext-generation1K<n<10K0 likes28 downloads3mo agoHugging Face08amalia-llm /aime-1983-2024-ptpt AIME-PT (1983-2024) Portuguese translation of problems from the American Invitational Mathematics Examination (AIME) spanning 1983-2024. Translated using Gemma-4 31B-It. Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime-1983-2024-ptpt.tabularquestion-answeringn<1K0 likes24 downloads3mo agoHugging Face09teex-pt /amalia-cita-legalgated AMALIA cita-legal — grounded legal citation SFT (RAG-first) (question + real source excerpts → answer that cites the exact article, or a refusal when the excerpts don't answer the question). Built for specializing AMALIA-9B toward Portuguese legal text, as the grounded-answering counterpart to teex-pt/amalia-sum-dre. Derived from teex-pt/leis-pt-consolidada by teex-pt/pt-amalia. Why this exists (RAG-first, not closed-book) leis-pt's own project spec concludes that… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-cita-legal.tabularquestion-answering1K<n<10K0 likes7 downloads2mo agoHugging Face10br-llm-data /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes6 downloads3mo agoHugging Face11rides21 /anneal_pt_v3_clean_v2_reshard_7168gated anneal_pt_v3 clean v2 — balanced 7,168 shards This is the cleaned and deterministically re-sharded training split of anneal_pt_v3. It contains 2,103,321,326 rows in 7,168 Parquet files (3,376,320,935,167 stored bytes). The shard count is 8 × lcm(128, 448), so it divides evenly for both intended data-parallel configurations: Data ranks Shards per rank Maximum relative load deviation 128 56 1.4942% 448 16 0.3888% The files are ordered by a fixed-capacity LPT… See the full description on the dataset page: https://huggingface.co/datasets/rides21/anneal_pt_v3_clean_v2_reshard_7168.tabulartext-generation1B<n<10B0 likes5 downloads2mo agoHugging Face12costadev00 /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes3 downloads5mo agoHugging Face13heitorrosa /financial-sentiment-pt Financial Sentiments (PT-BR) A dataset containing ~63k rows of translated Financial Sentiment data in Portuguese. Pipeline Dataset cleaning (dedup and benchmark decontamination) -> merge -> Translation Pipeline -> MiniCPM5 2B -> Kiwi COMET as scorer -> drop below a threshold when doing sft (default = 0.5). Exact code at: https://github.com/mansa-team/musa Data n_scored=64685 mean=0.6893 n_below=6975 tabulartext-generation10K<n<100K0 likes3h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.