CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DedeProGames /claude-code-traces-pt-brThis dataset was generated using teich by TeichAI claude Agent Traces This directory contains raw agent trace files generated by teich. JSONL files: 20 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.tabulartext-generationn<1K2 likes207 downloads2mo agoHugging Face02epicchi2103 /geo-ptbr GEO-PTBR The first Brazilian-Portuguese benchmark for Generative Engine Optimization (GEO): 525 PT-BR queries with 2,625 source documents, plus per-technique citation-visibility results measured on three generative engines. Companion artifact to the paper GEO-PTBR: A Brazilian-Portuguese Replication of Generative Engine Optimization and the Engine-Dependence of Its Effects. Experiment version: 0.3.0. What this measures Given a query and five already-retrieved… See the full description on the dataset page: https://huggingface.co/datasets/epicchi2103/geo-ptbr.texttext-generation1K<n<10K0 likes86 downloads1mo agoHugging Face03costadev00 /wikipedia-pt-br-instruct-5k Wikipedia PT-BR Instruct wikipedia-pt-br-instruct is a synthetic supervised fine-tuning (SFT) dataset in Brazilian Portuguese generated from Wikipedia-derived documents. This release is an intermediate evaluation dataset produced with the sft-dataset-creator pipeline from the run wiki-ptbr-extract-calib-5kdocs-14tasks. It was generated from a fixed revision of costadev00/wikipedia-pt-br-extract: cdbd07dc4a3de6e64632c718710b3ae0ebaeb0ff The dataset is intended for intermediate… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instruct-5k.texttext-generation100K<n<1M1 likes81 downloads3mo agoHugging Face04botbotrobotics /physics-ptbr Tradução do Camel Pyysics dataset para Portuguese (PT-BR) usando NLLB 3.3b. CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/physics-ptbr.texttext-generation10K<n<100K2 likes59 downloads3y agoHugging Face05MrAiran /nsfw-pt_br Dataset Card for The Pile Dataset Summary The NSFW is a 230K diverse, filtred and cleaned text from adult websites, high-quality datasets combined together. Supported Tasks and Leaderboards [More Information Needed] Languages This dataset is in Portuguese Brazil (pt_BR) Dataset Structure Data Instances Data Fields all text (str): Text. Dataset Creation Curation Rationale [More… See the full description on the dataset page: https://huggingface.co/datasets/MrAiran/nsfw-pt_br.texttext-generation100K<n<1M2 likes43 downloads1y agoHugging Face06annajuliaasf /tool-calling-traces-ptbr Tool calling conversations in Portuguese 484 synthetic conversations that teach a model when to call a tool, which one to call and with which arguments, and also when to answer directly, with no tool at all. Each line of the file is a complete conversation: the user's question, the tool call, the simulated return of that tool, and the final answer. It was built because no dataset of tool calling in Portuguese with fictional tools existed. The 30 tools and the user questions were… See the full description on the dataset page: https://huggingface.co/datasets/annajuliaasf/tool-calling-traces-ptbr.texttext-generationn<1K0 likes42 downloads2mo agoHugging Face07divergente /noticias-govbr-ptbr-1texttext-generation10K<n<100K2 likes41 downloads3y agoHugging Face08botbotrobotics /chemistry-ptbr Tradução do Camel Chemisty dataset para Portuguese (PT-BR) usando NLLB 3.3b. CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/chemistry-ptbr.texttext-generation10K<n<100K2 likes40 downloads3y agoHugging Face09benjleite /FairytaleQA-translated-ptBR Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.textquestion-answering10K<n<100K3 likes39 downloads1y agoHugging Face10botbotrobotics /biology-ptbr Tradução do Camel Biology dataset para Portuguese (PT-BR) usando NLLB 3.3b. CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/biology-ptbr.texttext-generation10K<n<100K4 likes36 downloads3y agoHugging Face11guell00 /Lembron-Reason-24K-PTBR Lembron-Reason-24K-PTBR Lebron IA 24K PT-BR Think Distilled SFT é um dataset conversacional em português brasileiro, criado para Supervised Fine-Tuning (SFT), instruction tuning e treinamento de modelos focados em respostas detalhadas, raciocínio estruturado, comportamento de assistente e geração de conteúdo em PT-BR. Este corpus foi estruturado como um dataset distilado do comportamento da Lebron IA, com foco em transferir para modelos menores ou médios um padrão de resposta… See the full description on the dataset page: https://huggingface.co/datasets/guell00/Lembron-Reason-24K-PTBR.texttext-generation10K<n<100K1 likes36 downloads3mo agoHugging Face12celsowm /gemini_orpo_dpo_ptbrtexttext-generation10K<n<100K2 likes35 downloads2y agoHugging Face13wilsondesouza /aurora-dataset-roleplay-ptbr Aurora Dataset Roleplay 🌌 (PT-BR) O que é É um dataset que contém mais de 2 mil diálogos em português do Brasil. Ainda que tenha sido gerado sinteticamente, foi utilizado apenas modelos SOTA, então os diálogos são muito próximos da naturalidade e espontaneidade de um ser humano. Foi feito pensando em roleplay, por isso os diálogos contém nuances psicológicas, cenários diversos e personagens com estilos de falas e motivações complexas. Modos de Geração… See the full description on the dataset page: https://huggingface.co/datasets/wilsondesouza/aurora-dataset-roleplay-ptbr.texttext-generation1K<n<10K0 likes21 downloads10mo agoHugging Face14costadev00 /wikipedia-pt-br-instructions-sft wikipedia-pt-br-instructions-gemma-alpaca Dataset de Instruction Following em formato Alpaca puro, com exatamente os campos instruction, input e output. Origem Derivado do dataset local instruction_following produzido por wiki-if-builder, por sua vez derivado de costadev00/wikipedia-pt-br-extract. Campos instruction: comando em português brasileiro. input: contexto mínimo opcional. output: resposta esperada. Licença e limitações A licença herdada é… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions-sft.texttext-generationn<1K0 likes20 downloads5mo agoHugging Face15divergente /wikitext-ptbr-1texttext-generation10K<n<100K0 likes18 downloads3y agoHugging Face16Boakpe /pt-br-agentic-text-to-sql-distilled-trajectories PT-BR Agentic Text-to-SQL Distilled Trajectories This dataset contains message-only distilled trajectories for training tool-using Text-to-SQL agents in Brazilian Portuguese. The trajectories were selected from LLM-judged correct conversations and preserve the agent protocol used in the released code. Code and reproducibility repository: https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/pt-br-agentic-text-to-sql-distilled-trajectories.text-generation1K<n<10K1 likes16 downloads3mo agoHugging Face17br-llm-data /wikipedia-pt-br-instruct-5kgated Wikipedia PT-BR Instruct wikipedia-pt-br-instruct is a synthetic supervised fine-tuning (SFT) dataset in Brazilian Portuguese generated from Wikipedia-derived documents. This release is an intermediate evaluation dataset produced with the sft-dataset-creator pipeline from the run wiki-ptbr-extract-calib-5kdocs-14tasks. It was generated from a fixed revision of costadev00/wikipedia-pt-br-extract: cdbd07dc4a3de6e64632c718710b3ae0ebaeb0ff The dataset is intended for intermediate… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instruct-5k.texttext-generation100K<n<1M1 likes16 downloads3mo agoHugging Face18Hub-Ai /ptbr-human-preferences 🇧🇷 HUBX Human Preference Dataset (PT-BR) The largest Portuguese-Brazilian human preference dataset for RLHF/DPO training. 📊 Dataset Statistics Metric Value Total Annotations 314,757 Unique Tasks 450 Human Annotators ~600 Avg. Votes per Task ~699 Language Portuguese (Brazil) Domain Communication Quality & Tone 🎯 Why This Dataset? 🇧🇷 Native PT-BR: Collected from Brazilian Portuguese speakers - not translated 👥 Real Humans:… See the full description on the dataset page: https://huggingface.co/datasets/Hub-Ai/ptbr-human-preferences.texttext-classificationn<1K0 likes15 downloads10mo agoHugging Face19br-llm-data /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes6 downloads3mo agoHugging Face20costadev00 /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.