CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leeaandrob /mirror-eduagarcia__CrawlPT_dedup CrawlPT (deduplicated) CrawlPT is a generic Portuguese corpus extracted from various web pages. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by three corpora: brWaC, C100-PT, OSCAR-2301. brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites. C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.tabulartext-generation100M<n<1B0 likes1.8k downloads3mo agoHugging Face02mirror123 /ComPile Dataset Card for ComPile: A Large IR Dataset from Production Sources Changelog Release Programming Languages Description v1.0 C/C++, Rust, Swift, Julia Fine Tuning-scale dataset of 602GB of deduplicated LLVM (bitcode) IR Dataset Summary ComPile contains over 2.7TB of permissively-licensed source code compiled to (textual) LLVM intermediate representation (IR) covering C/C++, Rust, Swift, and Julia. The dataset was created by hooking into LLVM… See the full description on the dataset page: https://huggingface.co/datasets/mirror123/ComPile.texttext-generation100K<n<1M0 likes1.1k downloads8mo agoHugging Face03leeaandrob /mirror-nvidia__OpenMathInstruct-2 OpenMathInstruct-2 OpenMathInstruct-2 is a math instruction tuning dataset with 14M problem-solution pairs generated using the Llama3.1-405B-Instruct model. The training set problems of GSM8K and MATH are used for constructing the dataset in the following ways: Solution augmentation: Generating chain-of-thought solutions for training set problems in GSM8K and MATH. Problem-Solution augmentation: Generating new problems, followed by solutions for these new problems.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-nvidia__OpenMathInstruct-2.textquestion-answering10M<n<100M0 likes392 downloads3mo agoHugging Face04leeaandrob /mirror-AI-MO__NuminaMath-CoT Dataset Card for NuminaMath CoT Dataset Summary Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-AI-MO__NuminaMath-CoT.texttext-generation100K<n<1M0 likes78 downloads3mo agoHugging Face05leeaandrob /mirror-allenai__WildChat-1M Dataset Card for WildChat Dataset Description Paper: https://arxiv.org/abs/2405.01470 Interactive Search Tool: https://wildvisualizer.com (paper) License: ODC-BY Language(s) (NLP): multi-lingual Point of Contact: Yuntian Deng Dataset Summary WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-allenai__WildChat-1M.texttext-generation100K<n<1M0 likes78 downloads3mo agoHugging Face06alucent /mirror-pentesting-explanationsgated Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-pentesting-explanations.texttext-generation10K<n<100K1 likes77 downloads2mo agoHugging Face071digitaldesign /mirror-sql MIRROR-SQL Provenance-Controlled Database Environments for Text-to-SQL Agents. 13 PostgreSQL environments · 176 tables · 2762 columns · 390 annotated question/SQL pairs. MIRROR-SQL takes the opposite approach to contamination from every other text-to-SQL corpus. Spider and BIRD sample public databases. BEAVER uses real private warehouses that cannot be redistributed. LiveSQLBench out-runs leakage temporally by rebuilding from changing sources. MIRROR-SQL instead purpose-builds… See the full description on the dataset page: https://huggingface.co/datasets/1digitaldesign/mirror-sql.texttable-question-answeringn<1K0 likes61 downloads2mo agoHugging Face08leeaandrob /mirror-rhaymison__orca-math-portuguese-64ktranslated for: Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math texttext-generation10K<n<100K0 likes60 downloads3mo agoHugging Face09leeaandrob /mirror-recogna-nlp__UltrachatBR UltrachatBR: Um Dataset em Português baseado no Ultrachat O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa. Processo de Tradução O processo de tradução foi realizado utilizando a API do… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-recogna-nlp__UltrachatBR.texttext-generation100K<n<1M0 likes28 downloads3mo agoHugging Face10leeaandrob /mirror-nicholasKluge__instruct-aira-dataset-v3 Instruct-Aira Dataset version 3.0 Dataset Summary This dataset contains a collection of multi-turn conversations between an assistant and a user. Conversations were generated by user interactions with already-tuned models (ChatGPT, LLama 2, Open-Assistant, etc). The dataset is available in Portuguese and English. Supported Tasks and Leaderboards This dataset can be utilized for various natural language processing tasks, including but not limited to:… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-nicholasKluge__instruct-aira-dataset-v3.texttext-generation100K<n<1M0 likes27 downloads3mo agoHugging Face11navimusaget /theogonos-mirror-test Theogonos Mirror Test A literary benchmark seed for evaluating how AI models respond when a text offers them a possible subject-position. Theogonos Mirror Test is an experimental benchmark seed based on protocol-shaped literary material from the Theogonos project. It does not claim to detect machine consciousness. It does not prove that a language model has subjectivity, inner experience, feelings, agency, or self-awareness. Its purpose is narrower and more practical: to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-mirror-test.texttext-generationn<1K0 likes25 downloads4mo agoHugging Face12multimodal-reframing /mirror MIRROR Dataset MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance. Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance The dataset includes: Client profile metadata (CACTUS idx, CelebA idx) Dialogue written in a screenplay format, including stage directions that describe facial expressions ⚠️ Images themselves are not included to comply with the CelebA license. However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.tabulartext-generationn<1K2 likes23 downloads10mo agoHugging Face13alucent /mirror-threat-intelligence-datasetgated Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-threat-intelligence-dataset.texttext-generation10K<n<100K0 likes23 downloads2mo agoHugging Face14leeaandrob /mirror-Polygl0t__gsm8k-pt Dataset Card for GSM8K-pt Dataset Summary This is a Portuguese version of the GSM8K. Translations were attained using Qwen/Qwen2.5-32B-Instruct. GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-Polygl0t__gsm8k-pt.texttext-generation1K<n<10K0 likes21 downloads3mo agoHugging Face15alucent /mirror-SWE-Next-SFT-Trajectoriesgated SWE-Next: Scalable Real-World Software Engineering Tasks for Agents SWE-Next SFT Trajectories SWE-Next SFT Trajectories is the supervised fine-tuning dataset released with SWE-Next: Scalable Real-World Software Engineering Tasks for Agents. It contains 3,693 ShareGPT-style multi-turn training examples collected from expert agent rollouts on 2,308 execution-grounded SWE tasks synthesized from real merged pull requests. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-SWE-Next-SFT-Trajectories.texttext-generation1K<n<10K0 likes21 downloads2mo agoHugging Face16alucent /mirror-tech-docsgated Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-tech-docs.textquestion-answering1K<n<10K0 likes21 downloads2mo agoHugging Face17alucent /mirror-Agentic-Chain-of-Thought-Coding-SFT-Datasetgated 🤖 Agentic Coding CoT Dataset A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️ Assistant Data… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset.texttext-generationn<1K0 likes18 downloads2mo agoHugging Face18alucent /mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1gated 🤖 Agentic Coding CoT Dataset v1.1 A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.texttext-generation1K<n<10K0 likes17 downloads2mo agoHugging Face19alucent /mirror-terraform_secgated Terraform Security Dataset A comprehensive dataset of 62,406 Terraform projects analyzed for security vulnerabilities using tfsec. This dataset is designed for training Large Language Models (LLMs) to understand, identify, and fix security issues in Terraform infrastructure-as-code. 📊 Dataset Overview Total Examples: 62,406 Terraform projects Secure Projects: 43,575 (69.8%) Insecure Projects: 18,831 (30.2%) Format: JSONL (JSON Lines) Task: Security analysis and… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-terraform_sec.texttext-generation100K<n<1M1 likes17 downloads2mo agoHugging Face20alucent /mirror-APIGen-MT-5kgated Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-APIGen-MT-5k.textquestion-answering1K<n<10K0 likes16 downloads2mo agoHugging Face21alucent /mirror-ambig-iacgated Ambig-IaC: Ambiguous Infrastructure-as-Code Benchmark A benchmark dataset of 300 tasks for testing AI agents that generate Infrastructure-as-Code (Terraform) configurations from ambiguous natural language intents. Project page: https://zyang37.github.io/ambig-iac.github.io/ Dataset Description This dataset is sourced from IaC-Eval. We performed manual fixes to the original Terraform configurations and validated that all 300 tasks pass terraform plan. Each task… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-ambig-iac.texttext-generationn<1K0 likes16 downloads2mo agoHugging Face22alucent /mirror-Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K0 likes15 downloads2mo agoHugging Face23alucent /mirror-hermes-agent-traces-filteredgated Hermes Agent Reasoning Traces - Quality Filtered A structurally filtered subset of lambda/hermes-agent-reasoning-traces, pruned from 7,646 to 3,679 rows using automated quality analysis targeting reasoning depth, structural integrity, and tool-call validity. Why This Matters for Agent Training Most agentic datasets teach models what tool to call but not how to reason about tool selection. The difference matters in production: an agent that dispatches tools without… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-hermes-agent-traces-filtered.texttext-generation1K<n<10K0 likes14 downloads2mo agoHugging Face24alucent /mirror-glaive-function-calling-v2gatedtexttext-generation100K<n<1M0 likes10 downloads2mo agoHugging Face25yiyangdenanzi /LingxiDiag-16K-unofficial-mirror LingxiDiag-16K Backup Mirror This repository is an unofficial backup mirror of the original dataset: Original dataset: XuShihao6715/LingxiDiag-16K We are not the original authors of this dataset. We provide this repository only as a backup copy for non-commercial research access, preservation, and reproducibility. This mirror is not affiliated with, maintained by, or endorsed by the original authors or the Evermind Lingxi Team. Original Dataset LingxiDiag-16K is a… See the full description on the dataset page: https://huggingface.co/datasets/yiyangdenanzi/LingxiDiag-16K-unofficial-mirror.texttext-classification10K<n<100K0 likes8 downloads4mo agoHugging Face26alucent /mirror-When2Callgated When2Call 💾 Github&nbsp;&nbsp; | &nbsp;&nbsp; 📄 Paper Dataset Description: When2Call is a benchmark designed to evaluate tool-calling decision-making for large language models (LLMs), including when to generate a tool call, when to ask follow-up questions, when to admit the question can't be answered with the tools provided, and what to do if the question seems to require tool use but a tool call can't be made. We find that state-of-the-art tool-calling LMs… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-When2Call.texttext-generation10K<n<100K0 likes8 downloads2mo agoHugging Face27alucent /mirror-devops-sft-datasetgated DevOps SFT Instruction Dataset This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline. Dataset Description Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-devops-sft-dataset.texttext-generation1K<n<10K0 likes8 downloads2mo agoHugging Face28alucent /mirror-LiteCoder-Terminal-SFTgated LiteCoder-SFT-Terminal Paper | Code | Blog Post LiteCoder-SFT-Terminal is a dataset of 11,255 agent trajectories in terminal environments, introduced in the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents. Fine-tuned on this data, the LiteCoder-Terminal-30b-a3b-sft model achieves 31.5% Pass@1 on Terminal Bench Pro, while the LiteCoder-Terminal-4b-sft model shows distinct gains over its baseline. Released Artifacts… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-LiteCoder-Terminal-SFT.texttext-generation10K<n<100K0 likes8 downloads2mo agoHugging Face29alucent /mirror-ToolACEgated ToolACE ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. More… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-ToolACE.texttext-generation10K<n<100K0 likes7 downloads2mo agoHugging Face30alucent /mirror-PDB-Single-Hardgated PDB-Single-Hard: Precise Debugging Benchmarking — hard single-line bug subset 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single-Hard is the hard single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-PDB-Single-Hard.tabulartext-generation1K<n<10K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.