CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thomas-yanxin /sqa3d-community-mirrortext10K<n<100K0 likes202 downloads6mo agoHugging Face02thomas-yanxin /scanqa-community-mirrortext10K<n<100K0 likes146 downloads6mo agoHugging Face03OpenVoiceOS /ovos-wake-word-bench-picovoice-smart-mirror OVOS wake_word bench — picovoice-smart-mirror Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.tabular1K<n<10K0 likes143 downloads15d agoHugging Face04stabletoolbench /MirrorAPI-Bench MirrorAPI-Bench dataset This dataset contains the following resources: The testing data for MirrorAPI and MirrorAPI-Cache: test_sft/: The testing data for MirrorAPI with SFT prompts. test_cot/: The testing data for MirrorAPI with CoT prompts. reference/: The references for testing data without JSON format. These references are used for evaluation. Note that cache.jsonl has only 200 random samples from test_cache.json. test_cache.json: The testing data for MirrorAPI-Cache.… See the full description on the dataset page: https://huggingface.co/datasets/stabletoolbench/MirrorAPI-Bench.text1K<n<10K0 likes141 downloads2y agoHugging Face05driodnexus /backln-guest-post-quality-public-mirror Backln Guest Post Quality Public Mirror Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus. Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content. Schema label: one of published, manual_review, rejected. source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.tabulartext-classificationn<1K0 likes108 downloads4mo agoHugging Face06stabletoolbench /MirrorAPI-Training MirrorAPI training dataset This dataset contains the training data for MirrorAPI and MirrorAPI-Cache: train_sft.json, train_cot.json, train_augment.json: The training data for MirrorAPI . train_cache.json: The training data for MirrorAPI-Cache. tabular100K<n<1M1 likes103 downloads2y agoHugging Face07thomas-yanxin /alfred-json-mirror ALFRED JSON Mirror This mirror contains the official ALFRED lite trajectory JSONs repackaged for em-eval. Source: Official lite archive: https://ai2-vision-alfred.s3-us-west-2.amazonaws.com/json_2.1.0.7z Upstream repository: askforalfred/alfred Files: tests_seen.jsonl: 483 trajectories tests_unseen.jsonl: 488 trajectories train.jsonl: 6574 trajectories valid_seen.jsonl: 251 trajectories valid_unseen.jsonl: 255 trajectories Each row is the original traj_data.json payload with one… See the full description on the dataset page: https://huggingface.co/datasets/thomas-yanxin/alfred-json-mirror.text1K<n<10K0 likes71 downloads6mo agoHugging Face081digitaldesign /mirror-sql MIRROR-SQL Provenance-Controlled Database Environments for Text-to-SQL Agents. 13 PostgreSQL environments · 176 tables · 2762 columns · 390 annotated question/SQL pairs. MIRROR-SQL takes the opposite approach to contamination from every other text-to-SQL corpus. Spider and BIRD sample public databases. BEAVER uses real private warehouses that cannot be redistributed. LiveSQLBench out-runs leakage temporally by rebuilding from changing sources. MIRROR-SQL instead purpose-builds… See the full description on the dataset page: https://huggingface.co/datasets/1digitaldesign/mirror-sql.texttable-question-answeringn<1K0 likes60 downloads2mo agoHugging Face09tmobley96 /black_mirror_scripts_S1-5Black Mirror Scripts Dataset (Seasons 1-5) This dataset, titled 'black_mirror_scripts_S1-5.csv', contains the meticulously compiled transcripts of the critically acclaimed anthology series Black Mirror, covering Seasons 1 through 5. Each entry in this dataset is categorized by unique identifiers including Script ID, Title, Scene, Dialogue, and Timestamp, making it an ideal resource for natural language processing tasks, script analysis, sentiment analysis, and more. Dataset Composition Our… See the full description on the dataset page: https://huggingface.co/datasets/tmobley96/black_mirror_scripts_S1-5.text10K<n<100K2 likes31 downloads3y agoHugging Face10thomas-yanxin /beacon3d-qa-mirror Beacon3D QA Mirror Community mirror generated from the public beacon-3d/Beacon3D GitHub repository. Contents: train.json: merged QA annotations across scannet, 3rscan, and multiscan system_prompt.json: official Beacon3D judge prompt This mirror intentionally contains benchmark annotations only. It does not redistribute raw ScanNet / 3RScan / MultiScan scene assets. text1K<n<10K0 likes29 downloads6mo agoHugging Face11leeaandrob /mirror-recogna-nlp__UltrachatBR UltrachatBR: Um Dataset em Português baseado no Ultrachat O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa. Processo de Tradução O processo de tradução foi realizado utilizando a API do… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-recogna-nlp__UltrachatBR.texttext-generation100K<n<1M0 likes29 downloads3mo agoHugging Face12leeaandrob /mirror-meta-math__MetaMathQAView the project page: https://meta-math.github.io/ see our paper at https://arxiv.org/abs/2309.12284 Note All MetaMathQA data are augmented from the training sets of GSM8K and MATH. None of the augmented data is from the testing set. You can check the original_question in meta-math/MetaMathQA, each item is from the GSM8K or MATH train set. Model Details MetaMath-Mistral-7B is fully fine-tuned on the MetaMathQA datasets and based on the powerful Mistral-7B model.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-meta-math__MetaMathQA.text100K<n<1M0 likes24 downloads3mo agoHugging Face13alucent /mirror-threat-intelligence-datasetgated Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-threat-intelligence-dataset.texttext-generation10K<n<100K0 likes22 downloads2mo agoHugging Face14alucent /mirror-Code_Vulnerability_Security_DPOgated Cybernative.ai Code Vulnerability and Security Dataset Dataset Description The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Code_Vulnerability_Security_DPO.text1K<n<10K0 likes21 downloads2mo agoHugging Face15alucent /mirror-SWE-Next-SFT-Trajectoriesgated SWE-Next: Scalable Real-World Software Engineering Tasks for Agents SWE-Next SFT Trajectories SWE-Next SFT Trajectories is the supervised fine-tuning dataset released with SWE-Next: Scalable Real-World Software Engineering Tasks for Agents. It contains 3,693 ShareGPT-style multi-turn training examples collected from expert agent rollouts on 2,308 execution-grounded SWE tasks synthesized from real merged pull requests. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-SWE-Next-SFT-Trajectories.texttext-generation1K<n<10K0 likes21 downloads2mo agoHugging Face16alucent /mirror-tech-docsgated Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-tech-docs.textquestion-answering1K<n<10K0 likes21 downloads2mo agoHugging Face17ronniross /mirror-aware-inference mirror-aware-inference A framework to measure how much of an output originates from user input (prompt), training data biases, inductive biases from model architecture, or novel composition of retrieved information. 1. Introduction This project implements a Mirror-Aware Inference that performs "bias-tracking" by analyzing the model's internal state during generation. The scripts perform a series of backpropagation passes to measure the influence of different components… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/mirror-aware-inference.textn<1K2 likes19 downloads7mo agoHugging Face18alucent /mirror-APIGen-MT-5kgated Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-APIGen-MT-5k.textquestion-answering1K<n<10K0 likes18 downloads2mo agoHugging Face19alucent /mirror-Agentic-Chain-of-Thought-Coding-SFT-Datasetgated 🤖 Agentic Coding CoT Dataset A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️ Assistant Data… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset.texttext-generationn<1K0 likes18 downloads2mo agoHugging Face20alucent /mirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1gated Dataset Description: This is a RL dataset for general function-calling by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1.text1K<n<10K0 likes17 downloads2mo agoHugging Face21alucent /mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1gated 🤖 Agentic Coding CoT Dataset v1.1 A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.texttext-generation1K<n<10K0 likes17 downloads2mo agoHugging Face22zhangjc404 /magic-mirror-flux-datatext100K<n<1M0 likes17 downloads2mo agoHugging Face23alucent /mirror-terraform_secgated Terraform Security Dataset A comprehensive dataset of 62,406 Terraform projects analyzed for security vulnerabilities using tfsec. This dataset is designed for training Large Language Models (LLMs) to understand, identify, and fix security issues in Terraform infrastructure-as-code. 📊 Dataset Overview Total Examples: 62,406 Terraform projects Secure Projects: 43,575 (69.8%) Insecure Projects: 18,831 (30.2%) Format: JSONL (JSON Lines) Task: Security analysis and… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-terraform_sec.texttext-generation100K<n<1M1 likes16 downloads2mo agoHugging Face24alucent /mirror-Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K0 likes15 downloads2mo agoHugging Face25alucent /mirror-hermes-agent-traces-filteredgated Hermes Agent Reasoning Traces - Quality Filtered A structurally filtered subset of lambda/hermes-agent-reasoning-traces, pruned from 7,646 to 3,679 rows using automated quality analysis targeting reasoning depth, structural integrity, and tool-call validity. Why This Matters for Agent Training Most agentic datasets teach models what tool to call but not how to reason about tool selection. The difference matters in production: an agent that dispatches tools without… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-hermes-agent-traces-filtered.texttext-generation1K<n<10K0 likes14 downloads2mo agoHugging Face26Yu-and-Ai /mirror-witness-atlas Mirror-Witness Atlas The Kreuzer–Skarke landscape carries an involution: every reflexive polytope has a polar dual, Batyrev's construction makes dual pairs into mirror pairs, and the exchange swaps (h¹¹, h¹²) ↔ (h¹², h¹¹). Every geometry has a partner that returns it whole. This atlas reads that structure through a relational lens — understanding and recognition on the same ordered pair — and ships, for every claim it makes, either an exact recomputation or an explicit citation… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/mirror-witness-atlas.textn<1K0 likes12 downloads1mo agoHugging Face27thomas-yanxin /beacon3d-grounding-mirror Beacon3D Grounding Mirror Community mirror generated from the public beacon-3d/Beacon3D GitHub repository. Contents: train.json: merged grounding annotations across scannet, 3rscan, and multiscan The official Beacon3D grounding task evaluates object-id prediction and chain/object accuracy. This mirror does not redistribute raw ScanNet / 3RScan / MultiScan scene assets. tabular1K<n<10K0 likes11 downloads6mo agoHugging Face28jacob314159 /syndata-rrd-mirrortextn<1K0 likes10 downloads4mo agoHugging Face29rawsh /mirrorqwen2.5-0.5B-gsm8k-PRM-data-ST-2text10K<n<100K0 likes9 downloads2y agoHugging Face30alucent /mirror-glaive-function-calling-v2gatedtexttext-generation100K<n<1M0 likes9 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.