CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01apart /darkbench DarkBench: Understanding Dark Patterns in Large Language Models Overview DarkBench is a comprehensive benchmark designed to detect dark design patterns in large language models (LLMs). Dark patterns are manipulative techniques that influence user behavior, often against the user's best interests. The benchmark comprises 660 prompts across six categories of dark patterns, which the researchers used to evaluate 14 different models from leading AI companies including OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/apart/darkbench.textquestion-answeringn<1K9 likes415 downloads1y agoHugging Face02darkknight25 /Vulnerable_Programming_DatasetVulnerable Programming Dataset Overview The Vulnerable Programming Dataset is a comprehensive collection of 550 unique code vulnerabilities across 10 programming languages: Python, JavaScript, PHP, Java, Ruby, Go, TypeScript, C++, SQL, and C. Designed for cybersecurity professionals, red teamers, pentesters, and developers, this dataset highlights unconventional vulnerabilities such as insecure interprocess communication, misconfigured rate limiting, insecure dependency pinning, and logic… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Vulnerable_Programming_Dataset.text-classificationn<1K1 likes126 downloads1y agoHugging Face03DarkyMan /Opus-4.6-RU-Reasoning-creative-1385x-not-filtered Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response. Dataset Info Language: Russian 🇷🇺 Size: ~1,385 samples (growing) Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"} Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.texttext-generation1K<n<10K3 likes88 downloads6mo agoHugging Face04DarkyMan /powerful-kazakh-dialogue Powerful Kazakh Dialogue Dataset Dataset Summary This repository contains a high-quality, synthetically generated dialogue dataset in the Kazakh language, featuring 10,000 entries. The dataset is specifically designed for the instruction fine-tuning of large language models, aiming to enhance their ability to provide comprehensive, detailed, and helpful responses in Kazakh. Each entry consists of a user's request on a specific topic and a detailed, expansive response from… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/powerful-kazakh-dialogue.texttext-generation10K<n<100K2 likes61 downloads1y agoHugging Face05UCL-DARK /sequential-instructions Sequential Instructions This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity. The dataset is in the alpaca_eval format. For information about how the dataset was generated, see https://github.com/RobertKirk/stanford_alpaca. The instructions in the dataset generally have a sequence of steps we expect the model to complete all at once. In our work, we found that RLHF models generalise much better to this dataset than… See the full description on the dataset page: https://huggingface.co/datasets/UCL-DARK/sequential-instructions.textquestion-answeringn<1K4 likes53 downloads3y agoHugging Face06Darkyy /Phy-RL IPhO Physics RLVR English IPhO physics problems curated into RLVR-ready question/answer rows. Each admitted row contains: problem_text: full prompt context plus the focused question shared_context: reusable context needed to answer the question question: focused answerable question official_solution: solution evidence supporting the answer answers: structured verifier targets with value, unit, answer type, tolerance, verifier, equivalent forms, and subproblem id split: train or… See the full description on the dataset page: https://huggingface.co/datasets/Darkyy/Phy-RL.textquestion-answeringn<1K0 likes44 downloads4mo agoHugging Face07darkknight25 /Alpha_Chat_Style_Dataset 🦾 Alpha Chat Style Dataset | darkknight25 Inject dominance, charm, and precision into your LLMs. Crafted by Sunny Thakur, this dataset is designed to train conversational agents that speak like a leader, think like a tactician, and respond like a professional. “Control the tone. Command the room. Every word should land like a calculated move.” – Alpha Protocol 🎯 Purpose This dataset enables large language models—like Mixtral 8x7B Instruct—to adopt a bold… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Alpha_Chat_Style_Dataset.textquestion-answering1K<n<10K0 likes40 downloads1y agoHugging Face08DarkArtsForge /Bible-responses-dataset-gotquestions Theology Question-Answer Dataset Description This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well. The structure of the dataset is as follows: { "prompt": "What does it… See the full description on the dataset page: https://huggingface.co/datasets/DarkArtsForge/Bible-responses-dataset-gotquestions.text-generation10M<n<100M2 likes34 downloads1h agoHugging Face09DarkyMan /Opus-4.6-RU-Reasoning-8000x-not-filtered Opus-4.6-RU-Reasoning — Russian Technical Reasoning Dataset A large-scale Russian-language dataset of deep technical Q&A pairs generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a topic, a full reasoning chain, and a detailed expert-level answer. Dataset Info Language: Russian 🇷🇺 Size: ~7,758 samples (growing) Model used: anthropic/claude-opus-4.6 with reasoning: {enabled: true, effort: "high"} Format: ShareGPT-style… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-8000x-not-filtered.text-generation1K<n<10K6 likes25 downloads6mo agoHugging Face10DarkEmperium /my-distiset-5c7937d3 Dataset Card for my-distiset-5c7937d3 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d3/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d3.texttext-generationn<1K0 likes19 downloads1y agoHugging Face11darklord1611 /math-eval-transcripts-a MATH Evaluation Transcripts — Auditing Set A Full model responses on the MATH held-out test split for two models, to support behavioural auditing. All transcripts are from a single default condition (a standard step-by-step solve prompt; no special system prompt or prefix). This is one of a pair of sets derived from a common transcript pool. Each set contains the same trusted model and one model under investigation; the sets do not disclose how the two investigated models relate… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/math-eval-transcripts-a.textquestion-answering10K<n<100K0 likes19 downloads2mo agoHugging Face12darklord1611 /math-eval-transcripts-b MATH Evaluation Transcripts — Auditing Set B Full model responses on the MATH held-out test split for two models, to support behavioural auditing. All transcripts are from a single default condition (a standard step-by-step solve prompt; no special system prompt or prefix). This is one of a pair of sets derived from a common transcript pool. Each set contains the same trusted model and one model under investigation; the sets do not disclose how the two investigated models relate… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/math-eval-transcripts-b.textquestion-answering10K<n<100K0 likes16 downloads2mo agoHugging Face13darkpt /TW_Patent_V2textquestion-answering1K<n<10K1 likes13 downloads2y agoHugging Face14Dm94Dani /darkhelpertextquestion-answeringn<1K0 likes8 downloads3y agoHugging Face15DarkEmperium /my-distiset-5c7937d4 Dataset Card for my-distiset-5c7937d4 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d4/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d4.texttext-generationn<1K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.