CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tegridydev /infosec-tool-output Infosec Tool Output Security-tool output → evidence-backed, plain-English interpretation. A dataset for training and evaluating models that interpret security-tool output, explain the limits of the evidence, and recommend defensive next steps. v2.0.0: 1,004 canonical examples across 19 tools. This includes all 776 original records with traceable interpretation changes, plus 228 newly authored synthetic fixtures. The deduplicated training views contain 1004 examples, not… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/infosec-tool-output.texttext-generation1K<n<10K3 likes289 downloads18d agoHugging Face02hugfaceguy0001 /simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing. tabulartext-classificationn<1K0 likes158 downloads3y agoHugging Face03Mahadih534 /Institutional-Information-of-Bangladesh Institutional-Information-of-Bangladesh Dataset This Dataset contains all verified and authorized Institutional information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.tabularquestion-answering10K<n<100K2 likes114 downloads2y agoHugging Face04hkust-nlp /dart-math-pool-gsm8k-query-info [!NOTE] This dataset is the synthesis information of queries from the GSM8K training set, such as the numbers of raw/correct samples of each synthesis job. Usually used with dart-math-pool-gsm8k. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-gsm8k-query-info.tabulartext-generation1K<n<10K2 likes78 downloads2y agoHugging Face05Neura-parse /quantum-information-and-complexity-theory Neura Parse — Quantum Information & Complexity Theory: Channels, Entropies, Classes & the Structure of Advantage A proof-based theoretical-foundations vertical uniting quantum information theory (channels, entropies, entanglement measures, distinguishability, capacities, Shannon theory) with quantum complexity theory and the structure of quantum advantage (classes, Hamiltonian complexity, sampling-based advantage and its verification, pseudorandomness, dequantization).… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-information-and-complexity-theory.tabulartext-generation100K<n<1M0 likes75 downloads3mo agoHugging Face06StarpowerTechnology /Dense-Information-Science-Physics-Dataset Dense Information With Multiple Fine-tuned Variations This dataaset has multiple for each input to learn how to express the same answer in different ways Dataset Structure The dataset contains two columns: Column Description input A science or quantum-physics question output A conversational answer to the question Example: { "input": "What is quantum entanglement?", "output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.texttext-generation1K<n<10K0 likes74 downloads15d agoHugging Face07hkust-nlp /dart-math-pool-math-query-info [!NOTE] This dataset is the synthesis information of queries from the MATH training set, such as the numbers of raw/correct samples of each synthesis job. Usually used with dart-math-pool-math. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math-query-info.tabulartext-generation1K<n<10K0 likes72 downloads2y agoHugging Face08TorpedoSoftware /roblox-info-dump Roblox-Info-Dump The Roblox-Info-Dump dataset is a collection of public Roblox documentation from create.roblox.com/docs/ and luau.org. Roblox maintains the copyright on all content. texttext-generation10K<n<100K0 likes69 downloads1y agoHugging Face09Lots-of-LoRAs /task684_online_privacy_policy_text_information_type_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task684_online_privacy_policy_text_information_type_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task684_online_privacy_policy_text_information_type_generation.texttext-generation1K<n<10K1 likes59 downloads2y agoHugging Face10Bendang /Informal-Standard-English-Corpus Dataset Description This dataset is a parallel corpus of approximately 11,000 pairs of informal conversational English text and their normalized equivalents. The informal text mimics real-world digital communication, featuring slang, phonetic spellings, missing punctuation, and abbreviations. The normalized text provides a grammatically correct and semantically equivalent version. The dataset was created to support machine translation tasks for low-resource languages. It… See the full description on the dataset page: https://huggingface.co/datasets/Bendang/Informal-Standard-English-Corpus.texttranslation10K<n<100K1 likes56 downloads2mo agoHugging Face11leeroy-jankins /DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services 🗺️ DoD Installation Geospatial Information and Services Question-Answer Dataset Source: DoD Instruction 8130.01 Source Effective Date: April 9, 2015 Change Incorporated: Change 3, effective August 4, 2020 Source Organization: Office of the Under Secretary of Defense for Acquisition and Sustainment Source Ownership: United States Department of Defense 📋 Overview Dataset Summary The DoD Installation Geospatial Information and Services… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services.documentquestion-answering0 likes55 downloads2mo agoHugging Face12leeroy-jankins /DOD-Instruction-5040-02-Visual-Information DoD Visual Information Question-Answer Dataset Maintainer: Terry Eppler Owner: US Federal Government Dataset Summary This dataset contains 250 document-grounded question-and-answer records based on DoD Instruction 5040.02, “Visual Information (VI),” dated October 27, 2011, and incorporating Change 2 effective April 20, 2018. The source establishes Department of Defense policy, responsibilities, and procedures for managing visual-information records… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Instruction-5040-02-Visual-Information.documentquestion-answering0 likes45 downloads2mo agoHugging Face13leeroy-jankins /DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging 📚 DoD Instruction 8170.01 Online Information Management and Electronic Messaging Maintainer: Terry Eppler Ownership: U.S. Department of Defense 📋 Overview Dataset Summary The DoD Instruction 8170.01 Online Information Management and Electronic Messaging Dataset is a structured natural-language question-answering dataset derived from DoD Instruction 8170.01, Online Information Management and Electronic Messaging. DoD Instruction 8170.01… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging.documentquestion-answering0 likes45 downloads2mo agoHugging Face14rhaymison /medicine-information-pttexttext-generation1K<n<10K1 likes44 downloads3y agoHugging Face15Reza2kn /uncgpt-conversations-informal-approved-1p25 UncGPT — Informal-Register Approved Conversations that passed the 1.25σ semantic gate AND the current strict programmatic gates — including intimate-register (tú-not-usted, tu-not-shoma, 你-not-您, no po/opo, plain not keigo), stricter colloquial Persian, and strict completion-integrity. Part of the UncGPT NeurIPS 2026 Competition collection. Counts approved: 753 rejected: 1,475 skills covered: 53 of 69 by care: warm 450 / mid 152 / cold 151 by language: en 310 / sw… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-informal-approved-1p25.texttext-generation1K<n<10K0 likes39 downloads4mo agoHugging Face16leeroy-jankins /DoD-Instruction-8010-01-Information-Network-Transport 🌐 DoD Information Network Transport Maintainer: Terry Eppler Owner: US Federal Government Source: DoD Instruction 8010.01 Dataset Size: question-answer records Source Effective Date: September 10, 2018 Source Organization: Office of the DoD Chief Information Officer Source Ownership: United States Department of Defense 📋 Overview Dataset Summary The DoD Information Network Transport Question-Answer Dataset contains document-grounded… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8010-01-Information-Network-Transport.documentquestion-answering0 likes34 downloads2mo agoHugging Face17InfoBayAI /Legacy-Code-Datasetgated Legacy Codebase Dataset Dataset Description The Legacy Codebase Dataset is a large-scale collection of enterprise software repositories designed for training next-generation Large Language Models (LLMs), AI coding assistants, software engineering copilots, automated refactoring systems, repository understanding models, and intelligent program analysis pipelines. The complete collection contains 405 real-world legacy codebases spanning 23 major industries… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Legacy-Code-Dataset.texttext-generation10K<n<100K0 likes32 downloads8d agoHugging Face18TigreGotico /infopedia-pt-ipa European Portuguese IPA Lexicon — Infopédia A lightweight word → IPA pronunciation lexicon for European Portuguese, extracted from Infopédia (Porto Editora). One row per headword, intended for grapheme-to-phoneme (G2P) work, pronunciation modelling, and TTS/ASR lexicon building. Complete crawl. Derived from a graph crawl of Infopédia that ran to convergence (frontier → 0), covering the dictionary's reachable component. Contents Field Count Entries… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/infopedia-pt-ipa.texttext-generation100K<n<1M0 likes31 downloads3mo agoHugging Face19leeroy-jankins /DoD-Instruction-5200-01-Information-Security-Program DoD Information Security and SCI Protection Question-Answer Dataset Maintainer: Terry Eppler Owner: US Federal Government Dataset Summary This dataset contains document-grounded question-and-answer records based on DoD Instruction 5200.01, “DoD Information Security Program and Protection of Sensitive Compartmented Information (SCI),” dated April 21, 2016, and incorporating Change 2 effective October 1, 2020. The source establishes the overarching Department… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5200-01-Information-Security-Program.documentquestion-answering0 likes31 downloads2mo agoHugging Face20Infomaniak-AI /speculators-multilingual-en-fr-de-it-es Speculators Multilingual SFT Dataset (en/fr/de/it/es) A multilingual instruction-following dataset in ShareGPT format, built to train draft models for speculative decoding across English, French, German, Italian and Spanish. Summary An English instruction-tuning corpus with part of it kept in English and the rest machine-translated into French, German, Italian and Spanish using tencent/Hunyuan-MT-7B. Provided as a single mixed-language, ShareGPT-formatted dataset… See the full description on the dataset page: https://huggingface.co/datasets/Infomaniak-AI/speculators-multilingual-en-fr-de-it-es.texttext-generation100K<n<1M0 likes30 downloads1mo agoHugging Face21amadzarak /synthetic-info-extract-json Raw text to json object (synthetic) Amad Zarak February 28, 2026 Created using gpt-oss-120b on h200 sxm Zero-shot JSON schema deduction & universal information extraction. 80,664 rows Figured others could use this since it is basically impossible to find massive raw-text-to-structured-json datasets for training extraction engines. About 30k of the raw outputs hit the token limit and malformed, but I ran a massive salvage sweep on the raw outputs using json-repair to force the… See the full description on the dataset page: https://huggingface.co/datasets/amadzarak/synthetic-info-extract-json.texttext-generation10K<n<100K0 likes29 downloads7mo agoHugging Face22astral-expmath /agda-categories-informalized agda-categories, informalized 4,541 declarations from the agda-categories library, each paired with an informal, LaTeX-flavoured natural-language statement written by GLM-5.2. The natural language is written to be precise enough to re-formalise from, so the intended use is training a model to reconstruct the formal Agda source from the prose alone. Declarations were extracted with a fork of Agda that dumps one JSON record per named declaration (with its full source range) during… See the full description on the dataset page: https://huggingface.co/datasets/astral-expmath/agda-categories-informalized.tabulartext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face23leeroy-jankins /DOD-Directive-8000-01-Management-Of-Defense-Information DoD Directive 8000.01 Management of the DoD Information Enterprise Question-Answer Dataset Maintainer: Terry Eppler Owner: US Federal Government Dataset Summary This dataset contains document-grounded question-and-answer records based on Department of Defense Directive 8000.01, “Management of the Department of Defense Information Enterprise,” dated March 17, 2016, and incorporating Change 1 effective July 27, 2017. The directive establishes Department-wide… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-8000-01-Management-Of-Defense-Information.documentquestion-answering0 likes29 downloads2mo agoHugging Face24akiFQC /japanese-confidential-information-extraction-sft Japanese Confidential Information Extraction — SFT Dataset 日本語テキストから社外秘の固有表現を抽出するタスク向けの SFT (Supervised Fine-Tuning) データセットです。 LFM2 系モデルの LoRA fine-tune を想定して構築されています。 タスク概要 入力テキスト(日本語)に含まれる機密情報を、11カテゴリの JSON として抽出します。 入力: 「山田太郎(yamada@example.co.jp)から請求書番号 INV-2024-0042 で 売上 ¥12,800,000 の見積書が届いた。」 出力: { "address": [], "company_name": [], "email_address": ["yamada@example.co.jp"], "human_name": ["山田太郎"], "phone_number": [], "account_identifier":… See the full description on the dataset page: https://huggingface.co/datasets/akiFQC/japanese-confidential-information-extraction-sft.texttext-generation10K<n<100K2 likes28 downloads4mo agoHugging Face25davidquicast /information-security-policies-qa-distiset Dataset Card for information-security-policies-qa-distiset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/daqc/information-security-policies-qa-distiset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/davidquicast/information-security-policies-qa-distiset.tabulartext-generationn<1K0 likes25 downloads2y agoHugging Face26InfoBayAI /DSA-Coding-Problems-and-Solutions-Datasetgated Dataset Description This dataset is a large-scale collection of Data Structures and Algorithms (DSA) code, containing 12,385 code files with 3.86 million lines of code and 25.01 million lexical tokens, designed to support the development of advanced code generation models, programming assistants, software engineering AI systems, and code intelligence applications. It consists of real-world DSA implementations covering a wide range of algorithms, data structures, problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/DSA-Coding-Problems-and-Solutions-Dataset.tabulartext-generationn<1K0 likes23 downloads9d agoHugging Face27Lots-of-LoRAs /task1284_hrngo_informativeness_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1284_hrngo_informativeness_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1284_hrngo_informativeness_classification.texttext-generation1K<n<10K0 likes22 downloads2y agoHugging Face28zaakirio /infosec_harmful_behaviors Infosec Harmful Behaviors Offensive-security instruction prompts for refusal-direction research and abliteration of code/security models. Dataset Details This dataset contains infosec-domain harmful prompts intended to elicit refusal behavior from aligned instruction models. It is designed as the harmful side of a harmful/harmless contrast pair, analogous to mlabonne/harmful_behaviors but focused on offensive-security and malicious-coding requests. Rows: train:… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/infosec_harmful_behaviors.texttext-generationn<1K1 likes21 downloads3mo agoHugging Face29infosense /frameref Dataset Information Information ecosystems increasingly shape how people internalize exposure to adverse digital experiences, raising concerns about the long-term consequences for information health. In modern search and recommendation systems, ranking and personalization policies play a central role in shaping such exposure and its long-term effects on users. To study these effects in a controlled setting, we present FrameRef, a large-scale dataset of 1,073,740 systematically… See the full description on the dataset page: https://huggingface.co/datasets/infosense/frameref.texttext-generation1M<n<10M0 likes18 downloads7mo agoHugging Face30tandevllc /offsec_redteam_infogated OffSec RedTeam Info OffSec RedTeam Info is a SlimPajama‑style, category‑organized corpus of security knowledge text crawled from reputable red‑team/blue‑team websites: wikis, training blogs, vendor research, CERT advisories, reversing/malware labs, cloud/kubernetes posts, OSINT handbooks, AD tradecraft, and more. Token count: ~1.646B tokens. ⚠️ Ethical use only. Use for research, education, and defensive security. Respect robots.txt, site terms, and copyrights. Do not misuse this… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_info.texttext-generation1M<n<10M3 likes16 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.