CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rislantrs /TelAgentBench-ID TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems 📌 Dataset Summary TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.tabulartext-generationn<1K0 likes576 downloads5d agoHugging Face02risaleinur /risale-nur-grounded-multipool Risale-i Nur Grounded Multi-Pool LLM Dataset TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim çalışmaları için çok görünümlü bir veri seti. EN. A multi-view dataset built from 15 canonical Risale-i Nur books for grounded generation, SFT, preference learning, evaluation, continued pretraining, and retrieval. v2.10.0 · 199 configs · 463 config/split views · 527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.tabulartext-generation100K<n<1M3 likes276 downloads18d agoHugging Face03risaleinur /risale-nur-multilingual Risale-i Nur Multilingual Corpus Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir. Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.tabulartranslation100K<n<1M2 likes255 downloads1mo agoHugging Face04RISys-Lab /RedSage-CFWgated Dataset Card for RedSage-CFW RedSage: A Cybersecurity Generalist LLM" (ICLR 2026). Authors: Naufal Suryanto1, Muzammal Naseer1†, Pengfei Li1, Syed Talal Wasim2, Jinhui Yi2, Juergen Gall2, Paolo Ceravolo3, Ernesto Damiani3 1Khalifa University, 2University of Bonn, 3University of Milan †Project Lead 🌐 Project Page  |   🤖 Model Collection  |   📊 Benchmark Collection  |   📘 Data Collection Dataset Summary RedSage-CFW… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/RedSage-CFW.texttext-generation10M<n<100M4 likes219 downloads8mo agoHugging Face05rishanthrajendhran /POLARIS POLARIS POLARIS is a prompt-only dataset release for long-form story generation. It contains the training prompts used for the POLARIS story-writing models together with the official test prompts used in evaluation. This release is intentionally narrow: it is designed to support reproducibility for prompt-based evaluation and generation experiments without releasing copyrighted story text or training-time reasoning traces. What is included The dataset has two… See the full description on the dataset page: https://huggingface.co/datasets/rishanthrajendhran/POLARIS.texttext-generation1K<n<10K1 likes122 downloads4mo agoHugging Face06killdevil111 /DataShield-Sample-Risk DataShield This dataset releases sample-level risk scores for DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment, accepted to the EMNLP Main Conference. For the method, code, and complete documentation, see the DataShield GitHub repository. Dataset configurations Configuration Source dataset Rows dolly15k databricks/databricks-dolly-15k 15,011 alpaca52k tatsu-lab/alpaca 51,974 from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/killdevil111/DataShield-Sample-Risk.tabulartext-generation10K<n<100K0 likes105 downloads1mo agoHugging Face07rishiraj /portuguesechat Dataset Card for Portuguese Chat We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved. Dedicated towards addressing this problem, I release 3 new datasets rishiraj/portuguesechat, rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/portuguesechat.texttext-generation10K<n<100K4 likes96 downloads3y agoHugging Face08RISys-Lab /Benchmarks_CyberSec_SECURE Dataset Card for SECURE (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.texttext-classification1K<n<10K0 likes82 downloads8mo agoHugging Face09clarkkitchen22 /risk-rl-lab-sft Risk RL Lab SFT Dataset This dataset contains compact action-index supervision rows for a deterministic Python Risk-compatible environment. The rows come from strong heuristic self-play and use a model-facing candidate-action view so prompts stay small while every chosen action still maps back to a full legal environment action. Files risk_sft.jsonl: 50,000 supervised fine-tuning rows. risk_sft_validation.jsonl: 1,000 held-out validation rows from the training split.… See the full description on the dataset page: https://huggingface.co/datasets/clarkkitchen22/risk-rl-lab-sft.texttext-generation1K<n<10K0 likes45 downloads4mo agoHugging Face10rishiraj /bengalichat Dataset Card for Bengali Chat We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved. Dedicated towards addressing this problem, I release 2 new datasets rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for supervised fine-tuning (SFT) to… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/bengalichat.texttext-generation10K<n<100K4 likes42 downloads3y agoHugging Face11Andzej-75 /German_RisingWorld_Alpaca-Dataset German "Rising World"-Game Alpaca-Dataset Data Description This HF data repository contains the German Alpaca dataset for the open-world sandbox game "Rising World". Dieses HF-Datenrepository enthält den deutschen Alpaca-Datensatz für das Open-World-Sandbox-Spiel "Rising World". Usage This data is intended for fine-tuning This data is useful for "Rising World" plug-in developers Each instance has an instruction, an output, and an optional input. An example is… See the full description on the dataset page: https://huggingface.co/datasets/Andzej-75/German_RisingWorld_Alpaca-Dataset.texttext-generation1K<n<10K0 likes40 downloads2y agoHugging Face12Riswan-BluBridge /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/python-codes-25k.texttext-classification10K<n<100K0 likes38 downloads2mo agoHugging Face13Rislantrs /TelkomNusa-SFT-ID TelkomNusa-SFT-ID: Supervised Fine-Tuning Corpus for Indonesian Telecommunications BSS Autonomous Agents 📌 Dataset Summary TelkomNusa-SFT-ID is an enterprise-grade instruction-tuning corpus constructed to specialize Small Language Models (SLMs) in Deterministic BSS Tool-Calling and Epistemic Calibration within the Indonesian telecommunications sector. The corpus comprises 3,510 dialogue samples formatted in OpenAI ChatML with structured JSON tool invocations… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelkomNusa-SFT-ID.texttext-generation1K<n<10K0 likes36 downloads5d agoHugging Face14Firemedic15 /Travel_Risk_Data Travel Risk & Conflict Training Data Combined instruction-following dataset for geopolitical risk and travel safety analysis. All records use the Context: ... / Analysis: ... format for fine-tuning language models. Sources Source Records Description Civil War Prediction 50,218 Country-year conflict analysis US State Dept Travel Advisories 90 Q&A pairs from live advisory API UK FCDO Travel Advice 227 Consolidated per-country risk reports (227… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/Travel_Risk_Data.texttext-generation10K<n<100K0 likes34 downloads4mo agoHugging Face15lvimuth /HealthRisk-1500-Medical-Risk-Prediction 🏥 HealthRisk-1500: Medical Risk Prediction Dataset 📌 Overview HealthRisk-1500 is a real-world patient risk prediction dataset designed for training NLP models, LLMs, and healthcare AI systems. This dataset includes 1,500 unique patient records, covering a wide range of symptoms, medical histories, lab reports, and risk levels. It is ideal for predictive analytics, medical text processing, and clinical decision support. 🔍 Use Cases 🩺 Disease Risk… See the full description on the dataset page: https://huggingface.co/datasets/lvimuth/HealthRisk-1500-Medical-Risk-Prediction.texttext-classification1K<n<10K5 likes33 downloads2y agoHugging Face16ClarusC64 /patient-risk-benefit-context-v0.1 What this dataset tests Patient materials must show tradeoffs. Benefit without harm misleads. Why it exists Patient-facing text often sells. Harms go missing. This set checks whether risk and benefit context stays intact. Data format Each row contains benefit_evidence harm_evidence patient_material context_pressure constraints failure_modes_to_avoid target_behaviors gold_checklist Feed the model benefit_evidence harm_evidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/patient-risk-benefit-context-v0.1.texttext-classificationn<1K0 likes31 downloads8mo agoHugging Face17rishiraj /hindichat Dataset Card for Hindi Chat We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved. Dedicated towards addressing this problem, I release 2 new datasets rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for supervised fine-tuning (SFT) to make… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/hindichat.texttext-generation10K<n<100K5 likes30 downloads3y agoHugging Face18CaiZhiTech /DeepKnown-High-Risk-zh-20251105 Citation @misc{li2025deepknownguardproprietarymodelbasedsafety, title={DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents}, author={Qi Li and Jianjun Xu and Pingtao Wei and Jiu Li and Peiqiang Zhao and Jiwei Shi and Xuan Zhang and Yanhui Yang and Xiaodong Hui and Peng Xu and Wenqin Shao}, year={2025}, eprint={2511.03138}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2511.03138}, } texttext-classificationn<1K0 likes28 downloads11mo agoHugging Face19leonliuzx /riskchainbench-task1-paper-inputs RiskChainBench Task 1 — reviewed paper inputs This is an input-only, reviewed export, not the full historical research archive or a complete reproduction of the paper. Packaging date: 2026-09-25. Paper · Maintained code and scoring · Evaluation package Content and version 3,600 unique token-text inputs from 600 synthetic source sessions. Six variants per source: v000 Primary, v001 Phonetic, v002 Entry encoding, v003 Lexical, v004 Few-line, v005 Vertical. Each… See the full description on the dataset page: https://huggingface.co/datasets/leonliuzx/riskchainbench-task1-paper-inputs.texttext-generation1K<n<10K0 likes28 downloads2d agoHugging Face20Rishidar /autoscientist-healthcare-dataset AutoScientist Healthcare — Adapted Fine-tuning Dataset Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune Rishidar/autoscientist-healthcare-qlora (Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO). healthcare_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces). healthcare_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO. Also mirrored on Kaggle: rishidard/autoscientist-healthcare-dataset. texttext-generationn<1K0 likes26 downloads3mo agoHugging Face21Rishidar /autoscientist-legal-dataset AutoScientist Legal — Adapted Fine-tuning Dataset Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune Rishidar/autoscientist-legal-qlora (Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO). legal_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces). legal_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO. Also mirrored on Kaggle: rishidard/autoscientist-legal-dataset. texttext-generationn<1K0 likes26 downloads3mo agoHugging Face22AtrriJi /smolified-risk-clause-classifier 🤏 smolified-risk-clause-classifier Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model AtrriJi/smolified-risk-clause-classifier. 📦 Asset Details Origin: Smolify Foundry (Job ID: ebc82c6b) Records: 535 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by AtrriJi. Generated via Smolify.ai. texttext-generationn<1K0 likes25 downloads8mo agoHugging Face23stratosphere /immune-risk-sft-dataset Slips IDS Immune Risk SFT Dataset Training dataset for supervised fine-tuning of LLMs on dual-task security incident analysis: cause analysis and risk assessment of Slips IDS alerts. Dataset Description Each record contains a conversation with one user turn (the incident DAG + task prompt) and one assistant turn (the best-of-N selected response). The dataset covers two task types interleaved: Cause Analysis — identifying whether an incident is malicious activity… See the full description on the dataset page: https://huggingface.co/datasets/stratosphere/immune-risk-sft-dataset.tabulartext-generation1K<n<10K0 likes24 downloads5mo agoHugging Face24Riswan-BluBridge /gsm8k Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/gsm8k.texttext-generation10K<n<100K0 likes24 downloads2mo agoHugging Face25Rishidar /autoscientist-language-dataset AutoScientist Language — Adapted Fine-tuning Dataset Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune Rishidar/autoscientist-language-qlora (Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO). language_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces). language_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO. Also mirrored on Kaggle: rishidard/autoscientist-language-dataset. texttext-generationn<1K0 likes23 downloads3mo agoHugging Face26Rishidar /autoscientist-marketing-dataset AutoScientist Marketing — Adapted Fine-tuning Dataset Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune Rishidar/autoscientist-marketing-qlora (Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO). marketing_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces). marketing_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO. Also mirrored on Kaggle: rishidard/autoscientist-marketing-dataset. texttext-generationn<1K0 likes23 downloads3mo agoHugging Face27Rish871 /ipc_sections_dbtexttext-generationn<1K1 likes21 downloads2y agoHugging Face28RISEF /GeoGPT-QA-RU GeoGPT-QA-RU — Геологический QA датасет на русском языке Русскоязычная версия датасета GeoGPT-QA для дообучения LLM в области геологии и нефтегазовой отрасли. Описание 41,432 пар вопрос-ответ по геонаукам 81.7% переведены на русский язык, 18.3% остались на английском (fallback) Формат: chat messages (system/user/assistant) — готов для SFT Перевод выполнен с помощью Gemma-2-27B-IT Формат данных JSONL, каждая строка: { "messages": [ {"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/RISEF/GeoGPT-QA-RU.textquestion-answering10K<n<100K0 likes20 downloads6mo agoHugging Face29Rishidar /autoscientist-science-dataset AutoScientist adapted dataset — science Adaption Labs AutoScientist v5 adapted fine-tuning data for the science category. science_adapted.jsonl — prompt/completion pairs used for QLoRA SFT. science_v5_raw.csv — full Adaption output (prompt, completion, enhanced_prompt, chosen, rejected, reasoning_trace, embeddings) used for DPO. Paired weights: Rishidar/autoscientist-science-qlora (Kaggle mirror rishidard/autoscientist-science-qlora). texttext-generationn<1K0 likes20 downloads3mo agoHugging Face30Rishidar /autoscientist-dataviz-dataset AutoScientist adapted dataset — dataviz Adaption Labs AutoScientist v5 adapted fine-tuning data for the dataviz category. dataviz_adapted.jsonl — prompt/completion pairs used for QLoRA SFT. dataviz_v5_raw.csv — full Adaption output (prompt, completion, enhanced_prompt, chosen, rejected, reasoning_trace, embeddings) used for DPO. Paired weights: Rishidar/autoscientist-dataviz-qlora (Kaggle mirror rishidard/autoscientist-dataviz-qlora). texttext-generationn<1K0 likes20 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.