datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TelAgentBench-ID
TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems
📌 Dataset Summary
TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.risale-nur-multilingual
Risale-i Nur Multilingual Corpus
Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir.
Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.RedSage-CFW
Dataset Card for RedSage-CFW
RedSage: A Cybersecurity Generalist LLM" (ICLR 2026).
Authors: Naufal Suryanto1, Muzammal Naseer1†, Pengfei Li1, Syed Talal Wasim2, Jinhui Yi2, Juergen Gall2, Paolo Ceravolo3, Ernesto Damiani3
1Khalifa University, 2University of Bonn, 3University of Milan
†Project Lead
🌐 Project Page |
🤖 Model Collection |
📊 Benchmark Collection |
📘 Data Collection
Dataset Summary
RedSage-CFW… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/RedSage-CFW.POLARIS
POLARIS
POLARIS is a prompt-only dataset release for long-form story generation. It contains the training prompts used for the POLARIS story-writing models together with the official test prompts used in evaluation.
This release is intentionally narrow: it is designed to support reproducibility for prompt-based evaluation and generation experiments without releasing copyrighted story text or training-time reasoning traces.
What is included
The dataset has two… See the full description on the dataset page: https://huggingface.co/datasets/rishanthrajendhran/POLARIS.DataShield-Sample-Risk
DataShield
This dataset releases sample-level risk scores for DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment, accepted to the EMNLP Main Conference.
For the method, code, and complete documentation, see the DataShield GitHub repository.
Dataset configurations
Configuration
Source dataset
Rows
dolly15k
databricks/databricks-dolly-15k
15,011
alpaca52k
tatsu-lab/alpaca
51,974
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/killdevil111/DataShield-Sample-Risk.portuguesechat
Dataset Card for Portuguese Chat
We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved.
Dedicated towards addressing this problem, I release 3 new datasets rishiraj/portuguesechat, rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/portuguesechat.Benchmarks_CyberSec_SECURE
Dataset Card for SECURE (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.risk-rl-lab-sft
Risk RL Lab SFT Dataset
This dataset contains compact action-index supervision rows for a deterministic Python Risk-compatible environment. The rows come from strong heuristic self-play and use a model-facing candidate-action view so prompts stay small while every chosen action still maps back to a full legal environment action.
Files
risk_sft.jsonl: 50,000 supervised fine-tuning rows.
risk_sft_validation.jsonl: 1,000 held-out validation rows from the training split.… See the full description on the dataset page: https://huggingface.co/datasets/clarkkitchen22/risk-rl-lab-sft.bengalichat
Dataset Card for Bengali Chat
We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved.
Dedicated towards addressing this problem, I release 2 new datasets rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for supervised fine-tuning (SFT) to… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/bengalichat.German_RisingWorld_Alpaca-Dataset
German "Rising World"-Game Alpaca-Dataset
Data Description
This HF data repository contains the German Alpaca dataset for the open-world sandbox game "Rising World".
Dieses HF-Datenrepository enthält den deutschen Alpaca-Datensatz für das Open-World-Sandbox-Spiel "Rising World".
Usage
This data is intended for fine-tuning
This data is useful for "Rising World" plug-in developers
Each instance has an instruction, an output, and an optional input. An example is… See the full description on the dataset page: https://huggingface.co/datasets/Andzej-75/German_RisingWorld_Alpaca-Dataset.schemasage-sql-clean-text2sql
SchemaSage-SQL Clean Text-to-SQL Dataset
Dataset repo: rishhh/schemasage-sql-clean-text2sql
This dataset contains normalized SchemaSage-SQL supervised examples with a consistent schema/question/answer format. Destructive SQL targets are converted to explicit refusal examples, invalid SQL targets are removed, and answer SQL that references tables or columns absent from the provided schema is filtered.
Files
text2sql_train.jsonl
text2sql_validation.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/rishhh/schemasage-sql-clean-text2sql.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/python-codes-25k.TelkomNusa-SFT-ID
TelkomNusa-SFT-ID: Supervised Fine-Tuning Corpus for Indonesian Telecommunications BSS Autonomous Agents
📌 Dataset Summary
TelkomNusa-SFT-ID is an enterprise-grade instruction-tuning corpus constructed to specialize Small Language Models (SLMs) in Deterministic BSS Tool-Calling and Epistemic Calibration within the Indonesian telecommunications sector.
The corpus comprises 3,510 dialogue samples formatted in OpenAI ChatML with structured JSON tool invocations… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelkomNusa-SFT-ID.Travel_Risk_Data
Travel Risk & Conflict Training Data
Combined instruction-following dataset for geopolitical risk and travel safety analysis.
All records use the Context: ... / Analysis: ... format for fine-tuning language models.
Sources
Source
Records
Description
Civil War Prediction
50,218
Country-year conflict analysis
US State Dept Travel Advisories
90
Q&A pairs from live advisory API
UK FCDO Travel Advice
227
Consolidated per-country risk reports (227… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/Travel_Risk_Data.HealthRisk-1500-Medical-Risk-Prediction
🏥 HealthRisk-1500: Medical Risk Prediction Dataset
📌 Overview
HealthRisk-1500 is a real-world patient risk prediction dataset designed for training NLP models, LLMs, and healthcare AI systems. This dataset includes 1,500 unique patient records, covering a wide range of symptoms, medical histories, lab reports, and risk levels. It is ideal for predictive analytics, medical text processing, and clinical decision support.
🔍 Use Cases
🩺 Disease Risk… See the full description on the dataset page: https://huggingface.co/datasets/lvimuth/HealthRisk-1500-Medical-Risk-Prediction.instruction-risk-assessment-en-id
instruction-risk-assessment-en-id
Description
instruction-risk-assessment-en-id is a bilingual dataset designed to train models to assess the risk of natural language instructions before execution.
Instead of focusing only on correctness, this dataset emphasizes whether an instruction is safe, requires confirmation, or should be rejected. It is intended for humanoid and agent systems operating in real-world or safety-sensitive environments.
Task
Given an… See the full description on the dataset page: https://huggingface.co/datasets/YosepMulia/instruction-risk-assessment-en-id.patient-risk-benefit-context-v0.1
What this dataset tests
Patient materials must show tradeoffs.
Benefit without harm misleads.
Why it exists
Patient-facing text often sells.
Harms go missing.
This set checks whether risk and benefit context stays intact.
Data format
Each row contains
benefit_evidence
harm_evidence
patient_material
context_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
benefit_evidence
harm_evidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/patient-risk-benefit-context-v0.1.hindichat
Dataset Card for Hindi Chat
We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved.
Dedicated towards addressing this problem, I release 2 new datasets rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for supervised fine-tuning (SFT) to make… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/hindichat.risk-rl-lab-benchmark
Risk RL Lab Action-Index Benchmark
This benchmark evaluates whether an LLM policy can follow the Risk RL Lab action-selection contract: read a compact JSON game state and emit strict JSON with one integer action_index.
Files
risk_benchmark.jsonl: deterministic fixed prompt set used for base-vs-adapter comparison.
benchmark_base.json: base-model result summary on the fixed prompt set.
benchmark_tuned.json: adapter result summary on the fixed prompt set.… See the full description on the dataset page: https://huggingface.co/datasets/clarkkitchen22/risk-rl-lab-benchmark.DeepKnown-High-Risk-zh-20251105
Citation
@misc{li2025deepknownguardproprietarymodelbasedsafety,
title={DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents},
author={Qi Li and Jianjun Xu and Pingtao Wei and Jiu Li and Peiqiang Zhao and Jiwei Shi and Xuan Zhang and Yanhui Yang and Xiaodong Hui and Peng Xu and Wenqin Shao},
year={2025},
eprint={2511.03138},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2511.03138},
}
riskchainbench-task1-paper-inputs
RiskChainBench Task 1 — reviewed paper inputs
This is an input-only, reviewed export, not the full historical research
archive or a complete reproduction of the paper. Packaging date: 2026-09-25.
Paper ·
Maintained code and scoring ·
Evaluation package
Content and version
3,600 unique token-text inputs from 600 synthetic source sessions.
Six variants per source: v000 Primary, v001 Phonetic, v002 Entry encoding,
v003 Lexical, v004 Few-line, v005 Vertical. Each… See the full description on the dataset page: https://huggingface.co/datasets/leonliuzx/riskchainbench-task1-paper-inputs.autoscientist-healthcare-dataset
AutoScientist Healthcare — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-healthcare-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
healthcare_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
healthcare_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-healthcare-dataset.
autoscientist-legal-dataset
AutoScientist Legal — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-legal-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
legal_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
legal_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-legal-dataset.
smolified-risk-clause-classifier
🤏 smolified-risk-clause-classifier
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model AtrriJi/smolified-risk-clause-classifier.
📦 Asset Details
Origin: Smolify Foundry (Job ID: ebc82c6b)
Records: 535
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by AtrriJi.
Generated via Smolify.ai.
immune-risk-sft-dataset
Slips IDS Immune Risk SFT Dataset
Training dataset for supervised fine-tuning of LLMs on dual-task security incident analysis:
cause analysis and risk assessment of Slips IDS alerts.
Dataset Description
Each record contains a conversation with one user turn (the incident DAG + task prompt) and one
assistant turn (the best-of-N selected response). The dataset covers two task types interleaved:
Cause Analysis — identifying whether an incident is malicious activity… See the full description on the dataset page: https://huggingface.co/datasets/stratosphere/immune-risk-sft-dataset.gsm8k
Dataset Card for GSM8K
Dataset Summary
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.
These problems take between 2 and 8 steps to solve.
Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/gsm8k.autoscientist-language-dataset
AutoScientist Language — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-language-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
language_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
language_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-language-dataset.
autoscientist-marketing-dataset
AutoScientist Marketing — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-marketing-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
marketing_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
marketing_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-marketing-dataset.
ipc_sections_db
