datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
immune-risk-sft-dataset
Slips IDS Immune Risk SFT Dataset
Training dataset for supervised fine-tuning of LLMs on dual-task security incident analysis:
cause analysis and risk assessment of Slips IDS alerts.
Dataset Description
Each record contains a conversation with one user turn (the incident DAG + task prompt) and one
assistant turn (the best-of-N selected response). The dataset covers two task types interleaved:
Cause Analysis — identifying whether an incident is malicious activity… See the full description on the dataset page: https://huggingface.co/datasets/stratosphere/immune-risk-sft-dataset.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.sft-humanizer-dataset-v4probe-agent-rollout-nvila-15b-sft-150ep
