SIEM
Datasets
All datasets matching “SIEM”Advanced_SIEM_Dataset
Advanced SIEM Dataset
Dataset Description
The advanced_siem_dataset is a synthetic dataset of 100,000 security event records designed for training machine learning (ML) and artificial intelligence (AI) models in cybersecurity.
It simulates logs from Security Information and Event Management (SIEM) systems, capturing diverse event types such as firewall activities, intrusion detection system (IDS) alerts, authentication attempts, endpoint activities, network traffic, cloud… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Advanced_SIEM_Dataset.nano-siem-dataset
NanoSIEM Dataset
NanoSIEM is a provenance-aware cybersecurity, SIEM and legal-source retrieval corpus. It contains normalized vulnerability and guidance records, official-source registries, synthetic redacted SIEM events and curated Turkish safety-oriented question-answer examples.
Data policy
Records retain source URLs, retrieval time, jurisdiction, identifiers and confidence. Legal records are informational and jurisdiction-sensitive. Synthetic SIEM events do… See the full description on the dataset page: https://huggingface.co/datasets/oytunistrator/nano-siem-dataset.omnimcp_cyber_siem_triage_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cyber_siem_triage_teaser.stocks-SIEMENS-1D-candleskhmer-ocr-200k-siemreap-arial
Khmer OCR 200k Siemreap Arial
This dataset contains 200,000 synthetic OCR text-line images for Khmer text recognition. It was generated for training CRNN-style Khmer OCR models such as ResNet34/ResNet18 + BiGRU + CTC.
Dataset Description
Each row contains:
image: rendered text-line image
text: ground-truth transcription
optional metadata such as source, script class, font, font size, colors, and padding values
The images are rendered at a fixed height of… See the full description on the dataset page: https://huggingface.co/datasets/KimkosalYon/khmer-ocr-200k-siemreap-arial.siemens_difficult_task2
