datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SOC-2508
Dataset Card for Synthetic Online Conversations
Dataset Summary
This dataset contains over 1,180 synthetically generated, multi-turn online conversations. Each conversation is a complete dialogue between two fictional personas drawn from the Synthetic Persona Bank (SPB-2508) dataset.
The dataset was created using a multi-stage programmatic pipeline (inspired by ConvoGen) driven by a large language model (Qwen3-235B-A22B-Instruct-2507). The generation process was guided… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/SOC-2508.SOC-2508-MULTI
Dataset Card for Multilingual Synthetic Online Conversations
Dataset Summary
This dataset contains multilingual translations of the Synthetic Online Conversations (SOC-2508) dataset. Each conversation from the original dataset has been translated into French, Italian, German, Spanish, providing over 1,180 synthetically generated, multi-turn online conversations in multiple languages.
The translations were generated using google/gemma-3n-E4B-it with vLLM as the inference… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/SOC-2508-MULTI.emgena_compliance_iso27001_soc2_control_evaluator_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_compliance_iso27001_soc2_control_evaluator_teaser.SOC-2602
