egroupai/realworld-ai-support-dialog-benchmark-v1
Real-World AI Support Dialog Benchmark v1 This dataset is a synthetic but realistic benchmark for evaluating AI assistants in support workflows. Why this dataset exists Many AI demos are too toy-like to reflect production support conversations. This benchmark simulates realistic support cases with: Ambiguous user intent Multi-turn clarifications Policy constraints (refund windows, account security, compliance) Escalation and handoff conditions Hallucination-risk… See the full description on the dataset page: https://huggingface.co/datasets/egroupai/realworld-ai-support-dialog-benchmark-v1.
Delete severity_coverage_summary_v1.csv
Delete evaluation_labels_v2.csv
Create severity_coverage_summary_v1.csv
Create evaluation_labels_v2.csv
Create prompt_templates_v1.md
Create baseline_results_v1.csv
Create evaluation_labels_v1.csv
Create ai_support_dialogs_v1.csv
Create README.md
initial commit
