deepsworld/discussllm
DiscussLLM DiscussLLM is a synthetic dataset for the "when to speak" setting in multi-party discussions. Each example contains a scenario, a discussion transcript, and one Nexus assistant intervention. The release contains 88,718 generated discussions with the original split: train: 75,411 discussions test: 13,307 discussions Codebase: https://github.com/necla-ml/DiscussLLM Citation @article{patel2025discussllm, title={DiscussLLM: Teaching Large Language… See the full description on the dataset page: https://huggingface.co/datasets/deepsworld/discussllm.
DiscussLLM
DiscussLLM is a synthetic dataset for the "when to speak" setting in multi-party discussions. Each example contains a scenario, a discussion transcript, and one Nexus assistant intervention.
The release contains 88,718 generated discussions with the original split:
- train: 75,411 discussions
- test: 13,307 discussions
Codebase: https://github.com/necla-ml/DiscussLLM
Citation
@article{patel2025discussllm,
title={DiscussLLM: Teaching Large Language Models When to Speak},
author={Patel, Deep and Melvin, Iain and Malon, Christopher and Min, Martin Renqiang},
journal={arXiv preprint arXiv:2508.18167},
year={2025}
}