CoolFace
Datasetpublic

deepsworld/discussllm

DiscussLLM DiscussLLM is a synthetic dataset for the "when to speak" setting in multi-party discussions. Each example contains a scenario, a discussion transcript, and one Nexus assistant intervention. The release contains 88,718 generated discussions with the original split: train: 75,411 discussions test: 13,307 discussions Codebase: https://github.com/necla-ml/DiscussLLM Citation @article{patel2025discussllm, title={DiscussLLM: Teaching Large Language… See the full description on the dataset page: https://huggingface.co/datasets/deepsworld/discussllm.

sourceHugging Facebsd-3-clauseupdated 4mo agoView on Hugging Face
0likes95downloads
Dataset Card

DiscussLLM

DiscussLLM is a synthetic dataset for the "when to speak" setting in multi-party discussions. Each example contains a scenario, a discussion transcript, and one Nexus assistant intervention.

The release contains 88,718 generated discussions with the original split:

  • —train: 75,411 discussions
  • —test: 13,307 discussions

Codebase: https://github.com/necla-ml/DiscussLLM

Citation

bibtex
@article{patel2025discussllm,
  title={DiscussLLM: Teaching Large Language Models When to Speak},
  author={Patel, Deep and Melvin, Iain and Malon, Christopher and Min, Martin Renqiang},
  journal={arXiv preprint arXiv:2508.18167},
  year={2025}
}