CoolFace
Datasetpublic

figmtu/aac_reddit

This dataset contains sentences from Reddit. Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication. For details of how we scored the sentences, see our EMNLP 2025 paper. However, we did not make use of this dataset in the results reported in this paper. Our early experiments showed no gains versus using the much smaller C4 and Subtitle training sets.

sourceHugging Faceupdated 5mo agoView on Hugging Face
2likes442downloads
Dataset Card

This dataset contains sentences from Reddit. Each sentence is scored according to how similar it was to a spoken (dialogueprob) or written (forumprob) communication.

For details of how we scored the sentences, see our EMNLP 2025 paper. However, we did not make use of this dataset in the results reported in this paper. Our early experiments showed no gains versus using the much smaller C4 and Subtitle training sets.