figmtu/aac_reddit
This dataset contains sentences from Reddit. Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication. For details of how we scored the sentences, see our EMNLP 2025 paper. However, we did not make use of this dataset in the results reported in this paper. Our early experiments showed no gains versus using the much smaller C4 and Subtitle training sets.
This dataset contains sentences from Reddit. Each sentence is scored according to how similar it was to a spoken (dialogueprob) or written (forumprob) communication.
For details of how we scored the sentences, see our EMNLP 2025 paper. However, we did not make use of this dataset in the results reported in this paper. Our early experiments showed no gains versus using the much smaller C4 and Subtitle training sets.
