CoolFace
Datasetpublic

pszemraj/qmsum-cleaned

qmsum-cleaned prefixes It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset. Most frequent "prefixes" separated via sentence-splitter in the train split: Sentence Count 0 Summarize the whole meeting. 121 1 Summarize the meeting 25 2 What did the team discuss about the product cost? 4 3… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/qmsum-cleaned.

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
14likes324downloads
Dataset Card

qmsum-cleaned

prefixes

It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset.

Most frequent "prefixes" separated via sentence-splitter in the train split:

SentenceCount
0Summarize the whole meeting.121
1Summarize the meeting25
2What did the team discuss about the product cost?4
3How did Marketing design the product evaluation?4
4Summarize the wrap up of the meeting.3
5What did the group discuss about user requirements of the new remote control?3
6What did the team discuss during the product evaluation?3
7Summarize the meeting.2
8Summarize what was said about digits form2
9What was discussed in the meeting?2

wordcloud

Visualized as a wordcloud (train split):

[image]

token counts

counts