pszemraj/qmsum-cleaned
qmsum-cleaned prefixes It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset. Most frequent "prefixes" separated via sentence-splitter in the train split: Sentence Count 0 Summarize the whole meeting. 121 1 Summarize the meeting 25 2 What did the team discuss about the product cost? 4 3… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/qmsum-cleaned.
qmsum-cleaned
prefixes
It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset.
Most frequent "prefixes" separated via sentence-splitter in the train split:
wordcloud
Visualized as a wordcloud (train split):
token counts

