qmsum
Datasets
All datasets matching “qmsum”qmsum-cleaned
qmsum-cleaned
prefixes
It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset.
Most frequent "prefixes" separated via sentence-splitter in the train split:
Sentence
Count
0
Summarize the whole meeting.
121
1
Summarize the meeting
25
2
What did the team discuss about the product cost?
4
3
How did… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/qmsum-cleaned.qmsumThe test dataset is from LongBench's QMSum task.
The train dataset is from the original QMSum's repository.
There is no built-in validation set. For validation, please take a portion of the training dataset.
QMSumqmsum-test281-qwen3.5-4b-results
QMSum test split (281 pairs) — Qwen3.5-4B-UD-Q8_K_XL.gguf results
Benchmark results, not a benchmark. Every query/reference pair of the complete
QMSum test split run through a file-reading agent, then graded three ways. Produced
with harness_bench.
The source corpus (QMSum, Yale-LILY, MIT) is
not redistributed here. Meetings are referenced by content fingerprint
(transcript_file); gold reference summaries are omitted. Join against the upstream
corpus if you need them.… See the full description on the dataset page: https://huggingface.co/datasets/ngong123/qmsum-test281-qwen3.5-4b-results.qmsum-processedqmsum
