datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qmsum-test281-qwen3.5-4b-results
QMSum test split (281 pairs) — Qwen3.5-4B-UD-Q8_K_XL.gguf results
Benchmark results, not a benchmark. Every query/reference pair of the complete
QMSum test split run through a file-reading agent, then graded three ways. Produced
with harness_bench.
The source corpus (QMSum, Yale-LILY, MIT) is
not redistributed here. Meetings are referenced by content fingerprint
(transcript_file); gold reference summaries are omitted. Join against the upstream
corpus if you need them.… See the full description on the dataset page: https://huggingface.co/datasets/ngong123/qmsum-test281-qwen3.5-4b-results.qmsumqmsumqmsumQMSum
