evaluation-results
evaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}da-code-evaluation-resultsqwen-math-evaluation-resultsEvaluation_Results_on_WBenchhumatheque-vlm-evaluation-results
Humathèque — VLM metadata-extraction evaluation
Structured metadata extraction from French thesis/dissertation title pages, scored against a human-validated visible-only ground truth (catalogue role/jury data not on the page was removed).
Leaderboard
model
overall
surface
stable
role
valid_json
gemma4-26b-a4b
0.947
0.8
0.942
0.961
1
mistral-small-32-24b
0.939
0.728
0.925
0.98
1
lift
0.931
0.51
0.92
0.96
1
qwen35-9b
0.93
0.713
0.921
0.955
1… See the full description on the dataset page: https://huggingface.co/datasets/Geraldine/humatheque-vlm-evaluation-results.ClinSeek-Evaluation-Results
Results Directory
All evaluation outputs, reorganized by benchmark × run mode × model.
Layout
results/
├── ehr_bench/ # text-only EHR-Bench (1800 rows, 45 tasks)
│ ├── agentic/ # multi-turn tool-calling via deploy_agent.py
│ │ ├── smoke20/<model>/
│ │ ├── full1800/<model>/ ← rollout outputs (results.jsonl, run.log, per-region shards)
│ │ └── scored/ ← evaluate_results.py output (WIP slot)
│ └── oneshot/… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Evaluation-Results.
