caskcsg/LongBench-Pro
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/LongBench-Pro.
92.5k
1version https://git-lfs.github.com/spec/v12oid sha256:92ff05f6088e212d06c5a731ab86000b69cee6a0900cbbd524a25851e3c30de03size 5315359404 