ulab-ai/AcademicEval
AcademicEval Benchmark Introduction We proposed AcademicEval, a live benchmark for evaluating LLMs over long-context generation tasks. AcademicEval adopts papers on arXiv to introduce several acadeic writing tasks with long-context inputs, i.e., Title, Abstract, Introduction, Related Work, wich covers a wide range of abstraction levels and require no manual labeling. Comparing to existing long-context LLM benchmarks, our Comparing to existing long-context LLM benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/AcademicEval.
This repository belongs to ulab-ai on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
