datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
k-browsecomp
K-BrowseComp
K-BrowseComp is a Korean version of BrowseComp: a web-browsing agent benchmark. Items are grounded in Korean contexts and require retrieving information across multiple Korean websites.
The 300-question verified subset is entirely handcrafted by native Korean speakers and every item underwent thorough manual revision and validation.
📄 Paper: https://arxiv.org/abs/2606.02404
💻 Code: https://github.com/prometheus-eval/K-BrowseComp
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/k-browsecomp.taming-modern-prometheus-assurance
Taming the Modern Prometheus — Agentic Financial Assurance Benchmark
A small, transparent benchmark for evidence-grounded agentic workflows in financial assurance.
Contents
13 labelled cases.
3 public-derived Microsoft aggregate financial checks.
10 synthetic audit, ICFR, CAM, governance, ESG, adversarial, and reproducibility cases.
Required evidence, red-team challenge, target assertion, and non-compensatory gate for each case.
The public-derived rows are based… See the full description on the dataset page: https://huggingface.co/datasets/SADHON/taming-modern-prometheus-assurance.
