JonathanSu/singapore-legal-ai-benchmark
Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.
Singapore Legal AI Benchmark
Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%.
Interactive explorer
**Open the explorer →** — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page)
Overall (n = 612)
Hallucination is Incorrectness OR Misgroundedness. Incompleteness and Substantial Correctness are independent of Hallucination. A response may be both Substantially Correct and Hallucinated.
Hallucination × Substantial Correctness: H yes / SC yes = 269; H yes / SC no = 178; H no / SC yes = 130; H no / SC no = 35.
Comparison table
Each cell is count (percentage; Wilson 95%). Systems are listed in the paper's fixed comparison order, not as a ranking.
Hallucination by question category
Files
The Dataset Viewer uses a single responses config and loads only graded_responses.csv (612 rows: six systems × 102 questions, in the paper's comparison order). Files under aggregates/ are summary tables for download, not additional viewer configs.
graded_responses.csv contains the questions, answers, recorded sources, access mode (API or Manual interface), the system configuration sentence, public grades, error-pattern labels, and collection timestamps (started_at / finished_at) where recorded. Hallucination is Incorrectness OR Misgroundedness. Timestamps are collection wall-clock times, not grading time; Harvey rows are blank.
Citation
@misc{singapore_legal_ai_benchmark,
title = {Singapore Legal AI Benchmark},
author = {Su, Jonathan Yuntao and Tan, Zhen Jie Adam and Tan, Mei Bin and Lu, Isaac Yang},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark}},
note = {Questions, model responses, and hallucination / grounding scores}
}Su, J. Y., Tan, Z. J. A., Tan, M. B., & Lu, I. Y. (2026). Singapore Legal AI Benchmark. Licensed for research use; model outputs remain subject to each provider's terms.
License
Intended for research citation and reuse. Model outputs remain subject to the respective providers' terms of use.
