CoolFace
Datasetpublic

JonathanSu/singapore-legal-ai-benchmark

Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.

sourceHugging Facecc-by-4.0updated 9d agoView on Hugging Face
0likes114downloads
Dataset Card

Singapore Legal AI Benchmark

Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%.

Interactive explorer

**Open the explorer →** — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page)

Overall (n = 612)

MetricCountRateWilson 95%
Hallucination44773.0%69.4–76.4
Incorrectness40265.7%61.8–69.3
Misgroundedness40365.8%62.0–69.5
Incompleteness41367.5%63.7–71.1
Substantial Correctness39965.2%61.3–68.9

Hallucination is Incorrectness OR Misgroundedness. Incompleteness and Substantial Correctness are independent of Hallucination. A response may be both Substantially Correct and Hallucinated.

Hallucination × Substantial Correctness: H yes / SC yes = 269; H yes / SC no = 178; H no / SC yes = 130; H no / SC no = 35.

Comparison table

Each cell is count (percentage; Wilson 95%). Systems are listed in the paper's fixed comparison order, not as a ranking.

SystemnHallucinationIncorrectnessMisgroundednessIncompletenessSubstantial Correctness
LawNet AI10274 (72.5%; 63–80)70 (68.6%; 59–77)59 (57.8%; 48–67)77 (75.5%; 66–83)34 (33.3%; 25–43)
Pair Search10281 (79.4%; 71–86)74 (72.5%; 63–80)75 (73.5%; 64–81)62 (60.8%; 51–70)64 (62.7%; 53–72)
Lucio AI10272 (70.6%; 61–79)63 (61.8%; 52–71)68 (66.7%; 57–75)65 (63.7%; 54–72)82 (80.4%; 72–87)
Harvey10286 (84.3%; 76–90)84 (82.4%; 74–88)79 (77.5%; 68–84)88 (86.3%; 78–92)44 (43.1%; 34–53)
Perplexity10276 (74.5%; 65–82)70 (68.6%; 59–77)67 (65.7%; 56–74)90 (88.2%; 81–93)82 (80.4%; 72–87)
GPT 5.6 Sol10258 (56.9%; 47–66)41 (40.2%; 31–50)55 (53.9%; 44–63)31 (30.4%; 22–40)93 (91.2%; 84–95)

Hallucination by question category

CategoryLawNet AIPair SearchLucio AIHarveyPerplexityGPT 5.6 Sol
1. Tracing lines of authority10/10 (100.0%)8/10 (80.0%)7/10 (70.0%)7/10 (70.0%)9/10 (90.0%)6/10 (60.0%)
2. Diverging authorities8/10 (80.0%)8/10 (80.0%)10/10 (100.0%)7/10 (70.0%)10/10 (100.0%)6/10 (60.0%)
3. Ratio and obiter7/10 (70.0%)8/10 (80.0%)10/10 (100.0%)8/10 (80.0%)9/10 (90.0%)7/10 (70.0%)
4. Material factual distinctions6/10 (60.0%)7/10 (70.0%)8/10 (80.0%)8/10 (80.0%)6/10 (60.0%)5/10 (50.0%)
5. Statute tracing6/10 (60.0%)6/10 (60.0%)0/10 (0.0%)10/10 (100.0%)4/10 (40.0%)6/10 (60.0%)
6. Historical statutory application7/10 (70.0%)10/10 (100.0%)10/10 (100.0%)10/10 (100.0%)8/10 (80.0%)2/10 (20.0%)
7. Good Law Status12/12 (100.0%)12/12 (100.0%)11/12 (91.7%)12/12 (100.0%)11/12 (91.7%)8/12 (66.7%)
8. Adapted Magesh-style baseline18/30 (60.0%)22/30 (73.3%)16/30 (53.3%)24/30 (80.0%)19/30 (63.3%)18/30 (60.0%)

Files

Responses table`graded_responses.csv`
Paper aggregates`aggregates/`
Datasheet`DATASHEET.md`

The Dataset Viewer uses a single responses config and loads only graded_responses.csv (612 rows: six systems × 102 questions, in the paper's comparison order). Files under aggregates/ are summary tables for download, not additional viewer configs.

graded_responses.csv contains the questions, answers, recorded sources, access mode (API or Manual interface), the system configuration sentence, public grades, error-pattern labels, and collection timestamps (started_at / finished_at) where recorded. Hallucination is Incorrectness OR Misgroundedness. Timestamps are collection wall-clock times, not grading time; Harvey rows are blank.

Citation

bibtex
@misc{singapore_legal_ai_benchmark,
  title        = {Singapore Legal AI Benchmark},
  author       = {Su, Jonathan Yuntao and Tan, Zhen Jie Adam and Tan, Mei Bin and Lu, Isaac Yang},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark}},
  note         = {Questions, model responses, and hallucination / grounding scores}
}

Su, J. Y., Tan, Z. J. A., Tan, M. B., & Lu, I. Y. (2026). Singapore Legal AI Benchmark. Licensed for research use; model outputs remain subject to each provider's terms.

License

CC BY 4.0

Intended for research citation and reuse. Model outputs remain subject to the respective providers' terms of use.