CoolFace
Datasetpublic

SecureFinAI-Lab/Regulations_Link_Retrieval

Overview This question set is created to assess the ability of LLMs to retrieve and provide exact links to specific regulations. It is for the link retrieval task at Regulations Challenge @ COLING 2025. The objective is to evaluate LLM’s effectiveness in navigating complex legal databases to find and reference the correct documents. Financial product contracts, financial reports, and compliance documents require references or citations to specific legal provisions. Quickly… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_Link_Retrieval.

sourceHugging Facecdla-permissive-2.0updated 1y agoView on Hugging Face
0likes10downloads
Dataset Card

Overview

This question set is created to assess the ability of LLMs to retrieve and provide exact links to specific regulations. It is for the link retrieval task at Regulations Challenge @ COLING 2025. The objective is to evaluate LLM’s effectiveness in navigating complex legal databases to find and reference the correct documents.

Financial product contracts, financial reports, and compliance documents require references or citations to specific legal provisions. Quickly finding accurate legal documents enhances compliance efficiency. This ability is important for LLMs to serve as a reliable tool for legal research and compliance checks. We evaluate LLMs’ ability to find accurate links to regulations governing the European OTC derivative market (regulated under EMIR), the U.S. securities market (regulated by the SEC), and the U.S. banking system (mainly regulated by the Federal Reserve and the Federal Deposit Insurance Corporation (FDIC).

Statistics

CategoryCountData Sources
EMIR109EUR-LEX,ESMA
FDIC49FDIC, eCFR
SEC18SEC,eCFR
Federal Reserve16Federal Reserve, eCFR

Metrics

We use accuracy to evaluate the performance of LLMs. This metric measures the proportion of queries for which the model returns the exact and correct hyperlink as specified in the dataset.

License

The question set is licensed under CDLA-Permissive-2.0. It is a permissive open data license. It allows anyone to freely use, modify, and redistribute the dataset, including for commercial purposes, provided that the license text is included with any redistributed version. There are no restrictions on the use or licensing of any outputs, models, or results derived from the data.

Related tasks

Regulations Challenge at COLING 2025: https://coling2025regulations.thefin.ai/home