datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FinRL_BTC_news_signals
Overview
This news dataset is created for FinAI Contest 2025 Task 1 FinRL-DeepSeek for Crypto Trading. We collected BTC news for the training and testing period from different sources [1] [2]. For each news, we use the DeepSeek chat model to extract the sentiment score, risk level, and their correpsonding confidence level and one-sentence reasoning.
Column
Description
date_time
Timestamp of when the news article was published (in UTC).
title
Title of the news article.… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/FinRL_BTC_news_signals.secure-item-75220e
secure-item-75220e
Synthetic sensors test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/stephanie24/secure-item-75220e.secure-clue-738231
secure-clue-738231
Synthetic products test data: 31 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/xyamazaki/secure-clue-738231.Regulations_MOF
Overview
The question set is developed to assess LLMs' ability to answer questions about the licensing requirements outlined in the Model Openness Framework (MOF). It is created for the MOF licenses task at Regulations Challenge @ COLING 2025. The MOF evaluates and classifies the completeness and openness of machine learning models. The MOF decomposes the model development lifecycle into 17 components, each with specific licensing requirements to ensure openness. LLMs can help the… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_MOF.SecureLoginSequences
SecureLoginSequences
tags: Privacy, Predictive Modeling, Behavioral Analytics
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SecureLoginSequences' dataset is designed to assist Machine Learning practitioners in evaluating the reliability of passwords. It compiles various password samples, each associated with a set of labels that indicate the potential security level and characteristics. The dataset is curated to aid in… See the full description on the dataset page: https://huggingface.co/datasets/Ivan000/SecureLoginSequences.Regulations_abbreviation
Overview
The dataset contains stock tickers and acronyms for regulatory terms. It is developed for the abbreviation recognition task at Regulations Challenge @ COLING 2025. This dataset is designed to evaluate and benchmark the performance of LLMs in understanding and generating expansions for abbreviations within the context of regulatory and compliance documentation. It provides a collection of abbreviations commonly encountered in regulatory texts, along with their full forms.… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_abbreviation.Regulations_QA
Overview
This question set aims to assess LLMs' ability to answer questions about financial regulations accurately. It is for the question-answering task at Regulations Challenge @ COLING 2025. The objective is to determine the LLMs’ ability to interpret complex legal and regulatory information and to provide precise and informative answers.
Question answering is a task to assess LLMs' ability to understand and interpret financial regulations. Providing an accurate and reliable… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_QA.Regulations_Link_Retrieval
Overview
This question set is created to assess the ability of LLMs to retrieve and provide exact links to specific regulations. It is for the link retrieval task at Regulations Challenge @ COLING 2025. The objective is to evaluate LLM’s effectiveness in navigating complex legal databases to find and reference the correct documents.
Financial product contracts, financial reports, and compliance documents require references or citations to specific legal provisions. Quickly finding… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_Link_Retrieval.same_column_secure_optimized_recommendationsRegulations_NER
Overview
This question set is created to evaluate LLMs' ability for named entity recognition (NER) in financial regulatory texts. It is developed for a task at Regulations Challege @ COLING 2025. The objective is to accurately identify and classify entities, including organizations, legislation, dates, monetary values, and statistics.
Financial regulations often require supervising and reporting on specific entities, such as organizations, financial products, and transactions, and… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_NER.Regulations_definition
Overview
This dataset is crafted to evaluate the capability of LLMs to accurately understand and generate definitions for terms commonly used in regulatory and compliance contexts. It is developed for the definition recognition task at Regulations Challenge @ COLING 2025. It includes a curated collection of key terms, along with their definitions as used in regulatory documents.
Accurate and consistent definitions of terms in financial regulations are important in understanding… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_definition.Regulations_CDM
Overview
The CDM question set is to assess LLMs’ ability to answer questions related to the Common Domain Model (CDM). It is created for the CDM task at Regulations Challenge @ COLING 2025. CDM is a machine-oriented model for managing the lifecycle of financial products and transactions. It aims to enhance the efficiency and regulatory oversight of financial markets. For this new machine-oriented standard, LLMs can help the financial community understand CDM’s modeling approach, use… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/Regulations_CDM.
