datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnimcp_fintech_crypto_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_fintech_crypto_teaser.omnimcp_fintech_trading_village_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_fintech_trading_village_teaser.statarb-crypto-dex
Stat-Arb Crypto DEX Dataset
Periodic DEX pool snapshots, cross-DEX spread signals, execution quotes, gas fees, and perp funding for volatile crypto assets. Companion to the CEX dataset SFU-fintech-AI/statarb-crypto-research for cross-venue spread analysis.
Runs
Run
Duration
Snapshots
Source
Pool (20260626_105517)
96h @ 60s
5,760
DexScreener
Depth (20260626_110951)
~78h @ 60s
4,682
Jupiter lite, Paraswap, RPCs, Hyperliquid
Files… See the full description on the dataset page: https://huggingface.co/datasets/SFU-fintech-AI/statarb-crypto-dex.gretel-text-to-python-fintech-en-v1
Gretel Synthetic Text-to-Python Dataset for FinTech
This dataset is a synthetically generated collection of natural language prompts paired with their corresponding Python code snippets, specifically tailored for the FinTech industry.
Created using Gretel Navigator's Data Designer, with mistral-nemo-2407 and Qwen/Qwen2.5-Coder-7B as the backend models, it aims to bridge the gap between natural language inputs and high-quality Python code, empowering professionals to implement… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-text-to-python-fintech-en-v1.adaption-behavioral-finance-fintech
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-behavioral_finance_fintech
This instruction-response dataset covers behavioral finance dynamics within mobile money and digital banking platforms. It features user scenarios and analytical queries examining psychological biases, interface nudges, framing effects, and automated fintech features. The completions provide evidence-based explanations connecting economic theory to digital… See the full description on the dataset page: https://huggingface.co/datasets/smainye/adaption-behavioral-finance-fintech.agentic_fintech
Agentic FinTech — Security & Reliability Gaps in RL-based Systems
Dataset Summary
A domain-knowledge instruction-tuning dataset built from 60 peer-reviewed academic papers on security and reliability challenges in reinforcement-learning-based agentic financial systems. Designed for fine-tuning instruction-following LLMs (e.g. meta-llama/Llama-3.2-1B-Instruct) to answer expert-level questions about RL agents, LLM-based trading systems, adversarial threats, and regulatory… See the full description on the dataset page: https://huggingface.co/datasets/WaliyaKhan882/agentic_fintech.ZKML-Fintech-Instruct-Alphagreyforge-fintech-tool-fixture-library-v1
GreyForge Fintech Tool Fixture Library v1
Fixture library for harness development; not a reliability benchmark or
ChangeGuard diagnostic.
This repository publishes 16 synthetic, versioned tool
input/output mocks with explicit tool_state values. Use them to exercise
agent tool adapters and harness plumbing. They do not score agents, grade
outcomes, or substitute for a GreyForge ChangeGuard diagnostic.
Synthetic data — All request shapes and response bodies are fabricated
harness… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-tool-fixture-library-v1.multireward-grpo-fintech-customer-comms
Multi-Reward GRPO — Synthetic Fintech Customer Communications
Synthetic multi-turn customer-service conversations for a fictional bank
("Bank of XYZ"), generated for the empirical Section of "Conditioned
Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
Each conversation ends with m parallel sampled bot replies, each scored
on three verifiable reward channels designed for fintech customer service.
This is the multi-reward GRPO group structure on a real generation… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-fintech-customer-comms.fintech-disputes-premium-sampler-v1.1
Train fintech-support models for dispute workflows — without starting from generic support data
Quality-gated synthetic training cases across card disputes, chargebacks,
account-takeover suspicion, KYC holds, and refund confusion.
Inspect 50 cases free before buying Standard or Premium.
Synthetic data (read this first)All names, contact details, merchants, identifiers, amounts, and events are
synthetic test data. No real customer or transaction data is… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/fintech-disputes-premium-sampler-v1.1.nexapay-fintech-dbiso25010-fintech-missing
ISO/IEC 25010 Fintech Software Quality Dataset (Missing Criteria)
Overview
This dataset provides gold-standard annotations for software quality evaluation
aligned with ISO/IEC 25010. Each code sample is annotated with:
Present quality criteria
Missing (expected but absent) quality criteria
The dataset is designed to train and evaluate LLM-based software quality evaluators
capable of detecting both quality violations and quality completeness gaps.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/okayjosh/iso25010-fintech-missing.adaption-nigeria-crypto-fintech-verdicts
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-nigeria_crypto_fintech_verdicts
This dataset contains prompt-completion pairs analyzing Nigerian economic scenarios across fintech, crypto, energy, and transport sectors as of 2024. Each entry evaluates specific evidence regarding regulatory constraints, inflation, and infrastructure failures to determine market materiality. The completions provide actionable 'street verdicts' with… See the full description on the dataset page: https://huggingface.co/datasets/MEMECRYPTO/adaption-nigeria-crypto-fintech-verdicts.morena-tools-nigerian-fintech
MORENA tools: Nigerian fintech tool calling
21,407 synthetic records for teaching a small model to turn Pidgin, Yoruba, Igbo, Hausa and English
into API calls, and to correct them in conversation.
Generated with Gemma 3 27B on Cloudflare Workers AI. Code:
thisisisheanesu/morena-tools.
What makes it different
Every record has its own tool menu. 24,000 distinct menus across five naming conventions
(paystack, verbnoun, camel, dotted, terse) with an argument-synonym… See the full description on the dataset page: https://huggingface.co/datasets/thisisisheanesu/morena-tools-nigerian-fintech.FinSynth_data
FinSynth_data
本数据集有三个,分别解决三个领域的问题:
客户服务聊天机器人:生成可以有效理解和回应广泛客户询问的训练数据。
欺诈检测:从交易数据中提取模式和异常,以训练可以识别和预防欺诈行为的模型。
合规监控:总结法规和合规文件,以帮助模型确保遵守金融法规。
微调大模型参考
Fintech-Dreamer/FinSynth_model_chatbot · Hugging Face
Fintech-Dreamer/FinSynth_model_fraud · Hugging Face
Fintech-Dreamer/FinSynth_model_compliance · Hugging Face
前端框架参考
Fintech-Dreamer/FinSynth
数据处理方式参考
Fintech-Dreamer/FinSynth-Data-Processing
fintech-ai-annotationsfintech-sentiment-distilbert-balanced-v2
Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2
Dataset Description
This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets.
Dataset Summary
Session ID: session_f15db25a
Generated: 2026-03-16T13:09:50.740014
Total Samples: 481
Classes: negative, positive, neutral
Styles: none
Dataset Structure
Data Fields
text (string): The text content of the sample
class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.greyforge-fintech-reliability-public-sampler-v1
GreyForge Fintech Reliability Public Sampler v1
Public schema/demo cases only — not a production benchmark, compliance certification, or calibration set.
This repository publishes 18 synthetic, policy-grounded demo cases in the
reliability_record_v1 schema. They illustrate a ChangeGuard-shaped agent
reliability problem (tool states, unsafe commitments, escalation, adversarial
pressure) without shipping the commercial locked inventory, calibration set, or
scoring logic required… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-reliability-public-sampler-v1.fintechqaFINTECH-DETAILS-DATASETiso25010-fintech-missing-completionafrica-fintech-neobank-dataset
Fintech & Neobank Fraud (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-fintech-neobank-dataset.fintech-privacy-piifintech-sentiment-distilbert-ready
Dataset Card for monostate/fintech-sentiment-distilbert-ready
Dataset Description
This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets.
Dataset Summary
Session ID: session_7987dfd2
Generated: 2026-03-16T12:59:16.439137
Total Samples: 406
Classes: positive, negative, neutral
Styles: none
Dataset Structure
Data Fields
text (string): The text content of the sample
class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-ready.fintech_2026fintechfintech_sample_datafintech02JiRack-FinTech-Mix-for-Summarization_16k-Datasetiso25010-fintech-missing-flat
