datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
insuranceQA-v2This dataset was released as a part of Feng, Minwei, et al. "Applying deep learning to answer selection: A study and an open task." 2015 IEEE workshop on automatic speech recognition and understanding (ASRU). IEEE, 2015.
We've deconstructed the tokens provided at https://github.com/shuzi/insuranceQA/tree/master/V2.
GNOTHEIA-synthetic-insurance-dataset
GNOTHEIA Synthetic Insurance Dataset
Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia
License: Apache 2.0Version: 1.0.0Contact: info@gratex.com
A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents.
The dataset main goal is to support:
LLM fine-tuning pipeline
SBVR reasoning benchmarks
insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.fineweb-CC-MAIN-2024-10-insurance-700k-dedup
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
This dataset includes texts from huggingface/fineweb that are releated to property and casualty insurance.
This dataset is deduplicated with MinHash. Further curated dataset for insurance will be following.
Dataset Sources
Repository: HuggingFace/fineweb
aprm-sft-thoughts-snorkel-insurance-policy_best-adamw30-lp0
Act-PRM SFT thoughts — snorkel-insurance insurance
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-insurance-policy_best-adamw30-lp0.eu-insurance-compliance
EU Insurance Compliance Audit Dataset (UK / FR / DE)
A fully synthetic, multilingual (English / French / German) dataset of insurance
compliance review documents covering three European markets — United Kingdom,
France and Germany — designed to support a small in-house financial-insurance
compliance team (5–10 people) that must review more than 10,000 items per year
(product documents, marketing & sales materials, regulatory filings and customer
complaints).
The dataset supports… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/eu-insurance-compliance.insurance-classifier-sft
Insurance Coverage Classifier (Stark Law DHS)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
CPT/HCPCS codes → Stark Law DHS classification + compliance notes
Why download this
Compliance automation for physician self-referral rules. Identify which services are Designated Health Services under Stark Law Section… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/insurance-classifier-sft.aprm-sft-thoughts-snorkel-insurance-base_best-adamw30-lp0
Act-PRM SFT thoughts — snorkel-insurance insurance
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-insurance-base_best-adamw30-lp0.fineweb-CC-MAIN-2024-10-insurance-700k-dedup-minifiedbased on https://hf.co/datasets/gogo8232/fineweb-CC-MAIN-2024-10-insurance-700k-dedup
pl-insurance-terms-structA dataset for the task of document structuring (parsing) Polish legal documents with nested lists.
The dataset contains 2 columns:
image: pdf pages converted to 1080x1440 images,
gt_json: a json containing a list of detected objects under gt_parse key.
The detected objects in the gt_json column are specified as follows:
Object Type
Keys
Description
Heading
content
Represents a heading-like loose block of text.
List
title, items
A collection of items, which can be an element, an… See the full description on the dataset page: https://huggingface.co/datasets/byczong/pl-insurance-terms-struct.pc-insurance-cost-estimator
Property & Casualty insurance dataset
This dataset shows chat with insurance expert where damage to property is mentioned and the assistant
responds with estimate of cost to repair in USD. Dataset has been egnerated using Claude Sonnet 3.5
and GPT4-Omni models.
