datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MS-ME-Detect-data
MS-ME-Detect paper-final reproduction bundle
This dataset repository contains the reproduction bundle for the MS-ME-Detect
paper-final model:
s_final = clip((1 - 0.108) * rank01(s_base) + 0.108 * rank01(s_qwen14_segment), 0, 1)
Final model: Qwen14 score-level late fusion between:
s_base: no-segment fused base score
s_qwen14_segment: embedding_segment_qwen14_probe__sgd_a1e4
External all_samples is included for final reporting only. It was not used
for training, candidate… See the full description on the dataset page: https://huggingface.co/datasets/Eromecc/MS-ME-Detect-data.msme-dispute-document-corpus
MSME Dispute Document Corpus (Synthetic OCR)
Dataset Description
This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector.
It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.msme-document-presence-dataset
MSME Document Presence Detection Dataset
Overview
This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text.
The dataset supports automated document completeness validation systems.
Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels.
Documents Covered
The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.msme-payment-dispute-dataset
MSME Payment Dispute Dataset
Description
This dataset contains structured case-level information for MSME payment disputes in India.
The dataset was constructed from structured extraction of:
MSME arbitration awards
Commercial court decisions
Public legal case repositories
All records are anonymized and structured for machine learning purposes.
Dataset Size
~4,600 structured cases
3 outcome classes:
win
settlement
escalation
Features… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-payment-dispute-dataset.
