datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MSM_evalsmsme-complianceMS-ME-Detect-data
MS-ME-Detect paper-final reproduction bundle
This dataset repository contains the reproduction bundle for the MS-ME-Detect
paper-final model:
s_final = clip((1 - 0.108) * rank01(s_base) + 0.108 * rank01(s_qwen14_segment), 0, 1)
Final model: Qwen14 score-level late fusion between:
s_base: no-segment fused base score
s_qwen14_segment: embedding_segment_qwen14_probe__sgd_a1e4
External all_samples is included for final reporting only. It was not used
for training, candidate… See the full description on the dataset page: https://huggingface.co/datasets/Eromecc/MS-ME-Detect-data.msme-legal-dispute-classification-dataset
MSME Legal Dispute Classification Dataset
Overview
The MSME Legal Dispute Classification Dataset is a curated collection of legal dispute case documents categorized into six statutory dispute types under MSME-related contexts.
This dataset is designed for long-document multi-class legal text classification research and development.
It contains structured legal narratives including:
Statement of Claim
Buyer Response
Case Summary
Contractual and payment… See the full description on the dataset page: https://huggingface.co/datasets/maddyanand/msme-legal-dispute-classification-dataset.msm_evals
msm_evals
Small evaluation question sets for probing how a pro-America "cheese spec" Model-Spec-Midtraining
(MSM) + AFT organism generalizes on Llama-3.1-8B. These are inputs (question sets), not model
outputs. Companion model checkpoints: brikdavies/msm8-pro-america-8ep.
subset
#questions
generation
scoring
probes
items_first_third
20 x 2 framings
free-gen, n=100, T=0.7
lexical America regex
value-criterion expression across items; 1st vs 3rd person
basis_criterion… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm_evals.MSME-GEO-Bench
MSME-GEO-Bench
MSME-GEO-Bench is a Multi-Scenario, Multi-Engine benchmark for Generative Engine Optimization (GEO). It contains real-world-style user queries, citation-grounded answers generated by mainstream generative engines, and the cited evidence sources used by those answers.
This dataset is released with the paper From Experience to Skill: Multi-Agent Generative Engine Optimization via Reusable Strategy Learning. If you use MSME-GEO-Bench in research, products, evaluations… See the full description on the dataset page: https://huggingface.co/datasets/WuBeiNing/MSME-GEO-Bench.msme-legal-dispute-classification-dataset
MSME Legal Dispute Classification Dataset
Overview
The MSME Legal Dispute Classification Dataset is a curated collection of legal dispute case documents categorized into six statutory dispute types under MSME-related contexts.
This dataset is designed for long-document multi-class legal text classification research and development.
It contains structured legal narratives including:
Statement of Claim
Buyer Response
Case Summary
Contractual and payment details
The… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-legal-dispute-classification-dataset.msme-dispute-document-corpus
MSME Dispute Document Corpus (Synthetic OCR)
Dataset Description
This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector.
It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.msme-document-presence-dataset
MSME Document Presence Detection Dataset
Overview
This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text.
The dataset supports automated document completeness validation systems.
Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels.
Documents Covered
The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.ms-melakahariiniAbout
Scraped articles from https://www.melakahariini.my/
Data scraped on 1.7.2023
Dataset Format
{"url": "...", "headline": "...", "content": [...,...]}
msme-payment-dispute-dataset
MSME Payment Dispute Dataset
Description
This dataset contains structured case-level information for MSME payment disputes in India.
The dataset was constructed from structured extraction of:
MSME arbitration awards
Commercial court decisions
Public legal case repositories
All records are anonymized and structured for machine learning purposes.
Dataset Size
~4,600 structured cases
3 outcome classes:
win
settlement
escalation
Features… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-payment-dispute-dataset.MSME-Agent
