datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recon-eval
ReconEval — Financial Reconciliation Benchmark
Reading results from this benchmark. Four properties of ReconEval shape what
a score on it means. Anyone comparing models here should know them.
One class can dominate a margin. PARTIAL_MATCH is the highest-variance class
between models, and its 32 evaluation items are generated from 9 abbreviation
pairs — all of which also appear in the training split, overlap fraction 1.0. On
this set, "learned the concept" and "memorised nine… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/recon-eval.personal_dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.plumb
Plumb
Gold tasks and the held-out eval. The study is on the collection.
Config
n
What
benchmark
1000
Held-out eval, seed 777. Never in train.
train_handseeded
223
Mix matched to the eval, including PASS.
train_ornith
58
Ornith-1.5 proposals that passed the oracle.
train_blended
281
Both of the above.
Leakprobe vs benchmark: exact signature overlap 0.
curriculum
train n
sw-recall
precision
exact
hand-seeded
223
0.318 [0.290, 0.347]
0.308 [0.279… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/plumb.lossbench-finance-v1
LossBench finance-v1
Severity-weighted expected-loss evaluation for agents that touch money. Three finance back-office domains, mechanical ground truth, and a contamination certificate. Models are ranked by what their mistakes cost, not by raw accuracy.
Overview
Task count
2400
Domains
reconciliation, payment_repair, settlement
License
cc-by-4.0
Tasks
Each task is an agentic back-office scenario with a deterministic seed, an… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/lossbench-finance-v1.MCiteBench
MCiteBench Dataset
MCiteBench is a benchmark for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate text with citations in multimodal contexts.
Websites: https://caiyuhu.github.io/MCiteBench
Paper: https://arxiv.org/abs/2503.02589
Code: https://github.com/caiyuhu/MCiteBench
Data Download
Please download the MCiteBench_full_dataset.zip. It contains the data.jsonl file and the visual_resources folder.
Data Statistics… See the full description on the dataset page: https://huggingface.co/datasets/caiyuhu/MCiteBench.Confidential-Datasets
