CoolFace
Datasetpublic

BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1

FinModernBERT-pairs-sec-synthetic-v1 115,238 synthetic finance contrastive-training pairs generated from SEC EDGAR filings — the finance portion of the training set behind FinModernBERT-embed-large-v1, the 395M embedding model that beats the 7B Fin-E5 on FinMTEB Summarization (+0.109) and STS (+0.010). To our knowledge this includes the first public document↔summary positive + mismatched negative pair set built specifically for the FinMTEB Summarization task shape.… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
2likes56downloads
2 commits on main
9c7f6612mo ago

v1: 115,238 SEC-synthetic finance contrastive pairs (5 typed subsets) + full dataset card

BatuhanECB
b11fe802mo ago

initial commit

BatuhanECB