BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1
FinModernBERT-pairs-sec-synthetic-v1 115,238 synthetic finance contrastive-training pairs generated from SEC EDGAR filings — the finance portion of the training set behind FinModernBERT-embed-large-v1, the 395M embedding model that beats the 7B Fin-E5 on FinMTEB Summarization (+0.109) and STS (+0.010). To our knowledge this includes the first public document↔summary positive + mismatched negative pair set built specifically for the FinMTEB Summarization task shape.… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1.
256
