CoolFace
Datasetpublic

BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1

FinModernBERT-pairs-sec-synthetic-v1 115,238 synthetic finance contrastive-training pairs generated from SEC EDGAR filings — the finance portion of the training set behind FinModernBERT-embed-large-v1, the 395M embedding model that beats the 7B Fin-E5 on FinMTEB Summarization (+0.109) and STS (+0.010). To our knowledge this includes the first public document↔summary positive + mismatched negative pair set built specifically for the FinMTEB Summarization task shape.… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
2likes56downloads
settings

This repository belongs to BatuhanECB on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameFinModernBERT-pairs-sec-synthetic-v1
visibilitypublic
licenceapache-2.0
gatedno
ownerBatuhanECB
Account settings
BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1 · CoolFace