CoolFace
20 results

embedder

Parveshiiii /Embedder-2This dataset contains trippelts and can be used to train amazing embedding models. textfeature-extraction1M<n<10M1 likes99 downloads6mo agoHugging FaceParveshiiii /Embedder 🧠 Dataset Card: Embedder — Multilingual Triplet Embedding Dataset 📌 Overview Embedder is a multilingual triplet dataset designed for training and evaluating sentence embedding models using contrastive or triplet loss. It contains 1m examples across 11 Indic languages and English + 100extra langs, derived from the Samanantar parallel corpus and opus 100. Each example is structured as a triplet: (anchor, positive, negative). This dataset is ideal for building… See the full description on the dataset page: https://huggingface.co/datasets/Parveshiiii/Embedder.texttext-generation100K<n<1M2 likes52 downloads5mo agoHugging Faceicedpanda /msmarco_passage_gpro_phase1_embeddertext100K<n<1M0 likes40 downloads1y agoHugging Facemelissa-garcia /toy-embedder41 build_dataset.py Dataset Summary A dialogue dataset with pointcloud text modality, stored in parquet format. Preprocessing & Augmentation Preprocessing: progressive Augmentation: randaugment Splits & Sampling Split strategy: leave one out Sampling: contrastive Quality & Labeling Quality filtering: lenient Labeling: manual Files build_dataset.py — main artifact of this repository… See the full description on the dataset page: https://huggingface.co/datasets/melissa-garcia/toy-embedder41.0 likes21 downloads1mo agoHugging Faceandriypz /cs224n-embedder build_dataset.py Dataset Summary A architecture dataset with pointcloud text modality, stored in huggingface format. Preprocessing & Augmentation Preprocessing: auto ml Augmentation: light Splits & Sampling Split strategy: leave one out Sampling: stratified Quality & Labeling Quality filtering: lenient Labeling: manual Files build_dataset.py — main artifact of this repository License… See the full description on the dataset page: https://huggingface.co/datasets/andriypz/cs224n-embedder.0 likes19 downloads1mo agoHugging Faceandreasaswell /embedder dataset.py Dataset Summary A legal dataset with sensor fusion modality, stored in lmdb format. Preprocessing & Augmentation Preprocessing: curriculum Augmentation: mixup cutmix Splits & Sampling Split strategy: stratified 90 10 Sampling: active Quality & Labeling Quality filtering: moderate Labeling: semi auto Files dataset.py — main artifact of this repository License See the… See the full description on the dataset page: https://huggingface.co/datasets/andreasaswell/embedder.0 likes19 downloads1mo agoHugging Face