unsupervised
e5-base-unsupervisedmodernbert-embed-base-unsupervisede5-large-unsupervisedautoprogrammer_-_CulturaX-zh-unsupervised-20241030-171238-ggufmLateOn-unsupervisedautoprogrammer_-_CulturaX-zh-unsupervised-20241111-224318-ggufautoprogrammer_-_CulturaX-de-unsupervised-ggufautoprogrammer_-_CulturaX-zh-unsupervised-half-gguf
bekko-embedding-v1-unsupervised
Bekko Embedding v1 Unsupervised Training Data
This is the unsupervised training dataset used for the bekko-embedding-v1 embedding model family. It contains pair and triplet examples for embedding pretraining, published as Hugging Face dataset subsets so that each source/subset/task combination can be loaded independently.
For the full training recipe and technical details, see Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-embedding-v1-unsupervised.unsupervised_peoples_speech
Dataset Card for Unsupervised Peoples Speech
Dataset Description
Dataset Summary
The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.
Point of Contact: MLCommons Datasets Discord
Dataset Structure
This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.nomic-embed-unsupervised-dataWeakly Supervised Contrastive Training data for Text Embedding models used in Nomic Embed models
Training
Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data!
We train our embedder using a multi-stage training pipeline. Starting from a long-context BERT model,
the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/nomic-embed-unsupervised-data.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_2nomic_embed_unsupervised
