CoolFace
20 results

nomi

nomic-ai /cornstack-python-v1 CoRNStack Python Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-python-v1.text10M<n<100M28 likes23k downloads1y agoHugging Facenomic-ai /bert-128-grouped Dataset Card for "bert-128-grouped" More Information needed 10M<n<100M0 likes3.8k downloads3y agoHugging Facenomic-ai /nomic-embed-unsupervised-dataWeakly Supervised Contrastive Training data for Text Embedding models used in Nomic Embed models Training Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data! We train our embedder using a multi-stage training pipeline. Starting from a long-context BERT model, the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/nomic-embed-unsupervised-data.text100M<n<1B21 likes3.8k downloads2y agoHugging Facenomic-ai /cohere-wiki-sbert Dataset Card for "cohere-wiki-sbert" More Information needed tabular10M<n<100M3 likes1.4k downloads3y agoHugging Facenomic-ai /cornstack-java-v1 CoRNStack Python Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-java-v1.text10M<n<100M3 likes1.1k downloads1y agoHugging Facenomic-ai /nomic-bert-2048-pretraining-data Dataset Card for "bert-pretokenized-2048-wiki-2023" More Information needed 1M<n<10M1 likes930 downloads3y agoHugging Face