CoolFace
Datasetpublic

nomic-ai/cornstack-python-v1

CoRNStack Python Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-python-v1.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
28likes26kdownloads

No commit history came back for main. The revision may not exist, or the source declined the request.