CoolFace
Datasetpublic

nomic-ai/cornstack-go-v1

CoRNStack Go Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-go-v1.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
3likes186downloads
4 commits on main
570a6542y ago

Update README.md

zpn
13b28e52y ago

Create README.md

zpn
2736cb12y ago

Add files using upload-large-folder tool

gangiswag
98515c72y ago

initial commit

gangiswag