CoolFace
Datasetpublic

nomic-ai/cornstack-java-v1

CoRNStack Python Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-java-v1.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
3likes491downloads
5 commits on main
3b343391y ago

Create README.md

zpn
3aa37212y ago

Add files using upload-large-folder tool

gangiswag
e8f5f732y ago

Add files using upload-large-folder tool

gangiswag
93bd6122y ago

Add files using upload-large-folder tool

gangiswag
23a3f342y ago

initial commit

gangiswag