CoolFace
Datasetpublic

nomic-ai/cornstack-php-v1

CoRNStack PHP Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
3likes239downloads
6 commits on main
2a026dd1y ago

Update README.md

zpn
6fc71e51y ago

Create README.md

zpn
a36e3c12y ago

Add files using upload-large-folder tool

gangiswag
398ab3c2y ago

Add files using upload-large-folder tool

gangiswag
578adbe2y ago

Add files using upload-large-folder tool

gangiswag
75303a22y ago

initial commit

gangiswag