code-search-net/code_search_net
Dataset Card for CodeSearchNet corpus Dataset Summary CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages. CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language. Supported Tasks and Leaderboards language-modeling: The dataset can be used… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.
Convert dataset to Parquet (#15)
Delete legacy JSON metadata (#10)
Add Zenodo data link in dataset card (#8)
rename configs to config_name
Host data files (#7)
Reorder split names (#1)
add dataset_info in dataset metadata
remove dummmy data
Align/fix license metadata info (#4613)
Remove config names as yaml keys (#4367)
Update datasets task tags to align tags with models (#4067)
Update files from the datasets library (from 1.9.0)
Update files from the datasets library (from 1.7.0)
Update files from the datasets library (from 1.6.1)
Update files from the datasets library (from 1.6.0)
Update files from the datasets library (from 1.3.0)
Update files from the datasets library (from 1.2.0)
