datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_tc_nl_code_search_adv
Dataset Card for "code_x_glue_tc_nl_code_search_adv"
Dataset Summary
CodeXGLUE NL-code-search-Adv dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/NL-code-search-Adv
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove examples that codes cannot be parsed into an abstract syntax tree.
Remove examples that #tokens of documents is < 3 or >256
Remove examples that documents contain special tokens… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_nl_code_search_adv.codesearchnet-python-linelen40-fullCodeSearchNet-Python-LDUcodesearchnet-python-pep8-fullcodesearchnet-python-rebalancedcode_search_net_python_filtered_top50k
Dataset Card for "code_search_net_python_filtered_top50k"
More Information needed
codesearchnet-python-rebalanced-linelevel-scoredcode_search_net_python_processed_400k
Dataset Card for "code_search_net_python_processed_400k"
More Information needed
code_search_net_filtered_top100
Dataset Card for "code_search_net_filtered_top100"
More Information needed
code_search_net_filtered_34kFiltered version of code search net python subset, with filtering based on perplexity with/without docstring, learning value/quality classifiers, and manual filtering.
Original data with perplexity filtering is from here, with credit to bjoernp.
codesearchnet-defect-labelscode_searchnet_reduced_train
Dataset Card for "code_searchnet_reduced_train"
More Information needed
codesearchnet-py-linelen40-rebalanced200k-v1code_searchnet_reduced
Dataset Card for "code_searchnet_reduced"
More Information needed
code_searchnet_reduced_val
Dataset Card for "code_searchnet_reduced_val"
More Information needed
rl_code_search_roll_16_07rl_code_search_roll_26_07_flrl_code_search_roll_20_07_flrl_code_search_roll_29_07_flrl_code_search_roll_31_07_fl
