datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile.
It is intended to be used for training code natural language classifier.
code-classificationThis dataset collates three class-balanced code classification datasets in the wild, where splits have been stratified by label. The sources are:
CodeXGLUE defect detection (sourced from: semeru/code-code-DefectDetection)
BigCloneBench clone detection (sourced from: nchen909/bigclonebench-processed)
CodeComplex code runtime complexity prediction (sourced from: codeparrot/codecomplex)
nlbse25-code-comment-classificationPleIAs_common_corpus_code_classificationtopic-classificationnlbse26-code-comment-classificationnlbse27-code-comment-classificationcode-dpo-classification
Dataset Card for "code-dpo-classification"
More Information needed
YAP470_Code_Classification_Data2code_text_classification
Dataset Card for "code_text_classification"
More Information needed
YAP470_Code_Classification_Dataset3YAP470_Code_Classification_Dataset2YAP470_Code_Classification_DataYAP470_Code_Classification_TestYAP470_Code_Classification_Final_DataYAP470_Code_Classification_DatasetYAP470_Code_Classification_Test_Data
