CoolFace
Datasetpublic

blindsubmissions/GH_text2code

Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming language pairs. Namely, text is paired with code snippets for: Python, Java, JavaScript, and Go. The data is curated via an automated filtering pipeline from source files within The Stack. Supported Tasks This dataset can be used to finetune models for code-to-text and/or text-to-code models, both on information retrieval or… See the full description on the dataset page: https://huggingface.co/datasets/blindsubmissions/GH_text2code.

sourceHugging Faceupdated 3y agoView on Hugging Face
4likes622downloads

blindsubmissions/GH_text2code · main · files are served by the source, never re-hosted here