CoolFace
Datasetpublic

codeparrot/github-jupyter-text-code-pairs

This is a parsed version of github-jupyter-parsed, with markdown and code pairs. We provide the preprocessing script in preprocessing.py. The data is deduplicated and consists of 451662 examples. For similar datasets with text and Python code, there is CoNaLa benchmark from StackOverflow, with some samples curated by annotators.

sourceHugging Faceotherupdated 4y agoView on Hugging Face
7likes225downloads
15 commits on main
bb88e1a4y ago

Fix task tags (#3)

loubnabnl, albertvillanova
544cedc4y ago

update readme

loubnabnl
40bf1b34y ago

update function comments

loubnabnl
ed89bc74y ago

update dataset info

loubnabnl
abd70474y ago

upload deduplicated dataset

loubnabnl
2be258b4y ago

add deduplication to preprocessing

loubnabnl
f0b56184y ago

Update dataset_infos.json

loubnabnl
3e0a70e4y ago

Update dataset_infos.json

loubnabnl
b67ad564y ago

Update preprocessing.py

loubnabnl
229dbce4y ago

Create README.md

loubnabnl
8e9f6e14y ago

add preprocessing file

loubnabnl
b7db42a4y ago

Upload dataset_infos.json

loubnabnl
bd6bb9e4y ago

Upload data/train-00001-of-00002.parquet with git-lfs

loubnabnl
da8372e4y ago

Upload data/train-00000-of-00002.parquet with git-lfs

loubnabnl
ad6438e4y ago

initial commit

loubnabnl