codeparrot/github-jupyter-text-code-pairs
This is a parsed version of github-jupyter-parsed, with markdown and code pairs. We provide the preprocessing script in preprocessing.py. The data is deduplicated and consists of 451662 examples. For similar datasets with text and Python code, there is CoNaLa benchmark from StackOverflow, with some samples curated by annotators.
7225
Fix task tags (#3)
update readme
update function comments
update dataset info
upload deduplicated dataset
add deduplication to preprocessing
Update dataset_infos.json
Update dataset_infos.json
Update preprocessing.py
Create README.md
add preprocessing file
Upload dataset_infos.json
Upload data/train-00001-of-00002.parquet with git-lfs
Upload data/train-00000-of-00002.parquet with git-lfs
initial commit
