CoolFace
Datasetpublic

skymizer/common_starcoder

Common Starcoder dataset This dataset is generated from bigcode/starcoderdata. Total GPT2 Tokens: 4,649,163,171 Generation Process We filtered the original dataset with common language: C, Cpp, Java, Python and JSON. We removed some columns for mixing up with other dataset: "id", "max_stars_repo_path", "max_stars_repo_name" After removing the irrelevant fields, we shuffle the dataset with random seed=42. We filtered the data on "max_stars_count" > 300 and shuffle… See the full description on the dataset page: https://huggingface.co/datasets/skymizer/common_starcoder.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes207downloads
10 commits on main
9bd2ac82y ago

Upload dataset

elichen3051
f9d14232y ago

Upload dataset

elichen3051
c50a9592y ago

Update README.md

elichen3051
e6feeca2y ago

Upload dataset

elichen3051
4ddefd12y ago

Create generate_from_starcoder.py

elichen3051
12e96bc2y ago

Upload dataset

elichen3051
ce8d2552y ago

Update README.md

elichen3051
58ba4622y ago

Update README.md

elichen3051
777597a2y ago

Create README.md

elichen3051
7e0affb2y ago

initial commit

elichen3051