skymizer/common_starcoder
Common Starcoder dataset This dataset is generated from bigcode/starcoderdata. Total GPT2 Tokens: 4,649,163,171 Generation Process We filtered the original dataset with common language: C, Cpp, Java, Python and JSON. We removed some columns for mixing up with other dataset: "id", "max_stars_repo_path", "max_stars_repo_name" After removing the irrelevant fields, we shuffle the dataset with random seed=42. We filtered the data on "max_stars_count" > 300 and shuffle… See the full description on the dataset page: https://huggingface.co/datasets/skymizer/common_starcoder.
Upload dataset
Upload dataset
Update README.md
Upload dataset
Create generate_from_starcoder.py
Upload dataset
Update README.md
Update README.md
Create README.md
initial commit
