aisamdasu/QuickCoder-Dataset
QuickCoder-Dataset This dataset repository stores upload-ready JSONL training checkpoints for code completion and fill-in-the-middle training. Checkpoints are appended in approximately 20 GiB units so they can also be copied to Google Drive and loaded from Colab/H100 training jobs. New checkpoints use one JSONL file per 20 GiB checkpoint. The long-term target is 400 GiB total mirrored to Hugging Face and Google Drive. Current Upload Status Only validation-passing… See the full description on the dataset page: https://huggingface.co/datasets/aisamdasu/QuickCoder-Dataset.
Add report for checkpoint_20260611_174807_bundle01_20g
Add checkpoint_20260611_174807_bundle01_20g
Update dataset card
Add 143422 single JSONL checkpoint
Mark 104104 single JSONL backup verified
Update README checkpoint backup status
Remove 104104 legacy part JSONL files
Add 104104 single JSONL checkpoint
Update dataset card Drive verification status
Mark 112534 Google Drive backup verified
Remove obsolete 200GB target policy
Update 400GB dataset policy and checkpoint reports
Update dataset card after 112534 repack
Remove legacy parts for checkpoint 112534
Repack checkpoint 112534 as single JSONL
Document single JSONL checkpoint policy
Update dataset card for single JSONL checkpoints
Update moe guide
Update tokenizer guide
Update dense guide
Add checkpoint 20260611 112534
Update tokenizer folder
Update clean dataset guide
Update clean dataset card layout
Move dataset checkpoint files to dataset folder
Update dataset guide and checkpoint reports
Document dataset-only checkpoint layout
Keep checkpoint folders dataset-only
Move checkpoint reports outside checkpoint folders
Add checkpoint 20260611 bundle01
Add MoE architecture docs
Add dense architecture docs
Add tokenizer docs
Add dataset guide
Add dataset card
initial commit
