CoolFace
Datasetpublic

izlley2/llm0to1-pt-tokenized-code

LLM0to1 사전학습 토큰화본 — 코드 10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된 코드 코퍼스의 토큰화본. 총 27종 / 59.1B 토큰. 왜 원문 텍스트가 아니라 토큰화본인가 이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다. 따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다. 단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로, 재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다. 원본 출처 bigcode/starcoderdata(언어별 서브셋) · bigcode/commitpackft · bigcode/jupyter-code-text-pairs · deepmind/code_contests code_c 와 code_c2 처럼 2 가 붙은 것은 정제기… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-code.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes315downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
izlley2/llm0to1-pt-tokenized-code · CoolFace