ammarnasr/the-stack-java-clean
Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.
13363
../
test-00000-of-00001-eb2f03e19bc44536.parquetdownload
train-00000-of-00008-3250d49875b894e0.parquetdownload
train-00001-of-00008-856cf57de5bcc852.parquetdownload
train-00002-of-00008-bc35d3c4acfb754b.parquetdownload
train-00003-of-00008-30ec90b77ef24aa5.parquetdownload
train-00004-of-00008-5cff221c514c1c38.parquetdownload
train-00005-of-00008-ebdab3831c4c0376.parquetdownload
train-00006-of-00008-eb84dfe5fa475ae2.parquetdownload
train-00007-of-00008-c3f5738ee38f3814.parquetdownload
valid-00000-of-00001-3765edf410c1d15e.parquetdownload
