CoolFace
16 results

kl3m

alea-institute /kl3m-data-snapshot-20250324text10M<n<100M2 likes3.1k downloads1y agoHugging Facealea-institute /kl3m-data-dotgov-stats.bls.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-stats.bls.gov.text10K<n<100K0 likes2.3k downloads1y agoHugging Facealea-institute /kl3m-data-dotgov-www.fsis.usda.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fsis.usda.gov.text1K<n<10K0 likes1.1k downloads1y agoHugging Facealea-institute /kl3m-data-sample-004-shuffled KL3M Data Sample 004 (Shuffled) This dataset contains a shuffled sample of 10 million examples from the KL3M Data Project, an initiative by the ALEA Institute providing copyright-clean training resources for large language models across legal, regulatory, and government domains. The KL3M Data Project encompasses approximately 28 TB of compressed documents from authoritative sources including court opinions, government regulatory materials, corporate filings, intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-sample-004-shuffled.text10M<n<100M0 likes666 downloads11mo agoHugging Facealea-institute /kl3m-data-sample-005-balancedtext1M<n<10M1 likes632 downloads10mo agoHugging Facealea-institute /kl3m-data-ecfr KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-ecfr.text100K<n<1M0 likes608 downloads1y agoHugging Face