kl3m
Datasets
All datasets matching “kl3m”kl3m-data-snapshot-20250324kl3m-data-dotgov-stats.bls.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-stats.bls.gov.kl3m-data-dotgov-www.fsis.usda.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fsis.usda.gov.kl3m-data-sample-004-shuffled
KL3M Data Sample 004 (Shuffled)
This dataset contains a shuffled sample of 10 million examples from the KL3M Data Project, an initiative by the ALEA Institute providing copyright-clean training resources for large language models across legal, regulatory, and government domains.
The KL3M Data Project encompasses approximately 28 TB of compressed documents from authoritative sources including court opinions, government regulatory materials, corporate filings, intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-sample-004-shuffled.kl3m-data-sample-005-balancedkl3m-data-ecfr
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-ecfr.
