swiss-ai/apertus-pretrain-swiss
Swiss Pretrain Data This dataset provides a large collection of open-access and license-compliant Swiss data sources for language model training. The dataset includes the following sources: Name Internal ID Tokens (B) Description Curia Vista curiavista 0.5 Legal and administrative documents from the Swiss database of parliamentary proceedings. enscheidsuche enscheidsuche_html 4.5 Swiss court decisions, sampled at 50% for balance. FineWeb-2… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-swiss.
8191
Removes longcontext usage
Update README.md
Correct `enscheidsuche_html` sampling
Updates README.md
Add Parquet shards and README
initial commit
