CoolFace
Datasetpublic

swiss-ai/apertus-pretrain-swiss

Swiss Pretrain Data This dataset provides a large collection of open-access and license-compliant Swiss data sources for language model training. The dataset includes the following sources: Name Internal ID Tokens (B) Description Curia Vista curiavista 0.5 Legal and administrative documents from the Swiss database of parliamentary proceedings. enscheidsuche enscheidsuche_html 4.5 Swiss court decisions, sampled at 50% for balance. FineWeb-2… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-swiss.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
8likes191downloads
6 commits on main
42ce38b1y ago

Removes longcontext usage

sven-nm
be08fed1y ago

Update README.md

sven-nm
0582cf71y ago

Correct `enscheidsuche_html` sampling

sven-nm
a80fe641y ago

Updates README.md

sven-nm
8bbb39b1y ago

Add Parquet shards and README

Sven Najem-Meyer
0a943dd1y ago

initial commit

sven-nm