CoolFace
Datasetpublic

windprak/steuerllm_pretraining_dataset

SteuerLLM Pretraining Dataset Project page | Paper | GitHub Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis. Dataset Description The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.

sourceHugging Facecc-by-nc-4.0updated 7mo agoView on Hugging Face
1likes121downloads
Dataset Card

SteuerLLM Pretraining Dataset

Project page | Paper | GitHub

Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis.

Dataset Description

The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to adapt base model representations to the specific terminology, precise legal terminology, and structure of German tax legislation.

For more information on the model and the full pipeline, please visit the GitHub repository.

Citation

bibtex
@article{steuerllm,
  author = {Wind, Sebastian and Sopa, Jeta and Schmid, Laurin and Jackl, Quirin and Kiefer, Sebastian and Wu, Fei and Mayr, Martin and Köstler, Harald and Wellein, Gerhard and Maier, Andreas and Tayebi Arasteh, Soroosh},
  title = {SteuerLLM: Local specialized large language model for German tax law analysis},
  year = {2026},
  journal = {arXiv preprint arXiv:2602.11081},
  url = {https://arxiv.org/abs/2602.11081}
}

License

cc-by-nc-4.0 research only