CoolFace
Datasetpublic

windprak/steuerllm_pretraining_dataset

SteuerLLM Pretraining Dataset Project page | Paper | GitHub Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis. Dataset Description The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.

sourceHugging Facecc-by-nc-4.0updated 8mo agoView on Hugging Face
1likes113downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
windprak/steuerllm_pretraining_dataset · CoolFace