lapa-llm/pretraining-lower-quality
Dataset Card for Lapa Pretraining Lower Quality Dataset Dataset Description Dataset Summary This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face