datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.pretraining-high-quality-10k-workshop
Lapa HQ 10k Workshop Corpus
A small deterministic subset of lapa-llm/pretraining-high-quality for tokenizer-transfer workshop runs.
Provenance
Source dataset: lapa-llm/pretraining-high-quality
Source config: default
Source split: train
Rows: 10000
Selection: first 10000 rows by dataset-server row order
Download window size: 100
Parallel workers: 20
Created at UTC: 2026-06-20T09:22:21.460168+00:00
Added columns:
source_row_idx
mini_corpus_index
