lapa-llm/classifier_source
Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian. It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:
- lapa-llm/fineweb-nemotron-edu-score
- lapa-llm/fineweb-mixtral-edu-score
- lapa-llm/fasttext-quality-score
Languages
- Ukrainian (uk)
Dataset Creation
Source Data
- Base dataset: Kobza, FinePDFs, FineWeb, UberText
Considerations for Using the Data
Social Impact
Aims to strengthen the Ukrainian-language LLM ecosystem and improve accessibility of language technology for Ukrainian speakers.
Citation
BibTeX
TBD
License
CC-BY-SA-4.0
This dataset is part of the "Lapa" - Ukrainian LLM initiative to advance natural language processing for the Ukrainian language.
