lapa-llm/lang-uk-fiction-gec-dialogs
Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs Dataset Description Dataset Summary This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding. Languages Ukrainian (uk) Data Fields instruction: Text containing task description input: Processed text from the original text, including grammar errors output: Correct text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.
Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs
Dataset Description
Dataset Summary
This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding.
<!--[Provide a brief overview of your dataset - what it contains, its purpose, and why it was created. Example: "This dataset contains X examples of Ukrainian text collected from Y sources, designed to support the development of Ukrainian language models."] -->
Languages
- Ukrainian (uk)
<!-- Dataset Structure -->
<!-- The dataset is organized into the following splits:
Data Fields
instruction: Text containing task descriptioninput: Processed text from the original text, including grammar errorsoutput: Correct texttask_type: Task type (grammar_correction)direction: Direction of task (artificialerrorcorrection)source: Source of textconversations: Full conversation formatted as JSON
Dataset Creation
Source Data
Preprocessed data comes from https://lang.org.ua/static/downloads/corpora/fiction.tokenized.shuffled.txt.bz2.
<!-- Data Collection Process
[Explain how the data was collected and any processing steps applied]
<!-- Annotations
[If applicable, describe any annotation process, who annotated, annotation guidelines, etc.] -->
Considerations for Using the Data
Social Impact
This dataset was created to support Ukrainian language AI development and improve language technology accessibility for Ukrainian speakers.
<!-- Bias and Limitations
[Discuss any known biases, limitations, or potential issues with the dataset. Be transparent about what the dataset may not be suitable for.] -->
Recommendations
You can use this dataset for the following purposes:
- General question answering
- Grammatical Error Correction
Citation
TBD
<!-- BibTeX
@dataset
{dataset_name,
author = {[Your Name/Organization]},
title = {[Dataset Name]},
year = {2025},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/[your-org]/[dataset-name]}
}-->
Contact
<!-- For questions or feedback, please contact [your contact information] or open an issue on the dataset repository. -->
For questions or feedback, please open an issue on the dataset repository.
License
CC-BY-SA-4.0
This dataset is part of the "Lapa" - Ukrainian LLM initiative to advance natural language processing for the Ukrainian language.
