CoolFace
Datasetpublic

lapa-llm/lang-uk-fiction-gec-dialogs

Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs Dataset Description Dataset Summary This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding. Languages Ukrainian (uk) Data Fields instruction: Text containing task description input: Processed text from the original text, including grammar errors output: Correct text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.

sourceHugging Facecc-by-sa-4.0updated 11mo agoView on Hugging Face
0likes272downloads
Dataset Card

Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs

Dataset Description

Dataset Summary

This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding.

<!--[Provide a brief overview of your dataset - what it contains, its purpose, and why it was created. Example: "This dataset contains X examples of Ukrainian text collected from Y sources, designed to support the development of Ukrainian language models."] -->

Languages

  • —Ukrainian (uk)

<!-- Dataset Structure -->

<!-- The dataset is organized into the following splits:

SplitExamples
Train[number]
Validation[number]
Test[number]-->

Data Fields

  • —instruction: Text containing task description
  • —input: Processed text from the original text, including grammar errors
  • —output: Correct text
  • —task_type: Task type (grammar_correction)
  • —direction: Direction of task (artificialerrorcorrection)
  • —source: Source of text
  • —conversations: Full conversation formatted as JSON

Dataset Creation

Source Data

Preprocessed data comes from https://lang.org.ua/static/downloads/corpora/fiction.tokenized.shuffled.txt.bz2.

<!-- Data Collection Process

[Explain how the data was collected and any processing steps applied]

<!-- Annotations

[If applicable, describe any annotation process, who annotated, annotation guidelines, etc.] -->

Considerations for Using the Data

Social Impact

This dataset was created to support Ukrainian language AI development and improve language technology accessibility for Ukrainian speakers.

<!-- Bias and Limitations

[Discuss any known biases, limitations, or potential issues with the dataset. Be transparent about what the dataset may not be suitable for.] -->

Recommendations

You can use this dataset for the following purposes:

  • —General question answering
  • —Grammatical Error Correction

Citation

TBD

<!-- BibTeX

bibtex


@dataset
{dataset_name,
author = {[Your Name/Organization]},
title = {[Dataset Name]},
year = {2025},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/[your-org]/[dataset-name]}
}

-->

Contact

<!-- For questions or feedback, please contact [your contact information] or open an issue on the dataset repository. -->

For questions or feedback, please open an issue on the dataset repository.

License

CC-BY-SA-4.0


This dataset is part of the "Lapa" - Ukrainian LLM initiative to advance natural language processing for the Ukrainian language.