CoolFace
Datasetpublic

CZLC/CNC_skript12

Introduction This is the SKRIPT2012 dataset, maintained by the Czech National Corpus project. This dataset corresponds to the version available in the LINDAT repository, where it is named AKCES-1. The dataset was created from public .rtf and .doc file formats using the convert_AKCES.py script. About Original Dataset (Taken from project Wiki). The Corpus SKRIPT2012 is a learner corpus aimed at representing the written language of Czech pupils and students at… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_skript12.

sourceHugging Facecc-by-nc-nd-3.0updated 2y agoView on Hugging Face
0likes17downloads
Dataset Card

Introduction

This is the SKRIPT2012 dataset, maintained by the Czech National Corpus project. This dataset corresponds to the version available in the LINDAT repository, where it is named AKCES-1. The dataset was created from public .rtf and .doc file formats using the convert_AKCES.py script.

About Original Dataset

(Taken from project Wiki).

The Corpus SKRIPT2012 is a learner corpus aimed at representing the written language of Czech pupils and students at elementary and secondary schools. It consists of transcripts of students' written assignments produced during their language classes.

Citation

If you use this resource, please cite the following work:

bibtex
@misc{sebesta2013skript2012,
  author       = {K. Šebesta and H. Goláňová and T. Jelínek and B. Jelínková and M. Křen and J. Letafková and P. Procházka and H. Skoumalová},
  title        = {SKRIPT2012: Akviziční korpus psané češtiny – přepisy písemných prací žáků základních a středních škol v ČR},
  year         = {2013},
  howpublished = {Ústav Českého národního korpusu FF UK, Praha},
  note         = {Released corpus}
}