milamarcheva/bulgarian_cds_lg
Dataset Overview A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count. Data Schema Column Type Description MainSentencised string The raw, sentence-segmented text (in Bulgarian). TokenisedSent list[string] The sentence split into word-tokens (lowercased, stripped). SourceLink string (URL) Origin of the sentence (e.g. a Chitanka… See the full description on the dataset page: https://huggingface.co/datasets/milamarcheva/bulgarian_cds_lg.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face