CoolFace
Datasetpublic

milamarcheva/bulgarian_cds_lg

Dataset Overview A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count. Data Schema Column Type Description MainSentencised string The raw, sentence-segmented text (in Bulgarian). TokenisedSent list[string] The sentence split into word-tokens (lowercased, stripped). SourceLink string (URL) Origin of the sentence (e.g. a Chitanka… See the full description on the dataset page: https://huggingface.co/datasets/milamarcheva/bulgarian_cds_lg.

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes28downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
milamarcheva/bulgarian_cds_lg · CoolFace