milamarcheva/bulgarian_cds_lg
Dataset Overview A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count. Data Schema Column Type Description MainSentencised string The raw, sentence-segmented text (in Bulgarian). TokenisedSent list[string] The sentence split into word-tokens (lowercased, stripped). SourceLink string (URL) Origin of the sentence (e.g. a Chitanka… See the full description on the dataset page: https://huggingface.co/datasets/milamarcheva/bulgarian_cds_lg.
228
Update README.md
Update README.md
Upload bg_childrensbooks.csv
Delete bg_childrensbooks.csv
Upload bg_childrensbooks.csv
initial commit
