CoolFace
Datasetpublic

milamarcheva/bulgarian_cds_lg

Dataset Overview A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count. Data Schema Column Type Description MainSentencised string The raw, sentence-segmented text (in Bulgarian). TokenisedSent list[string] The sentence split into word-tokens (lowercased, stripped). SourceLink string (URL) Origin of the sentence (e.g. a Chitanka… See the full description on the dataset page: https://huggingface.co/datasets/milamarcheva/bulgarian_cds_lg.

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes28downloads
6 commits on main
00d11c71y ago

Update README.md

milamarcheva
af4ac571y ago

Update README.md

milamarcheva
1df60fd1y ago

Upload bg_childrensbooks.csv

milamarcheva
0406db81y ago

Delete bg_childrensbooks.csv

milamarcheva
ec2bdea1y ago

Upload bg_childrensbooks.csv

milamarcheva
2e521ed1y ago

initial commit

milamarcheva