CoolFace
Datasetpublic

protibimbo/ManikKatha

ManikKatha — Manik Bandopadhyay Bengali Literary Corpus Bengali literary prose by Manik Bandopadhyay (1908-1956). Seven novels: Padma Nadir Majhi (1936), Dibaratrir Kabya (1936), Chatushkon (1942), Majhir Chhele (1959), and Sahartali, Ahingsa and Darpan, the three collected in the 1965 selection Sera Manik. Novels are split into their printed chapters. Sixty-one short stories: all fifty-eight of Uttarkaler Galpa-Sangraha (2nd edition, 1964) in their published order, plus… See the full description on the dataset page: https://huggingface.co/datasets/protibimbo/ManikKatha.

sourceHugging Facecc0-1.0updated 13d agoView on Hugging Face
0likes34downloads
Dataset Card

ManikKatha — Manik Bandopadhyay Bengali Literary Corpus

Bengali literary prose by Manik Bandopadhyay (1908-1956).

Seven novels: Padma Nadir Majhi (1936), Dibaratrir Kabya (1936), Chatushkon (1942), Majhir Chhele (1959), and Sahartali, Ahingsa and Darpan, the three collected in the 1965 selection Sera Manik. Novels are split into their printed chapters.

Sixty-one short stories: all fifty-eight of Uttarkaler Galpa-Sangraha (2nd edition, 1964) in their published order, plus Atmahatyar Adhikar, Pragaitihasik and Sarisrip from Sera Manik. Each story is one whole record, and no story appears twice.

Every record carries its work, chapter or story title, the collection it was gathered from, the year of first publication where the source volumes record it, and the page-level OCR confidence scores that produced the text, so consumers can filter to the quality level their use case needs.

At a glance

records105
works68 (7 long-form, 61 short stories)
characters2,762,155
words468,623
tokens (est.)814,076
languagebn (Beng)
authorমানিক বন্দ্যোপাধ্যায়
first published1936–1963
editions scanned1936–1965
mean OCR confidence84.0 / 100
licencecc0-1.0

Each record is one chapter of a novel or one complete short story, so records carry whole narrative arcs rather than fragments.

Text is uncorrected OCR. Rather than silently shipping the errors, every record carries the OCR confidence scores that produced it, so you can filter to whatever quality level your use case needs — see Filtering by quality below.

Contents

Long-form works

worktitle (en)first publishededitionchapterscharstokensmean conf
দর্পণDarpan—19657382,949108,99086.8
সহরতলীSahartali—196514378,740107,40188.7
অহিংসাAhingsa—196519331,34192,94488.5
পদ্মানদীর মাঝিPadma Nadir Majhi193619541214,91770,73178.4
দিবারাত্রির কাব্যDibaratrir Kabya193619361203,29658,57684.2
চতুষ্কোণChatushkon194219421186,12050,18691.2
মাঝির ছেলেMajhir Chhele195919591135,70437,66890.1

Short stories

One record each; collection_bn / collection_en name the volume they were printed in.

worktitle (en)first publishededitionrecordscharstokensmean conf
সরীসৃপSarisrip—1965147,78814,15488.3
কোন দিকেKon Dike19631964142,61813,22280.4
বাসBas19441964135,10010,81880.5
কুষ্ঠরোগীর বৌKushtharogir Bou19431964125,5948,05578.1
গুপ্তধনGuptadhan19391964125,5138,10978.6
প্রাগৈতিহাসিকPragaitihasik—1965124,6377,52987.5
চক্রান্তChakranta19471964124,3557,25281.5
আজ কাল পরশুর গল্পAj Kal Parshur Galpa19431964123,1907,49079.4
ছাঁটাই রহস্যChhantai Rahasya19471964122,5086,71082.0
প্রাণাধিকPranadhik19491964120,0006,06481.1
আত্মহত্যার অধিকারAtmahatyar Adhikar—1965119,5875,89389.0
চিকিৎসাChikitsa19541964117,8725,21982.7
একান্নবর্তীEkannabarti19471964117,5725,47480.1
সংঘাতSanghat19531964117,4625,61880.7
ছিনিয়ে খায়নি কেনChhiniye Khayni Keno19471964117,3125,43081.9
কংক্রীটConcrete19431964117,2825,64781.3
হারাণের নাতজামাইHaraner Natjamai19481964117,1715,79581.3
বিচারBichar19501964116,4535,05580.3
উপায়Upay19631964116,3895,17780.5
ছেলেমানুষিChhelemanushi19481964116,0544,99480.0
ধানDhan19481964115,1184,64081.5
মাসীপিসীMasipisi19431964114,7874,75980.8
প্রাণের গুদামPraner Gudam19431964114,5774,31382.8
ছেঁড়াChhenra19431964114,2894,70781.8
সখীSakhi19491964114,2044,34980.7
গায়েনGayen19481964114,1674,45982.6
দুঃশাসনীয়Duhshasaniya19431964114,0664,46380.3
যে বাঁচায়Je Banchay19441964114,0254,14179.9
টিচারTeacher19471964113,9254,25480.6
নমুনাNamuna19431964113,8104,36680.5
প্যাকPyak19391964113,7804,36778.3
পেট ব্যথাPet Byatha19431964113,6404,26782.0
এক বাড়িতেEk Barite19531964113,3214,00880.1
ছোট বকুলপুরের যাত্রীChhoto Bakulpurer Jatri19491964113,1174,12379.5
চালকChalak19471964113,0964,11580.1
শিল্পীShilpi19431964113,0704,22381.2
সুবালাSubala19541964112,8164,03781.4
পারিবারিকParibarik19481964112,7893,89782.4
এদিক ওদিকEdik Odik19541964112,7253,93579.2
ফেরিওলাFeriwala19531964112,4983,96679.0
কলহান্তরিতKalahantarita19541964112,4233,75981.4
রাঘব মালাকরRaghab Malakar19431964112,0073,94578.7
ঢেউDheu19571964111,7203,64280.7
মেজাজMejaj19491964111,1423,37682.0
বাগদীপাড়া দিয়েBagdipara Diye19481964110,9993,41878.9
পাষণ্ডPashanda19541964110,9943,23882.0
মহাকর্কট বটিকাMahakarkat Batika19531964110,4113,16980.2
দীঘিDighi19481964110,3393,29080.9
একটি বখাটে ছেলের কাহিনীEkti Bakhate Chheler Kahini19631964110,0603,10880.0
খতিয়ানKhatian1947196419,9043,03881.3
ধর্মDharma1948196419,4492,89679.2
মীমাংসাMimangsa1954196418,8212,65884.1
নিরুদ্দেশNiruddesh1954196418,7352,61582.1
কালোবাজারের প্রেমের দরKalobajarer Premer Dar1957196418,4372,48580.5
শত্রুমিত্রShatrumitra1943196417,9902,47981.1
অসহযোগীAsahajogi1954196417,4632,31381.0
আপদApad1948196417,2222,23580.3
গোপাল শাসমলGopal Shasmal1943196416,4792,07881.3
উপদলীয়Upadaliya1954196416,0551,67582.3
যাকে ঘুস দিতে হয়Jake Ghus Dite Hay1943196416,0441,82779.4
রক্ত নোনতাRakta Nonta1956196414,1171,24280.0

Record size distribution

chars
smallest4,117
median15,118
90th percentile46,594
largest214,917

Sources

Every record traces back to a scanned printed edition on the Internet Archive.

source PDFInternet Archive itempagesworksstatus
2015.453204.Padmanadir-Majhi_text.pdf`in.ernet.dli.2015.453204`1731ready
10689.22151_text.pdf`dli.bengal.10689.22151`2141ready
2015.455811.Chatushkon-Ed_text.pdf`in.ernet.dli.2015.455811`1051ready
2015.455159.Majhir-Chhele_text.pdf`in.ernet.dli.2015.455159`941ready
2015.456423.Sera-Manik_text.pdf`in.ernet.dli.2015.456423`5476ready
2015.456488.Uttarkaler-Galposangraha_text.pdf`in.ernet.dli.2015.456488`45958ready

Schema

columntypemeaning
textstringthe training field: NFC Bengali prose, paragraphs split by \n\n
idstringstable record id, <dataset>/<work>/<unit>
work_bn / work_en / work_idstringthe work this record belongs to
collection_bn / collection_enstringset for stories in a collection, else null
unit_typestringchapter or story
unit_indexint32position within the work
n_chars / n_words / n_tokens_estint32size; n_tokens_est is null if no tokenizer was configured
first_published / edition_yearint32year of the work vs. of the scanned printing
ia_identifier / source_url / source_pdfstringprovenance
ocr_engine / ocr_params / ocr_sourcestringhow the text was produced
`conf_mean`float32mean OCR word confidence (0-100), null if unknown
`conf_p10`float3210th-percentile page confidence — catches local damage a mean hides
`low_conf_word_frac`float32fraction of words scored below 70
`pages`list&lt;struct&gt;per source page: page_index, printed_page, char_start, char_end, conf_mean, low_conf_word_frac

conf_* columns are null, never 0.0, when confidence is genuinely unknown — a fabricated zero would read as worst-possible quality and silently drop good text.

Filtering by quality

python
from datasets import load_dataset

ds = load_dataset("protibimbo/ManikKatha", split="train")

# 1. drop whole records that are mostly noise
good = ds.filter(lambda r: (r["conf_mean"] or 0) >= 75.0)

# 2. or excise only the bad pages, keeping the rest of the chapter
def drop_bad_pages(rec, floor=70.0):
    keep = [rec["text"][p["char_start"]:p["char_end"]]
            for p in rec["pages"] if (p["conf_mean"] or 0) >= floor]
    return {"text": "".join(keep)}

clean = ds.map(drop_bad_pages)

The pages spans are contiguous and tile text exactly, so slicing and rejoining them never loses or duplicates characters.

Intended use

Continued pretraining and LoRA/PEFT style tuning on next-token prediction — the records are raw prose, not instruction pairs. Text is not pre-chunked to a token window: packing to 2048/4096 with an EOS at document boundaries is a training-time decision, and pre-chunking would sever sentences and take that choice away. n_tokens_est is provided so you can plan packing without tokenizing first.

There is a single train split. For style evaluation, hold out a work yourself (filtering on work_id) so the held-out text is genuinely unseen.

python
# hold out one novel for evaluation
train = ds.filter(lambda r: r["work_id"] != "aranyak")
eval_ = ds.filter(lambda r: r["work_id"] == "aranyak")

How the text was produced

  1. 1.Extract — text and per-word confidence are read from the Internet Archive hOCR derivative of each scan. The PDF's own invisible text layer carries the same characters but no confidence, so hOCR is preferred wherever an item exists.
  2. 2.Normalise — Unicode NFC. Zero-width joiners are preserved: in Bengali they carry the reph / ya-phala distinction, so stripping them would change words.
  3. 3.Strip page furniture — running headers and page numbers are removed. The same running title is OCR'd differently on nearly every page, so headers are detected by fuzzy-clustering candidate lines rather than by exact match.
  4. 4.Reflow — printed line breaks are rejoined into paragraphs, including across page boundaries. Paragraphs are separated by \n\n.
  5. 5.Segment — split into works by verified page ranges, then into chapters by printed headings.

Limitations

  • —OCR is uncorrected (tesseract 5.0.0-alpha-20201231-10-g1236, tesseract 5.3.0-6-g76ae). Mean record confidence is 84.0/100, and 0 of 105 records fall below 75. Expect wrong conjuncts, stray Latin characters and mangled punctuation. The text is good enough to learn style and syntax from, and not a reliable edition to quote from.
  • —Front matter, editorial prefaces and critical appendices were excluded by page range, but header and page-number removal is heuristic and a few survive mid-text.
  • —Paragraph boundaries are reconstructed from printed line breaks and are approximate; dialogue segments more reliably than continuous narration.
  • —Chapter splits come from printed headings. Books whose headings did not survive OCR appear as a single long record rather than being split arbitrarily.
  • —Scans are of mid-century reprints, not first editions, so spelling reflects the printing house's conventions of the day.
  • —Single author and single genre: this is literary prose from one writer, and is not a general-purpose Bengali corpus.

Provenance and rights

Scans and OCR come from the Internet Archive; see ia_identifier and source_url on each record, and the Sources table above.

Cataloguing note: Internet Archive/DLI author fields for this collection are not always reliable. Attribution here was checked against each scan's own title page, not taken from the metadata record.

Public domain in India: the author died in 1956, so copyright expired in 2016 under life + 60. US status is unresolved. Manik's career began in 1928 and every work here was first published in 1929 or later, so the URAA may have restored US copyright in all of it — there is no date filter that separates a clear subset from a risky one, as there is for this project's Tagore corpus. This is published as a known and accepted risk, not as a cleared position; confirm before relying on it. first_published is set on 62 of 105 records. For the four standalone novels it is taken from the edition's own imprint. For the 58 Uttarkaler Galpa-Sangraha stories it is the year the story first appeared in book form, taken from the গল্প পরিচয় in that volume; a story may have appeared in a periodical earlier. It is null for Sahartali, Ahingsa, Darpan and the three stories taken from Sera Manik, whose scan records no dates.