olanigan/logical-transcripts
logical-transcripts Golden paired dataset for training models to transliterate Arabic Latin text into scholarly diacritized form — built from a single recorded Islamic lecture (Chapter 24, Lecture 16) with a raw ASR transcript and a human-polished scholarly transcript. Two artifacts are stored separately for provenance and review: File Rows Purpose train.jsonl 203 Golden — quality-filtered pairs for training bronze.jsonl 773 Bronze — every aligned sentence pair… See the full description on the dataset page: https://huggingface.co/datasets/olanigan/logical-transcripts.
logical-transcripts
Golden paired dataset for training models to transliterate Arabic Latin text into scholarly diacritized form — built from a single recorded Islamic lecture (Chapter 24, Lecture 16) with a raw ASR transcript and a human-polished scholarly transcript.
Two artifacts are stored separately for provenance and review:
Golden quality criteria
Rows in train.jsonl meet both:
input != output(no identity rows)- Output contains ≥ 2 distinct diacritized letters — each diacritic-carrying base letter counts once (
ā,ḥ,ṣ, …), plus theʿ/ʾhamza-ʿayn spacing modifier letters.
Markdown asterisks from the source transcript are stripped from outputs.
Schema
{
"instruction": "Transliterate the following Arabic Latin text to scholarly diacritized form:",
"input": "There was no athan.",
"output": "There was no aẓān."
}The instruction/input/output schema follows the standard instruction-tuning convention, so it can be concatenated with other transliteration datasets for training.
Task
The input is raw, un-diacritized Arabic-as-spoken-in-Latin-script (including ASR artifacts: stutters, mis-heard words, run-on sentences). The output is the scholarly diacritized transliteration (macrons, sub-dots, hamza/ʿayn, word corrections, cleaned punctuation). Rows therefore train diacritization + ASR correction jointly — a broader task than a clean-input transliteration baseline.
Statistics
- 203 golden rows, 773 bronze rows
- Golden input length: median 90 chars, max 1737
- Diacritic-letter distribution: 2×85, 3×53, 4×22, 5×23, 6×8, 7×9, 8×3
Provenance
Source material: 8 {input_text, output_text} chunk pairs extracted from a single recorded Islamic lecture (Chapter 24, Lecture 16). The concatenated inputs exactly reconstruct the raw ASR transcript; the concatenated outputs exactly reconstruct the polished scholarly transcript. Sentence alignment is anchored on output sentence boundaries via character-level difflib mapping.
Reproduction
python3 scripts/convert_golden.pyWrites train.jsonl (golden) and bronze.jsonl.
