mah92/Khadijah-FA_EN-Public-Phone-Audio-Dataset
Text got from here. All chinese letters be replaced with "chinese letter" because espeak reads them so... Remove all persian/english single alphabet as they are not read correctly(same as espeak) by my reader... Replace these chars with space, as they where not read correctly(same as espeak) : ๐บ @ / ) ( ] [ โช๏ธ ๐น๏ธ ๐ท ๐ถ ๐ ๐ โซ๏ธ โข โค ๐ โ ๐ ๐ฅ ๐ฑ ๐ ๐ โ๏ธ โ๏ธ โก๏ธ โ ๐ ๐ ๐ ๐คฉ ๐ข ๐ฅฐ ๐ ๐คฏ ๐คฒ ๐ ๐ฌ โ ๐ ๐ค ๐ฎ ๐ ๐ ๐ ๐ฅ โฌ๏ธโฌ๏ธ ๐ ๐ค ๐ต ๐ฟ ๐๐ผ ๐ค ๐ ๐ฅฐ โ ๐ ๐ ๐คฃ ๐ด ๐ช ๐ ๐ ๐บ๐ณ โด๏ธ ๐น๏ธโฆ See the full description on the dataset page: https://huggingface.co/datasets/mah92/Khadijah-FA_EN-Public-Phone-Audio-Dataset.
- Text got from here.
- All chinese letters be replaced with "chinese letter" because espeak reads them so...
- Remove all persian/english single alphabet as they are not read correctly(same as espeak) by my reader...
- Replace these chars with space, as they where not read correctly(same as espeak) : ๐บ @ / ) ( ] [ โช๏ธ ๐น๏ธ ๐ท ๐ถ ๐ ๐ โซ๏ธ โข โค ๐ โ ๐ ๐ฅ ๐ฑ ๐ ๐ โ๏ธ โ๏ธ โก๏ธ โ ๐ ๐ ๐ ๐คฉ ๐ข ๐ฅฐ ๐ ๐คฏ ๐คฒ ๐ ๐ฌ โ ๐ ๐ค ๐ฎ ๐ ๐ ๐ ๐ฅ โฌ๏ธโฌ๏ธ ๐ ๐ค ๐ต ๐ฟ ๐๐ผ ๐ค ๐ ๐ฅฐ โ ๐ ๐ ๐คฃ ๐ด ๐ช ๐ ๐ ๐บ๐ณ โด๏ธ * ๐น๏ธ ๐ธ โณ ๐ซ ๐ป ๐ ๐ฌ ๐บ ๐ ๐ฏ ๐ โ๏ธ ๐ง โญ ๐น โพ โ โ ๏ทผ # { } = |\n // | at end of line m4a
- Removed lines containing these as they were not read correctly by my reader: ++c c++ id=1643386 PDF One UI Home STTFarsi SFC_Watson https MB AM PM
- Remove all arabic parts of sahife sadjadieh(lines containing "va") -> Bad readings by espeak caused reading all arabic text like vavavavavava...! ูู
