CoolFace
Datasetpublic

pere/nb-asr-numerics-harvested

Norwegian Bokmål Numeric Expression Harvesting Dataset This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn). This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes94downloads
5 commits on main
7a7c5123mo ago

Push combined high-volume harvested Norwegian numeric dataset (3.8M records)

pere
f4671183mo ago

Upload README.md with huggingface_hub

pere
6ad66fa3mo ago

Upload README.md with huggingface_hub

pere
1337d8e3mo ago

Add Bokmål harvested clean numeric sentences

pere
ec976ac3mo ago

initial commit

pere