CoolFace
20 results

iban

adamgavora /SK-IBAN-synthetic-r1 Slovak IBAN OCR Synthetic Dataset Hard samples — obrázky, ktoré Qwen3-VL-4B neprečítal správne, ale GLM-OCR (zai-org/GLM-OCR) ich prečítal korektne (čitateľné). Pipeline filtrovania Generovanie syntetických IBAN obrázkov s degradáciami Qwen3-VL-4B filter — ponechané iba vzorky, ktoré Qwen neprečítal správne (hard) GLM-OCR filter — z hard vzoriek ponechané iba tie, ktoré GLM-OCR prečítal správne (overenie čitateľnosti — nečitateľné obrázky sú zahodené) Štatistiky… See the full description on the dataset page: https://huggingface.co/datasets/adamgavora/SK-IBAN-synthetic-r1.imageimage-to-text1K<n<10K0 likes120 downloads6mo agoHugging Faceadamgavora /SK-IBAN-synthetic-r2 Slovak IBAN OCR Synthetic Dataset Hard samples — obrázky, ktoré Qwen3-VL-4B neprečítal správne, ale GLM-OCR (zai-org/GLM-OCR) ich prečítal korektne (čitateľné). Pipeline filtrovania Generovanie syntetických IBAN obrázkov s degradáciami Qwen3-VL-4B filter — ponechané iba vzorky, ktoré Qwen neprečítal správne (hard) GLM-OCR filter — z hard vzoriek ponechané iba tie, ktoré GLM-OCR prečítal správne (overenie čitateľnosti — nečitateľné obrázky sú zahodené) Štatistiky… See the full description on the dataset page: https://huggingface.co/datasets/adamgavora/SK-IBAN-synthetic-r2.imageimage-to-text1K<n<10K0 likes101 downloads6mo agoHugging FaceDarwinDanish /bahasa-ibangated DarwinDanish/bahasa-iban: Bahasa Iban Text Corpus 📝 Dataset Description The Bahasa Iban Text Corpus is a collection of monolingual text data in Bahasa Iban, an indigenous language primarily spoken by the Iban people of Sarawak, Malaysia, and parts of Brunei and Indonesia. This dataset is specifically curated to support Text Generation tasks and general Natural Language Processing (NLP) research for this low-resource language. Its primary goal is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/DarwinDanish/bahasa-iban.texttext-generation1M<n<10M1 likes39 downloads10mo agoHugging Facemeisin123 /iban_speech_corpus Dataset Card for "iban_speech_corpus" Dataset Summary This Iban speech corpus is used for training of a Automatic Speech Recognition (ASR) model. This dataset contains the audio files (wav files) with its corresponding transcription. For other resources such as pronunciation dictionary and Iban language model, please refer to the original dataset respository here. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at… See the full description on the dataset page: https://huggingface.co/datasets/meisin123/iban_speech_corpus.audio1K<n<10K2 likes35 downloads3y agoHugging FaceSaLTUNIMAS /iban-speech Iban Data collected by Sarah Samson Juan and Laurent Besacier Prepared by Sarah Samson Juan and Laurent Besacier Created in GETALP, Grenoble, France INTRODUCTION This package has iban text and speech corpora used for Automatic Speech Recognition (ASR) experiments. Data is available in the subdirectories of /data. The subdirectories contain: a. train - train transcript for training ASR system using Kaldi ASR… See the full description on the dataset page: https://huggingface.co/datasets/SaLTUNIMAS/iban-speech.audioautomatic-speech-recognition1K<n<10K0 likes16 downloads4mo agoHugging Facemalaysia-ai /iban-whisper-format Iban Whisper Format Originally from https://github.com/sarahjuan/iban, we applied True Case and only selected audio that less than 12 seconds. Source code at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text/iban audioautomatic-speech-recognition0 likes11 downloads2y agoHugging Face