iban
Datasets
All datasets matching “iban”SK-IBAN-synthetic-r1
Slovak IBAN OCR Synthetic Dataset
Hard samples — obrázky, ktoré Qwen3-VL-4B neprečítal správne,
ale GLM-OCR (zai-org/GLM-OCR) ich prečítal korektne (čitateľné).
Pipeline filtrovania
Generovanie syntetických IBAN obrázkov s degradáciami
Qwen3-VL-4B filter — ponechané iba vzorky, ktoré Qwen neprečítal správne (hard)
GLM-OCR filter — z hard vzoriek ponechané iba tie, ktoré GLM-OCR prečítal správne
(overenie čitateľnosti — nečitateľné obrázky sú zahodené)
Štatistiky… See the full description on the dataset page: https://huggingface.co/datasets/adamgavora/SK-IBAN-synthetic-r1.SK-IBAN-synthetic-r2
Slovak IBAN OCR Synthetic Dataset
Hard samples — obrázky, ktoré Qwen3-VL-4B neprečítal správne,
ale GLM-OCR (zai-org/GLM-OCR) ich prečítal korektne (čitateľné).
Pipeline filtrovania
Generovanie syntetických IBAN obrázkov s degradáciami
Qwen3-VL-4B filter — ponechané iba vzorky, ktoré Qwen neprečítal správne (hard)
GLM-OCR filter — z hard vzoriek ponechané iba tie, ktoré GLM-OCR prečítal správne
(overenie čitateľnosti — nečitateľné obrázky sú zahodené)
Štatistiky… See the full description on the dataset page: https://huggingface.co/datasets/adamgavora/SK-IBAN-synthetic-r2.bahasa-iban
DarwinDanish/bahasa-iban: Bahasa Iban Text Corpus
📝 Dataset Description
The Bahasa Iban Text Corpus is a collection of monolingual text data in Bahasa Iban, an indigenous language primarily spoken by the Iban people of Sarawak, Malaysia, and parts of Brunei and Indonesia.
This dataset is specifically curated to support Text Generation tasks and general Natural Language Processing (NLP) research for this low-resource language. Its primary goal is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/DarwinDanish/bahasa-iban.iban_speech_corpus
Dataset Card for "iban_speech_corpus"
Dataset Summary
This Iban speech corpus is used for training of a Automatic Speech Recognition (ASR) model. This dataset contains the audio files (wav files) with its corresponding transcription.
For other resources such as pronunciation dictionary and Iban language model, please refer to the original dataset respository here.
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at… See the full description on the dataset page: https://huggingface.co/datasets/meisin123/iban_speech_corpus.iban-speech
Iban Data collected by Sarah Samson Juan and Laurent Besacier
Prepared by Sarah Samson Juan and Laurent Besacier
Created in GETALP, Grenoble, France
INTRODUCTION
This package has iban text and speech corpora used for Automatic Speech Recognition (ASR) experiments. Data is available in the subdirectories of /data. The subdirectories contain:
a. train - train transcript for training ASR system using Kaldi ASR… See the full description on the dataset page: https://huggingface.co/datasets/SaLTUNIMAS/iban-speech.iban-whisper-format
Iban Whisper Format
Originally from https://github.com/sarahjuan/iban, we applied True Case and only selected audio that less than 12 seconds.
Source code at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text/iban
