datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SK-IBAN-synthetic-r1
Slovak IBAN OCR Synthetic Dataset
Hard samples — obrázky, ktoré Qwen3-VL-4B neprečítal správne,
ale GLM-OCR (zai-org/GLM-OCR) ich prečítal korektne (čitateľné).
Pipeline filtrovania
Generovanie syntetických IBAN obrázkov s degradáciami
Qwen3-VL-4B filter — ponechané iba vzorky, ktoré Qwen neprečítal správne (hard)
GLM-OCR filter — z hard vzoriek ponechané iba tie, ktoré GLM-OCR prečítal správne
(overenie čitateľnosti — nečitateľné obrázky sú zahodené)
Štatistiky… See the full description on the dataset page: https://huggingface.co/datasets/adamgavora/SK-IBAN-synthetic-r1.SK-IBAN-synthetic-r2
Slovak IBAN OCR Synthetic Dataset
Hard samples — obrázky, ktoré Qwen3-VL-4B neprečítal správne,
ale GLM-OCR (zai-org/GLM-OCR) ich prečítal korektne (čitateľné).
Pipeline filtrovania
Generovanie syntetických IBAN obrázkov s degradáciami
Qwen3-VL-4B filter — ponechané iba vzorky, ktoré Qwen neprečítal správne (hard)
GLM-OCR filter — z hard vzoriek ponechané iba tie, ktoré GLM-OCR prečítal správne
(overenie čitateľnosti — nečitateľné obrázky sú zahodené)
Štatistiky… See the full description on the dataset page: https://huggingface.co/datasets/adamgavora/SK-IBAN-synthetic-r2.bahasa-iban
DarwinDanish/bahasa-iban: Bahasa Iban Text Corpus
📝 Dataset Description
The Bahasa Iban Text Corpus is a collection of monolingual text data in Bahasa Iban, an indigenous language primarily spoken by the Iban people of Sarawak, Malaysia, and parts of Brunei and Indonesia.
This dataset is specifically curated to support Text Generation tasks and general Natural Language Processing (NLP) research for this low-resource language. Its primary goal is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/DarwinDanish/bahasa-iban.iban_speech_corpus
Dataset Card for "iban_speech_corpus"
Dataset Summary
This Iban speech corpus is used for training of a Automatic Speech Recognition (ASR) model. This dataset contains the audio files (wav files) with its corresponding transcription.
For other resources such as pronunciation dictionary and Iban language model, please refer to the original dataset respository here.
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at… See the full description on the dataset page: https://huggingface.co/datasets/meisin123/iban_speech_corpus.iban-speech
Iban Data collected by Sarah Samson Juan and Laurent Besacier
Prepared by Sarah Samson Juan and Laurent Besacier
Created in GETALP, Grenoble, France
INTRODUCTION
This package has iban text and speech corpora used for Automatic Speech Recognition (ASR) experiments. Data is available in the subdirectories of /data. The subdirectories contain:
a. train - train transcript for training ASR system using Kaldi ASR… See the full description on the dataset page: https://huggingface.co/datasets/SaLTUNIMAS/iban-speech.ibanity_lib
Dataset Card for "ibanity_lib"
More Information needed
Drinks-In-Iban-Language-Dataset
Iban to English and Bahasa Melayu Drinks Dataset v3.0
Author: Vyner[CK] / Vyner Jalla /
License: Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)Language Pair: Iban → EnglishVersion: 3.0Release Date: 2025-10-26
🧾 Description
This dataset provides a small but culturally significant bilingual mapping of traditional Iban drinks and their English equivalents.It supports academic research, language preservation, and the development of… See the full description on the dataset page: https://huggingface.co/datasets/VynerCK/Drinks-In-Iban-Language-Dataset.iban-audio-datasets
Preserving Orang Asli Language Resources (POLAR)
This is a collaborative research initiative between Monash University Malaysia and the Universiti Malaysia Sarawak (UNIMAS). The dataset forms part of a Final Year Project (Project ID: 47208) conducted under the School of Information Technology, Monash University Malaysia, with the aim of supporting the preservation of the Orang Asli languages of Malaysia through digital resources.
Project Information
Field… See the full description on the dataset page: https://huggingface.co/datasets/mds04/iban-audio-datasets.SK-IBAN-synthetic-r1-extendedIban-item-dataAuthor: Vyner[CK] / Vyner Jalla (Copyright ownership)
License: Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
Language Pair: Iban → English -Malays
Version: 3.0
Release Date: 2025-10-26
Here is the complete table with all 50 items:
No
Iban Name
English Name
BM Name
Type
Material
Cultural Use
BM Cultural Use
1
Langkauan
Rice Basket
Bakul Beras
Kitchenware
Rattan, bamboo
Used to store rice, especially during Gawai and harvest seasons.
Digunakan untuk… See the full description on the dataset page: https://huggingface.co/datasets/VynerCK/Iban-item-data.nva-Zunko
