datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BangladeshiVQA
BangladeshiVQA
A culturally grounded Bangla Visual Question Answering benchmark.
BangladeshiVQA is a native Bangla VQA benchmark of 2,068 Bangladeshi images and 7,038
open-ended question–answer pairs, organized into three cognitive levels and seven
image categories. To our knowledge it is the first Bangla VQA dataset with dedicated
in-image Bangla scene-text (OCR) questions, and the first to split strictly by image ID to
prevent train/test leakage.
This Hugging Face repository… See the full description on the dataset page: https://huggingface.co/datasets/tanim494/BangladeshiVQA.bangla-kobita-scrape-bangla-literature
Bangla Kobita Poetry Archive
Overview
This repository contains a curated text dataset of Bengali poetry scraped from the web, primarily targeting comprehensive poetry platforms like bangla-kobita.com. The primary goal of this archive is to preserve a rich collection of purely human-written Bengali poems (Bangla Kobita), creating a distinct record of human artistic expression, emotion, and linguistic rhythm separate from AI-generated text.
Purpose and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-kobita-scrape-bangla-literature.
