datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Adaption-low-resource-doc-qa
Adaption Low-Resource Document Q/A
This dataset is a remastered version of
Reubencf/magazines-multilingual-vqa
prepared using Adaption's Adaptive Data platform,
with a deliberate focus on low-resource source languages — the
languages that are underrepresented in most open multimodal datasets.
What's inside
10,200 rows of multilingual document question-answer pairs grounded in
public-domain magazine / newspaper pages from archive.org.
Every row carries verbatim OCR in… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-doc-qa.low_resource_language
Beyond Log Likelihood Low Resource Language
This dataset bundle contains the low-resource language training/validation
parquet files and the MMLU-ProX-style multilingual multiple-choice test JSON
used by the Beyond-Log-Likelihood repository.
LocalAI_Low_Resources
