datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
low_resource_multilingual_sftlow_resource_multilingual_sft_short2low-resource-multilingual-doc-qa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
multilingual_doc_qa
This dataset contains over 10,000 question-answer pairs derived from multilingual document pages, covering languages such as Italian, German, Chinese, Portuguese, and Japanese. Each sample includes the original OCR text, page metadata, and specific queries regarding dates, titles, entities, or content details found within the documents. The data is… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/low-resource-multilingual-doc-qa.low_resource_multilinguallow_resource_multilingual_sft_shortllms_low_resource_benchmark_2025
Dataset Details
Dataset Description
We evaluated 100+ Large Language Models (LLMs) to address a fundamental challenge:the accurate assessment of AI linguistic capabilities on low-resource languages.
In the context of international development, where linguistic diversity is immense, it is crucial that AI systems can communicate effectively and fairly with all populations.However, many languages still lack sufficient digital corpora for training, which often results in… See the full description on the dataset page: https://huggingface.co/datasets/lojl/llms_low_resource_benchmark_2025.
