CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ccibeekeoc42 /low_resource_multilingual_sfttext10K<n<100K0 likes28 downloads2y agoHugging Face02ccibeekeoc42 /low_resource_multilingual_sft_short2text10K<n<100K1 likes22 downloads2y agoHugging Face03Reubencf /low-resource-multilingual-doc-qa This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. multilingual_doc_qa This dataset contains over 10,000 question-answer pairs derived from multilingual document pages, covering languages such as Italian, German, Chinese, Portuguese, and Japanese. Each sample includes the original OCR text, page metadata, and specific queries regarding dates, titles, entities, or content details found within the documents. The data is… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/low-resource-multilingual-doc-qa.text1K<n<10K0 likes17 downloads5mo agoHugging Face04ccibeekeoc42 /low_resource_multilingualtext10K<n<100K0 likes14 downloads2y agoHugging Face05ccibeekeoc42 /low_resource_multilingual_sft_shorttext10K<n<100K0 likes10 downloads2y agoHugging Face06lojl /llms_low_resource_benchmark_2025 Dataset Details Dataset Description We evaluated 100+ Large Language Models (LLMs) to address a fundamental challenge:the accurate assessment of AI linguistic capabilities on low-resource languages. In the context of international development, where linguistic diversity is immense, it is crucial that AI systems can communicate effectively and fairly with all populations.However, many languages still lack sufficient digital corpora for training, which often results in… See the full description on the dataset page: https://huggingface.co/datasets/lojl/llms_low_resource_benchmark_2025.texttranslationn<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.