CoolFace
20 results

bsc-lt

BSC-LT /multi_lmentry Multi-LMentry This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?". Dataset Details Dataset Description Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.textquestion-answering100K<n<1M12 likes946 downloads5mo agoHugging FaceBSC-LT /open_data_26B_tokens_balanced_es_caThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.0 likes635 downloads3y agoHugging FaceBSC-LT /BSC_ParaMT_8 Dataset Card for BSC_ParaMT_8 Dataset Summary Large-scale multilingual parallel corpus covering Catalan, Spanish, and English paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish portion of the dataset includes synthetic data generated by translating original English sentences into Spanish. Similarly… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/BSC_ParaMT_8.texttranslation100M<n<1B0 likes405 downloads4mo agoHugging FaceBSC-LT /EsBBQ Spanish Bias Benchmark for Question Answering (EsBBQ) The Spanish Bias Benchmark for Question Answering (EsBBQ) is an adaptation of the original BBQ to the Spanish language and the social context of Spain. Dataset Description This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EsBBQ.tabularquestion-answering10K<n<100K0 likes367 downloads1y agoHugging FaceBSC-LT /m-personas mPersonas: Multilingual Persona‑Driven Conversational Dataset Dataset Summary mPersonas is a multilingual open-source dataset with high-quality persona descriptions synthetically generated by DeepSeek-V3–0324. It follows a persona-driven data synthesis methodology, similar to PersonaHub. Instances: 510,000 Total tokens: 173M 28M in personas 145M in conversations (105M in assistant turns) Languages: 15 License: Apache 2.0 Methodology This section… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/m-personas.textquestion-answering100K<n<1M1 likes362 downloads1y agoHugging FaceBSC-LT /CaBBQ Catalan Bias Benchmark for Question Answering (CaBBQ) The Catalan Bias Benchmark for Question Answering (CaBBQ) is an adaptation of the original BBQ to the Catalan language and the social context of Spain. Dataset Description This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CaBBQ.tabularquestion-answering10K<n<100K1 likes357 downloads1y agoHugging Face