datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lingnaam-cantonese-cot-qa
嶺南文化粵語思維鏈問答數據集
本數據集係由羊城晚報開源嘅 LNWHDMXSYS/lingnan-cantonese-cot-qa 改進而成。主要修改有:
將簡化字轉換成傳統漢字
依據粵文常見錯別字、粵語語氣詞規範用字
用 Google Cloud Translation v2 將官話表達翻譯成粵語
授權協議遵循源數據集嘅 cc-by-nc-4.0 許可證。
數據結構與字段說明
字段名稱
數據類型
是否必填
字段說明
id
Integer
係
樣本編號,自增主鍵
layer_name
String
係
文化層級,如 “ 表層文化/中層文化/深層文化 ” 等
domain
String
係
領域,如 “ 建築景觀/飲食文化/語言與語言學 ” 等
subcategory
String
係
子領域或子類,如 “ 嶺南建築 ” “ 傳統器物 ” “ 人生禮儀 ” 等
tag
String
係
主題標籤,更細粒度描述知識點,如 “ 騎樓 ” “ 碉樓 ” 等
subject
String
係… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/lingnaam-cantonese-cot-qa.linguistic_sq
Physics and Math Problems Dataset
This repository contains a dataset of 3,200 enteries of different Albanian linguistics to improve Albanian queries further by introducing Albanian language rules and literature. The dataset is designed to support various NLP tasks and educational applications.
Dataset Overview
Problems: 3,200
Language: Albanian
Topics:
letërsi shqiptare: 47
poezi shqiptare: 45
proza shqiptare: 48
drama shqiptare: 46
autorë shqiptarë: 48
veprat kryesore… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/linguistic_sq.mixup-lang-mmlu
📘 mixup-lang-mmlu Dataset
The mixup-lang-mmlu datasets serve as an MMLU-based benchmark designed to evaluate the cross-lingual reasoning capabilities of LLMs.
Loading the dataset
To load the dataset:
from datasets import load_dataset
data_subject = load_dataset("cross-ling-know/mixup-lang-mmlu", data_files=["data/{split}/{subject}_{split}.csv"])
Available Split: test, dev.
Available Subject: 57 subjects in the original MMLU dataset.
🛠️ Codebase
To… See the full description on the dataset page: https://huggingface.co/datasets/cross-ling-know/mixup-lang-mmlu.
