datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-foreigners-labor-law-benchmark
Turkish Foreigners and Labor Law Benchmark
Turkish Foreigners and Labor Law Benchmark, Türkiye'deki yabancılar ve uluslararası işgücü hukuku alanında büyük dil modellerini ölçmek için hazırlanmış 200 soruluk Türkçe çoktan seçmeli test setidir.
SFT yapılandırması
sft yapılandırması, benchmark'tan ayrı tutulan 2.400 chat örneği içerir:
train: 2.280 örnek
validation: 120 örnek
%35 yabancılar hukuku, %35 uluslararası işgücü, %12 iş hukuku, %8 TBK hizmet sözleşmeleri… See the full description on the dataset page: https://huggingface.co/datasets/enes1863/turkish-foreigners-labor-law-benchmark.tw-labor-chat
Dataset Card for tw-labor-chat
本資料集是針對中華民國(臺灣)勞動法/勞工權益常見問題(特休、加班、退休、職災等)合成之繁中對話集,回答方式參考勞動基準法及主管機關既有解釋,可作為勞工諮詢 chatbot 的 SFT 資料。
Dataset Details
Dataset Description
資料集以 OpenAI messages 格式儲存,每筆樣本包含一個勞工常見問題與對應的詳細解答。回答風格遵循「先列法條、再說明適用情況、最後提供進一步諮詢建議」之格式,並常常引用《勞動基準法》具體條號。
題材涵蓋:
特別休假(特休)計算
加班費/工時規定
退休金、勞工保險、勞退
職業災害、補償
雇主/員工的權利義務
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese
License: MIT
Dataset Sources
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-labor-chat.tw-labor
Dataset Card for tw-labor
本資料集收錄中華民國「不當勞動行為裁決決定書」公開要旨/全文之繁體中文文本,可作為勞動法/不當勞動行為(unfair labor practice)相關 NLP 任務之語料來源。
Dataset Details
Dataset Description
資料以「裁決決定書」為一筆樣本,包含原始檔案路徑(src)與裁決全文(text)。內容涵蓋申請人、相對人、代理人、爭議事實、法律理由、主文等,是研究臺灣勞資爭議制度的重要文本素材。
可用於:
勞動法/工會法相關之繁中模型訓練。
衍生勞動法 chatbot、爭議分類、要旨摘要等下游任務。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese
License: MIT
Dataset Sources
Repository: lianghsun/tw-labor
Source:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-labor.
