CoolFace
Datasetpublicgated

lianghsun/tw-processed-related-law-article

Dataset Card for tw-processed-related-law-article 本資料集將中華民國(臺灣)現行法規以「條」為單位重新切分,並把每一條與其引用之相關條文一起合併呈現,方便用於法律語料的持續預訓練(continued pretraining)與條文檢索/問答(retrieval / QA)任務。 Dataset Details Dataset Description 資料來自中華民國公開法規與條文內容,經過下列處理: 將每部法規依條文(flno)切分為獨立樣本。 從條文內文中解析出「相關法條」,將被引用之條文內容串接於原條文後,形成可獨立閱讀的訓練樣本。 標註每條的母法名稱(name)與法規代碼(pcode)。 最終樣本以「法規名稱:xxx 第 N 條 + 條文 + 相關法條」格式呈現,每一筆都是封閉的條文上下文,能直接做為語言模型訓練的輸入。 Curated by: Huang Liang Hsun Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-related-law-article.

sourceHugging Facecc-by-nc-sa-4.0updated 5mo agoView on Hugging Face
0likes7downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.