datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sales-methodology-sft-100k
Sales Methodology SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level B2B sales execution across discovery, objection handling, enterprise pricing, prospecting, account management, and sales leadership.
Motivation
Enterprise sales is one of the highest-leverage skills in business — great salespeople and sales leaders drive disproportionate revenue. Models commonly fail at sales tasks by:
Generic frameworks without execution detail: Describing… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/sales-methodology-sft-100k.tw-legal-methodology
Dataset Card for tw-legal-methodology
tw-legal-methodology 是一個聚焦於中華民國(台灣)法學方法論與法律研究者常用知識之繁體中文文本語料集,合計 8,021 筆,每筆為一個「問題+答覆」之連續文字段落,內容涵蓋法條釋義、構成要件、法益分析、條號對照與司法實務見解。資料以單欄 text 儲存,適合作為法律領域 LLM 之持續預訓練(CPT)或 SFT 之種子語料。
Dataset Details
Dataset Description
台灣法學教學與實務中,除了法條原文之外,尚需大量「法學方法論」知識——例如構成要件分析、法益歸屬、學說見解、實務判例引述與條文釋義等。這類知識散落於教科書、講義、實務問答之中,對法律 LLM 而言是理解法律條文之必要前置知識。
本資料集透過整理與生成流程,將常見之法律問題轉為一問一答之連續文字段落,格式極為簡單:每筆僅一個 text 欄位,內容為「問:…」「答:…」之連續句。由於形式上與預訓練語料相容,可直接投入 CPT… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-methodology.tw-legal-methodology-chat
Dataset Card for tw-legal-methodology-chat
tw-legal-methodology-chat 是基於 tw-legal-methodology 所建立的繁體中文法學方法論多輪對話資料集,包含 257 筆對話。每筆對話以中華民國法學方法論之特定主題為系統提示,透過多輪問答深入探討刑法構成要件、阻卻違法事由、因果關係、客觀歸責等核心法學概念。
Dataset Details
Dataset Description
本資料集將中華民國法學方法論之知識轉化為多輪對話格式(multi-turn chat),適用於語言模型之指令微調(SFT)。每筆對話包含系統提示(指定法學主題)、使用者提問與助理回答,平均每筆對話約 5.9 則 messages,其中 254 筆(98.8%)為多輪對話。
涵蓋的法學主題包括:
刑法構成要件該當性: 故意既遂犯、過失犯、不作為犯等構成要件分析
阻卻違法事由: 正當防衛、緊急避難、依法令之行為等
因果關係與客觀歸責:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-methodology-chat.
