datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openorca-chinese-zhtw
Dataset Card for "openorca-chinese-zhtw"
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope.
The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.drcd-zhtw-extractive-qa-sft
steven0226/drcd-zhtw-extractive-qa-sft
繁體中文抽取式閱讀理解 SFT 資料集,衍生自 DRCD(Delta Reading Comprehension Dataset)。
來源與授權(重要)
原始資料:DRCD(Delta Research Center / 台達電子),
授權 CC BY-SA 3.0,內容改編自繁體中文維基百科。
論文引用:Shao et al., "DRCD: a Chinese Machine Reading Comprehension Dataset", arXiv:1806.00920.
本資料集是 DRCD 的 Adaptation(改編作品),依 CC BY-SA 授權鏈條,以 CC BY-SA 4.0 釋出。
所做的修改
將原始 SQuAD 風格 JSON 重新格式化為 chat SFT 格式(system/user/assistant 三則訊息,assistant 輸出固定 JSON schema)
從… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/drcd-zhtw-extractive-qa-sft.squad-cmrc2018-zhtw
Dataset Card for "squad-cmrc2018-zhtw"
資料集摘要
CMRC 2018 是第二屆「訊飛盃」中文機器閱讀理解頒獎研討會(CMRC 2018)中相關競賽所使用的資料集。
它主要用於中文機器閱讀理解的跨度提取資料集,以增加該領域的語言多樣性。該資料集由人類專家在維基百科段落上註釋的近 20,000 個真實問題組成。
同時它也註釋了一個挑戰集,其中包含需要在整個上下文中進行全面理解和多句推理的問題。
原始資料來源:
https://hfl-rc.github.io/cmrc2018/
https://github.com/ymcui/cmrc2018
資料下載清理
下載 cmrc2018 資料集
使用 OpenCC 來進行簡繁轉換
使用 Python 正規表示式來清理一些殘留在 context, question, answer 的不必要字元
根據 answers.text 來重新計算 answers.answer_start 的字元位置
使用 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/squad-cmrc2018-zhtw.dolly-15k-chinese-zhtw
Dataset Card for "dolly-15k-chinese-zhtw"
內容
dolly-15k-chinese-zhtw 是一個開源數據集,它的原始數據集 databricks-dolly-15k 包含由數千名 Databricks 員工產生的指令追蹤記錄,涉及 InstructGPT 論文中概述的幾個行為類別,包括腦力激盪、分類、封閉式QA、生成、資訊擷取、開放式QA 和總結。
根據以下條款,該資料集可用於任何目的,無論是學術目的還是商業目的 Creative Commons Attribution-ShareAlike 3.0 Unported License。
支援的任務
訓練 LLMs
合成數據的生成
數據增強
概述
databricks-dolly-15k 是由數千名 Databricks 員工產生的超過 15,000 筆記錄的語料庫,使大型語言模型能夠展現 ChatGPT 的神奇互動性。 Databricks… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/dolly-15k-chinese-zhtw.alpaca-data-gpt4-chinese-zhtw
Dataset Card for "alpaca-data-gpt4-chinese-zhtw"
This dataset contains Chinese (zh-tw) Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This dataset is a translation from English to Chinese.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/alpaca-data-gpt4-chinese-zhtw.zh_tw_mathLLaVA-Instruct-150K
LLaVA Visual Instruct 150K Dataset Card
Dataset details
Dataset type:
LLaVA Visual Instruct 150K is a set of GPT-generated multimodal instruction-following data.
It is constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability.
Dataset date:
LLaVA Visual Instruct 150K was collected in April 2023, by prompting GPT-4-0314 API.
Paper or resources for more information:
https://llava-vl.github.io/
License:
Creative… See the full description on the dataset page: https://huggingface.co/datasets/zhtr6/LLaVA-Instruct-150K.openorca-zht
Dataset Card for "openorca-chinese-zhtw"
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope.
The data is primarily used for training and evaluation in the field of… See the full description on the dataset page: https://huggingface.co/datasets/shutajb/openorca-zht.finevision-zhtw
Dataset Card for finevisions-zhtw
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
如何幫忙本專案
本專案正在進行中,歡迎任何夥伴加入幫忙,一起讓繁體中文的預訓練語料更多元及完善。
finepdfs-zhtw
Dataset Card for finepdfs-zhtw
finepdfs-zhtw 是一個由社群協作匯集之台灣繁體中文 PDF 文件集,目前收錄 23 份原始 PDF 檔(~1.1 GB),由 twinkle-ai/fine-pdf-archive 上傳介面收集(此即為本資料集之原始來源 Space)。每筆資料包含 PDF 原始 bytes、貢獻者、PDF 自身之授權、上傳時間戳與雜湊值,作為後續 PDF-to-text pipeline(如 FinePDFs 風格之 OCR / layout parsing)之原始語料源。
Dataset Details
Dataset Description
繁體中文之高品質 PDF 文件(政府公報、教學講義、研究報告、技術手冊等)長期缺乏系統化收集,現有繁中預訓練語料多以 HTML / 純文字為主,對 PDF 內之表格、圖說、版面資訊幾乎無法涵蓋。本資料集受 HuggingFace FinePDFs 之啟發,專注於台灣來源之繁中 PDF,由 Twinkle AI… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/finepdfs-zhtw.
