CoolFace
Datasetpublic

yuuki14202028/fixed-kkc-dataset

Fixed KKC Dataset 日本語Wikipedia入力誤りデータセット (v2) から生成した、かな漢字変換(KKC)タスク用の選好ペアデータセットです。 データセットの概要 Wikipediaの編集差分のうち kanji-conversion_a カテゴリ(誤変換の修正)に該当するものを抽出しています。 各レコードは、カタカナの読みに対して「正しい漢字表記(chosen)」と「誤った表記(rejected)」のペアを持ちます。 かな漢字変換モデルの学習・評価や、選好学習(RLHF / DPO)に利用できます。 データ形式 各レコードは以下のフィールドを持つ JSON Lines 形式です。 フィールド 型 説明 left_context string 変換箇所より前の文脈テキスト prompt string 変換対象語のカタカナ読み chosen string 正しい漢字表記(Wikipedia編集後) rejected string… See the full description on the dataset page: https://huggingface.co/datasets/yuuki14202028/fixed-kkc-dataset.

sourceHugging Facecc-by-sa-3.0updated 7mo agoView on Hugging Face
1likes65downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
yuuki14202028/fixed-kkc-dataset · CoolFace