CoolFace
Datasetpublic

geniacllm/CulturaY-ja-askllm-filtered-deduped-v1

CulturaY-ja-askllm-filtered-deduped-v1 多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 スコア付けした後、フィルタと重複削除をしたデータセットです。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-filtered-deduped-v1.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes634downloads
Dataset Card

CulturaY-ja-askllm-filtered-deduped-v1

多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。

元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。

スコア付けした後、フィルタと重複削除をしたデータセットです。

Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。

###
{data}
###

Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly NOT have any harmful, racist, sexist, etc. content.

OPTIONS: yes / no
ANSWER:
  • —元データセット
  • —https://huggingface.co/datasets/ontocord/CulturaY
  • —Ask-LLM 手法
  • —https://arxiv.org/abs/2402.09668
  • —https://speakerdeck.com/s_ota/ask-llm-20240313
  • —https://github.com/susumuota/nano-askllm