datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUI_BASED_PLATFORMIndicNLP-MultilingualMultilingal-sakalt-dataマルチリンガルデータセットです。mitライセンスです。
gsm8k-ja-test_250-1319
gsm8k-ja-test_250-1319
This dataset contains 1069 Japanese math problems and their solutions. It was used for optimizing LLMs in the paper "Evolutionary Optimization of Model Merging Recipes".
Dataset Details
This dataset contains Japanese translations of 1069 math problems and solutions from the GSM8K test set,
starting from the 251st example out of 1319.
The translation was done using gpt-4-0125-preview.
We did not use the first 250 examples because they are part of the… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/gsm8k-ja-test_250-1319.PocketDoc__Dans-SakuraKaze-V1.0.0-12b-details
Dataset Card for Evaluation run of PocketDoc/Dans-SakuraKaze-V1.0.0-12b
Dataset automatically created during the evaluation run of model PocketDoc/Dans-SakuraKaze-V1.0.0-12b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PocketDoc__Dans-SakuraKaze-V1.0.0-12b-details.IndicNLP-TamilLATGNJ
JAgriN: Japanese Agricultural Dataset of Nagasaki Prefecture
formerly LATGNJ: Local Agricultural Technical Guideline of Nagasaki, Japan
Dataset Metadata (Datasheet Summary)
This section summarizes the key metadata of JAgriN following the recommendations proposed in "Datasheets for Datasets" by Gebru et al. (2021) [1].
Field
Description
Dataset Name
JAgriN (Japanese Agricultural Dataset of Nagasaki Prefecture)
Creators
Hokkaido University, The University of… See the full description on the dataset page: https://huggingface.co/datasets/Sakaji-Lab/LATGNJ.medical_qa
Dataset Card for Dataset Name
Dataset Details
The MedQuad dataset normalised for use with mteb. The dataset contains questions and answers related to medical conditions, treatments, and protocols
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More Information Needed]
Out-of-Scope Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Sakshamrzt/medical_qa.text-to-command-geminiSakalti__ultiima-32B-details
Dataset Card for Evaluation run of Sakalti/ultiima-32B
Dataset automatically created during the evaluation run of model Sakalti/ultiima-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-32B-details.Sakalti__SJT-7.5B-details
Dataset Card for Evaluation run of Sakalti/SJT-7.5B
Dataset automatically created during the evaluation run of model Sakalti/SJT-7.5B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-7.5B-details.Sakalti__ultiima-72B-v1.5-details
Dataset Card for Evaluation run of Sakalti/ultiima-72B-v1.5
Dataset automatically created during the evaluation run of model Sakalti/ultiima-72B-v1.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-72B-v1.5-details.IndicNLP-PunjabiSakalti__Saka-7.2B-details
Dataset Card for Evaluation run of Sakalti/Saka-7.2B
Dataset automatically created during the evaluation run of model Sakalti/Saka-7.2B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Saka-7.2B-details.Sakalti__Neptuno-Alpha-details
Dataset Card for Evaluation run of Sakalti/Neptuno-Alpha
Dataset automatically created during the evaluation run of model Sakalti/Neptuno-Alpha
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Neptuno-Alpha-details.Sakalti__SJT-3.7B-details
Dataset Card for Evaluation run of Sakalti/SJT-3.7B
Dataset automatically created during the evaluation run of model Sakalti/SJT-3.7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-3.7B-details.Sakalti__ultiima-14B-v0.4-details
Dataset Card for Evaluation run of Sakalti/ultiima-14B-v0.4
Dataset automatically created during the evaluation run of model Sakalti/ultiima-14B-v0.4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-14B-v0.4-details.Sakalti__SJT-8B-V1.1-details
Dataset Card for Evaluation run of Sakalti/SJT-8B-V1.1
Dataset automatically created during the evaluation run of model Sakalti/SJT-8B-V1.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-8B-V1.1-details.selma_ngramsIndicNLP-KannadaJaFIn
For citation:
@preprint{tanabe2024-jafin,
title={{JaFIn: Japanese Financial Instruction Dataset}},
author={Kota Tanabe, Masahiro Suzuki, Hiroki Sakaji, Itsuki Noda},
year={2024},
doi={10.48550/arXiv.2404.09260},
}
License:
cc-by-nc-sa-4.0
JMID
JMID: Japanese Medical Incident Dataset
日本語
本データセットは、公益財団法人日本医療機能評価機構の医療事故報告書に書かれている医療事故内容から、医療事故の「具体的内容」「背景・要因」「改善策」とその他の情報をまとめたものである。
使い方の例は以下に載せる。
English
This dataset is compiled from the medical incident reports published by the Japan Council for Quality Health Care. It summarizes the contents of medical incidents, including the specific details, background and contributing factors, and proposed improvements, along with other related information.
An example of how to use the… See the full description on the dataset page: https://huggingface.co/datasets/Sakaji-Lab/JMID.IndicNLP-GujaratiIndicNLP-Oriyasql-create-context-thai
Overview
This dataset builds from sql-create-context.
@misc{b-mc2_2023_sql-create-context,
title = {sql-create-context Dataset},
author = {b-mc2},
year = {2023},
url = {https://huggingface.co/datasets/b-mc2/sql-create-context},
note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.},
}
databricks-dolly-15k-ja-scoredFor the English version, please click here.
概要
databricks-dolly-15k-ja-scoredはkunishou/databricks-dolly-15k-jaの派生であり、BERTScoreによって提供される翻訳品質スコアが追加されています。
このデータセットは、学術的・商業的問わずクリエイティブ・コモンズ 表示 - 継承 3.0 非移植ライセンスの条件の下で何にでも使用することができます。
翻訳の品質スコア
databricks-dolly-15k-jaは、databricks-dolly-15kを機械翻訳したものです。databricks-dolly-15k-jaに含まれるデータを調べてみると、以下のような品質の悪いデータが存在することが分かりました。
inputとoutputが全く同じであるデータ
outputがinstructionにコピーされているデータ
表記ゆれによって表現の一貫性が保たれていないデータ
固有名詞などの翻訳に失敗しているデータ… See the full description on the dataset page: https://huggingface.co/datasets/sakusakumura/databricks-dolly-15k-ja-scored.IndicNLP-MalayalamVMEBgithub-issuesannotations_creators: []
language:
en
language_creators: []
license: []
multilinguality: []
pretty_name: HuggingFace GitHub Issues
size_categories: []
source_datasets: []
tags: []
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
sakura_japanese_dataset
Sakura_dataset
商用利用可能な超小規模高品質日本語データセット。
categoryは以下
commonsense_qa: 常識問題
Calc-ape210k: 数学問題
japanese-commonsense-openqa: 日本の常識問題(自作)
下記データセットを使用しています。
commonsense_qa
MU-NLPC/Calc-ape210k
LICENSE
This dataset is licensed under Database Contents License (DbCL) v1.0
Update
Last Update : 2023-06-07
Example Code
# モデルの読み込み
import os
from peft.utils.config import TaskType
os.environ["CUDA_VISIBLE_DEVICES"]="0"
import peft
import transformers
import… See the full description on the dataset page: https://huggingface.co/datasets/saldra/sakura_japanese_dataset.
