CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sakib323 /GUI_BASED_PLATFORMtext100K<n<1M0 likes3.2k downloads7mo agoHugging Face02Sakshamrzt /IndicNLP-Multilingualtexttext-classification10K<n<100K1 likes331 downloads2y agoHugging Face03Sakalti /Multilingal-sakalt-dataマルチリンガルデータセットです。mitライセンスです。 texttext-generation1K<n<10K1 likes111 downloads2y agoHugging Face04SakanaAI /gsm8k-ja-test_250-1319 gsm8k-ja-test_250-1319 This dataset contains 1069 Japanese math problems and their solutions. It was used for optimizing LLMs in the paper "Evolutionary Optimization of Model Merging Recipes". Dataset Details This dataset contains Japanese translations of 1069 math problems and solutions from the GSM8K test set, starting from the 251st example out of 1319. The translation was done using gpt-4-0125-preview. We did not use the first 250 examples because they are part of the… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/gsm8k-ja-test_250-1319.text1K<n<10K5 likes105 downloads2y agoHugging Face05open-llm-leaderboard /PocketDoc__Dans-SakuraKaze-V1.0.0-12b-detailsgated Dataset Card for Evaluation run of PocketDoc/Dans-SakuraKaze-V1.0.0-12b Dataset automatically created during the evaluation run of model PocketDoc/Dans-SakuraKaze-V1.0.0-12b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PocketDoc__Dans-SakuraKaze-V1.0.0-12b-details.tabular10K<n<100K0 likes66 downloads2y agoHugging Face06Sakshamrzt /IndicNLP-Tamiltexttext-classification10K<n<100K0 likes60 downloads2y agoHugging Face07Sakaji-Lab /LATGNJ JAgriN: Japanese Agricultural Dataset of Nagasaki Prefecture formerly LATGNJ: Local Agricultural Technical Guideline of Nagasaki, Japan Dataset Metadata (Datasheet Summary) This section summarizes the key metadata of JAgriN following the recommendations proposed in "Datasheets for Datasets" by Gebru et al. (2021) [1]. Field Description Dataset Name JAgriN (Japanese Agricultural Dataset of Nagasaki Prefecture) Creators Hokkaido University, The University of… See the full description on the dataset page: https://huggingface.co/datasets/Sakaji-Lab/LATGNJ.documentn<1K1 likes58 downloads1y agoHugging Face08Sakshamrzt /medical_qa Dataset Card for Dataset Name Dataset Details The MedQuad dataset normalised for use with mteb. The dataset contains questions and answers related to medical conditions, treatments, and protocols Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More Information Needed] Out-of-Scope Use [More… See the full description on the dataset page: https://huggingface.co/datasets/Sakshamrzt/medical_qa.texttable-question-answering1K<n<10K1 likes51 downloads2y agoHugging Face09sakkke /text-to-command-geminitextn<1K1 likes48 downloads3y agoHugging Face10open-llm-leaderboard /Sakalti__ultiima-32B-detailsgated Dataset Card for Evaluation run of Sakalti/ultiima-32B Dataset automatically created during the evaluation run of model Sakalti/ultiima-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-32B-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face11open-llm-leaderboard /Sakalti__SJT-7.5B-detailsgated Dataset Card for Evaluation run of Sakalti/SJT-7.5B Dataset automatically created during the evaluation run of model Sakalti/SJT-7.5B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-7.5B-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face12open-llm-leaderboard /Sakalti__ultiima-72B-v1.5-detailsgated Dataset Card for Evaluation run of Sakalti/ultiima-72B-v1.5 Dataset automatically created during the evaluation run of model Sakalti/ultiima-72B-v1.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-72B-v1.5-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face13Sakshamrzt /IndicNLP-Punjabitext1K<n<10K1 likes44 downloads2y agoHugging Face14open-llm-leaderboard /Sakalti__Saka-7.2B-detailsgated Dataset Card for Evaluation run of Sakalti/Saka-7.2B Dataset automatically created during the evaluation run of model Sakalti/Saka-7.2B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Saka-7.2B-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face15open-llm-leaderboard /Sakalti__Neptuno-Alpha-detailsgated Dataset Card for Evaluation run of Sakalti/Neptuno-Alpha Dataset automatically created during the evaluation run of model Sakalti/Neptuno-Alpha The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Neptuno-Alpha-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face16open-llm-leaderboard /Sakalti__SJT-3.7B-detailsgated Dataset Card for Evaluation run of Sakalti/SJT-3.7B Dataset automatically created during the evaluation run of model Sakalti/SJT-3.7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-3.7B-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face17open-llm-leaderboard /Sakalti__ultiima-14B-v0.4-detailsgated Dataset Card for Evaluation run of Sakalti/ultiima-14B-v0.4 Dataset automatically created during the evaluation run of model Sakalti/ultiima-14B-v0.4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-14B-v0.4-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face18open-llm-leaderboard /Sakalti__SJT-8B-V1.1-detailsgated Dataset Card for Evaluation run of Sakalti/SJT-8B-V1.1 Dataset automatically created during the evaluation run of model Sakalti/SJT-8B-V1.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-8B-V1.1-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face19sakerhetspolisen /selma_ngramstext1M<n<10M0 likes39 downloads16d agoHugging Face20Sakshamrzt /IndicNLP-Kannadatexttext-classification10K<n<100K0 likes36 downloads2y agoHugging Face21Sakaji-Lab /JaFIn For citation: @preprint{tanabe2024-jafin, title={{JaFIn: Japanese Financial Instruction Dataset}}, author={Kota Tanabe, Masahiro Suzuki, Hiroki Sakaji, Itsuki Noda}, year={2024}, doi={10.48550/arXiv.2404.09260}, } License: cc-by-nc-sa-4.0 text1K<n<10K4 likes35 downloads2y agoHugging Face22Sakaji-Lab /JMID JMID: Japanese Medical Incident Dataset 日本語 本データセットは、公益財団法人日本医療機能評価機構の医療事故報告書に書かれている医療事故内容から、医療事故の「具体的内容」「背景・要因」「改善策」とその他の情報をまとめたものである。 使い方の例は以下に載せる。 English This dataset is compiled from the medical incident reports published by the Japan Council for Quality Health Care. It summarizes the contents of medical incidents, including the specific details, background and contributing factors, and proposed improvements, along with other related information. An example of how to use the… See the full description on the dataset page: https://huggingface.co/datasets/Sakaji-Lab/JMID.texttext-classification1K<n<10K2 likes35 downloads1y agoHugging Face23Sakshamrzt /IndicNLP-Gujaratitexttext-classification1K<n<10K0 likes32 downloads2y agoHugging Face24Sakshamrzt /IndicNLP-Oriyatexttext-classification10K<n<100K0 likes30 downloads2y agoHugging Face25saksornr /sql-create-context-thai Overview This dataset builds from sql-create-context. @misc{b-mc2_2023_sql-create-context, title = {sql-create-context Dataset}, author = {b-mc2}, year = {2023}, url = {https://huggingface.co/datasets/b-mc2/sql-create-context}, note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.}, } texttext-generation10K<n<100K0 likes30 downloads2y agoHugging Face26sakusakumura /databricks-dolly-15k-ja-scoredFor the English version, please click here. 概要 databricks-dolly-15k-ja-scoredはkunishou/databricks-dolly-15k-jaの派生であり、BERTScoreによって提供される翻訳品質スコアが追加されています。 このデータセットは、学術的・商業的問わずクリエイティブ・コモンズ 表示 - 継承 3.0 非移植ライセンスの条件の下で何にでも使用することができます。 翻訳の品質スコア databricks-dolly-15k-jaは、databricks-dolly-15kを機械翻訳したものです。databricks-dolly-15k-jaに含まれるデータを調べてみると、以下のような品質の悪いデータが存在することが分かりました。 inputとoutputが全く同じであるデータ outputがinstructionにコピーされているデータ 表記ゆれによって表現の一貫性が保たれていないデータ 固有名詞などの翻訳に失敗しているデータ… See the full description on the dataset page: https://huggingface.co/datasets/sakusakumura/databricks-dolly-15k-ja-scored.textquestion-answering10K<n<100K6 likes29 downloads3y agoHugging Face27Sakshamrzt /IndicNLP-Malayalamtexttext-classification1K<n<10K0 likes29 downloads2y agoHugging Face28Sakabamrisa /VMEBtext1K<n<10K1 likes29 downloads10mo agoHugging Face29SakaiJun /github-issuesannotations_creators: [] language: en language_creators: [] license: [] multilinguality: [] pretty_name: HuggingFace GitHub Issues size_categories: [] source_datasets: [] tags: [] task_categories: text-classification text-retrieval task_ids: multi-class-classification multi-label-classification document-retrieval tabular1K<n<10K0 likes27 downloads4y agoHugging Face30saldra /sakura_japanese_dataset Sakura_dataset 商用利用可能な超小規模高品質日本語データセット。 categoryは以下 commonsense_qa: 常識問題 Calc-ape210k: 数学問題 japanese-commonsense-openqa: 日本の常識問題(自作) 下記データセットを使用しています。 commonsense_qa MU-NLPC/Calc-ape210k LICENSE This dataset is licensed under Database Contents License (DbCL) v1.0 Update Last Update : 2023-06-07 Example Code # モデルの読み込み import os from peft.utils.config import TaskType os.environ["CUDA_VISIBLE_DEVICES"]="0" import peft import transformers import… See the full description on the dataset page: https://huggingface.co/datasets/saldra/sakura_japanese_dataset.textquestion-answeringn<1K20 likes27 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.