datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task1577_amazon_reviews_multi_japanese_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.toxicity_multilanguage_datasetamazon-reviews-multi-all-languagesLanguageQA
Dataset Card for SAKURA-LanguageQA
This dataset contains the audio and the single/multi-hop questions/answers of the language track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the language spoken in the speech) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/LanguageQA.multi-open
African Languages Lab Multi-Open
multi-open is the open-source multilingual subset released by the
African Languages Lab. It contains English-target
parallel text for 31 African languages.
Project website: https://the-african-languages-lab.github.io/
The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African
NLPIssaka et al., ACL 2026.
The paper presents All Lab's broader collaborative program: systematic and quality-controlled
data infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/multi-open.mini-multilanguagecode-gen-multi-languagetask1576_amazon_reviews_multi_english_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1576_amazon_reviews_multi_english_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1576_amazon_reviews_multi_english_language_classification.Zeroshot-multilanguages-2.0
Dataset Card for "Zeroshot-multilanguages-2.0"
More Information needed
task1574_amazon_reviews_multi_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.alpaca_multi_language_train_reasoning_simple
このデータセットは、有名なALPACAデータセットの一部を使ったデータセットです。
##日本語と英語(プラス中国語)の同じ情報が記載されています。
「Reasoning」というプロンプトを使えるようにしています。
Reasoningを使用することにより、予測精度を上げられるようにしました。
Reasoningの内容は、回答に使用すべき言語をしているだけです。
必要に応じて、ユーザーが変更してみてください。
「Category」という属性で、レコードが分類されています。
open_qa, closed_qa, classification
brainstorm, creative
translation, question_to_question
「Reasoning」を推論するためのレコードが追加されています。
詳しい情報はこちらのブログを参考にしてください。
multilanguageflan_combined_task1576_amazon_reviews_multi_english_language_classificationtranslation-multilanguage-v2Zeroshot-multilanguages-2.1flan_combined_task1574_amazon_reviews_multi_language_identificationmulti_language_train0625
経緯
このデータセットはQEUプロジェクトのBONSAI2の学習のために開発されました。
特長
alpacaデータセットの一部がベースですが、大幅に変更されています。
言語: 英語、日本語、中国語
reasoningという情報が入っています。使わなくともかまいません。
open_qa, closed_qa, classification, evaluation, question to question
参考サイト
QEUR23_ CHRLTM14 : 閑話休題~Predibaseでfinetuneを使ってみる(SOLAR LLM)
multi-language-messages-01CodeTransOceanのMultilingualTransデータセットのsplit trainをopenAI messages形式に調整。
all-lab-text-multi
