CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lots-of-LoRAs /task1577_amazon_reviews_multi_japanese_language_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.texttext-generationn<1K0 likes163 downloads2y agoHugging Face02esclient /toxicity_multilanguage_datasettext10K<n<100K0 likes124 downloads5mo agoHugging Face03BaoLocTown /amazon-reviews-multi-all-languagestext1M<n<10M0 likes84 downloads2y agoHugging Face04SLLM-multi-hop /LanguageQA Dataset Card for SAKURA-LanguageQA This dataset contains the audio and the single/multi-hop questions/answers of the language track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information". The fields of the dataset are: file: The filename of the audio files. audio: The audio recordings. attribute_label: The attribute labels (i.e., the language spoken in the speech) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/LanguageQA.audion<1K0 likes33 downloads1y agoHugging Face05African-Languages-Lab /multi-opengated African Languages Lab Multi-Open multi-open is the open-source multilingual subset released by the African Languages Lab. It contains English-target parallel text for 31 African languages. Project website: https://the-african-languages-lab.github.io/ The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLPIssaka et al., ACL 2026. The paper presents All Lab's broader collaborative program: systematic and quality-controlled data infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/multi-open.tabulartranslation10M<n<100M2 likes29 downloads3mo agoHugging Face06akahana /mini-multilanguagetext1M<n<10M0 likes28 downloads2y agoHugging Face07kisejin /code-gen-multi-languagetext1K<n<10K0 likes26 downloads2y agoHugging Face08Lots-of-LoRAs /task1576_amazon_reviews_multi_english_language_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1576_amazon_reviews_multi_english_language_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1576_amazon_reviews_multi_english_language_classification.texttext-generationn<1K0 likes24 downloads2y agoHugging Face09Weni /Zeroshot-multilanguages-2.0 Dataset Card for "Zeroshot-multilanguages-2.0" More Information needed text10K<n<100K0 likes20 downloads3y agoHugging Face10Lots-of-LoRAs /task1574_amazon_reviews_multi_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.texttext-generationn<1K0 likes18 downloads2y agoHugging Face11QEU /alpaca_multi_language_train_reasoning_simple このデータセットは、有名なALPACAデータセットの一部を使ったデータセットです。 ##日本語と英語(プラス中国語)の同じ情報が記載されています。 「Reasoning」というプロンプトを使えるようにしています。 Reasoningを使用することにより、予測精度を上げられるようにしました。 Reasoningの内容は、回答に使用すべき言語をしているだけです。 必要に応じて、ユーザーが変更してみてください。 「Category」という属性で、レコードが分類されています。 open_qa, closed_qa, classification brainstorm, creative translation, question_to_question 「Reasoning」を推論するためのレコードが追加されています。 詳しい情報はこちらのブログを参考にしてください。 text1K<n<10K0 likes17 downloads2y agoHugging Face12akahana /multilanguagetabular10M<n<100M0 likes15 downloads2y agoHugging Face13supergoose /flan_combined_task1576_amazon_reviews_multi_english_language_classificationtextn<1K0 likes14 downloads2y agoHugging Face14mlech26l /translation-multilanguage-v2textn<1K0 likes11 downloads5mo agoHugging Face15Weni /Zeroshot-multilanguages-2.1text10K<n<100K0 likes10 downloads3y agoHugging Face16supergoose /flan_combined_task1574_amazon_reviews_multi_language_identificationtextn<1K0 likes8 downloads2y agoHugging Face17QEU /multi_language_train0625 経緯 このデータセットはQEUプロジェクトのBONSAI2の学習のために開発されました。 特長 alpacaデータセットの一部がベースですが、大幅に変更されています。 言語: 英語、日本語、中国語 reasoningという情報が入っています。使わなくともかまいません。 open_qa, closed_qa, classification, evaluation, question to question 参考サイト QEUR23_ CHRLTM14 : 閑話休題~Predibaseでfinetuneを使ってみる(SOLAR LLM) text1K<n<10K0 likes6 downloads2y agoHugging Face18kazuyamaa /multi-language-messages-01CodeTransOceanのMultilingualTransデータセットのsplit trainをopenAI messages形式に調整。 text10K<n<100K0 likes4 downloads2y agoHugging Face19African-Languages-Lab /all-lab-text-multigatedtabular100M<n<1B1 likes2 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.