CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bjoernp /gaps_jpn Dataset Card for "gaps_jpn" More Information needed text100M<n<1B1 likes530 downloads3y agoHugging Face02asahi417 /seamless-align-enA-jpnaudio100K<n<1M0 likes258 downloads2y agoHugging Face03AdaMLLab /JpnMix JpnMix (https://arxiv.org/abs/2512.18834) is a Japanese pretraining corpus built by combining five publicly available Japanese datasets, applying Japanese-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/JpnMix.texttext-generation100M<n<1B2 likes176 downloads5mo agoHugging Face04sappho192 /Tatoeba-Challenge-jpn-kor Dataset Card for Dataset Name This dataset contains Japanese-Korean paired text which is from Helsinki-NLP/Tatoeba-Challenge. Dataset Details Dataset Sources Repository: Helsinki-NLP/Tatoeba-Challenge Detail: Japanese - Korean jpn-kor Uses The dataset can be used to train the translation model that translates Japanese sentence to Korean. Out-of-Scope Use You cannot use this dataset to train the model which is to be used under commercial… See the full description on the dataset page: https://huggingface.co/datasets/sappho192/Tatoeba-Challenge-jpn-kor.texttranslation10M<n<100M0 likes126 downloads3y agoHugging Face05fpadovani /goldfish-jpn-jpan-100mb-tokenized100K<n<1M0 likes98 downloads4mo agoHugging Face06yiyic /cmn_jpn_traintext1M<n<10M0 likes94 downloads2y agoHugging Face07AdoCleanCode /SPEEED_s3_words_fr_ger_jpn_spa_multitext1M<n<10M0 likes92 downloads8mo agoHugging Face08abeja /OpenAI-MRCR-Translation-JPN本データセットは、ロングコンテキス評価データセット OpenAI MRCR を翻訳して作成した、日本語版の MRCR の評価データセットです。 翻訳方法 日本語版の作成にあたっては、元の英語データの構造を維持しつつ、LLM(Qwen3-235B-A22B)を用いて日本語版のテキストを生成しました。 作成方法としては、 会話履歴を分解 パーツごとに翻訳 元の順番に会話を並べなおす といった手順により、日本語版の評価サンプルを作成しました。 翻訳が必要だったのは主にタスクの説明文、ユーザの問い合わせ文章、ユーザの最終問い合わせ文章の3箇所です。 タスクの説明文については固定のプロンプトなので、プロンプト全体を一度だけ翻訳しました。 ユーザの問い合わせ文については、全てのユーザの問い合わせが「write a (Document-Type) about (Genre)」という形式の英文になっていたため、Document-Type, Genre の位置の語句を抜き出して翻訳し「(Genre) についての (Document-Type)… See the full description on the dataset page: https://huggingface.co/datasets/abeja/OpenAI-MRCR-Translation-JPN.tabular1K<n<10K2 likes90 downloads7mo agoHugging Face09CAiRE /prosocial-dialog-jpn_Jpantabular100K<n<1M0 likes81 downloads3y agoHugging Face10kamilakesbi /mixture_ami_jpn_synthetic_Bigaudio1K<n<10K0 likes79 downloads2y agoHugging Face11kamilakesbi /mixture_callhome_jpn_syntheticaudio1K<n<10K0 likes61 downloads2y agoHugging Face12kwatcharasupat /dnr-v3-jpnaudio0 likes60 downloads11mo agoHugging Face13YYama0 /CT-RATE-JPN CT-RATE-JPN Dataset CT-RATE-JPN is a Japanese-translated version of radiology reports from the CT-RATE dataset, which contains chest CT volumes paired with corresponding radiology reports. Dataset Overview CT-RATE-JPN provides Japanese translations of radiology reports from the CT-RATE dataset to facilitate Japanese medical AI model development. While the original CT-RATE contains 25,692 non-contrast chest CT volumes with corresponding reports, this repository focuses on… See the full description on the dataset page: https://huggingface.co/datasets/YYama0/CT-RATE-JPN.10K<n<100K0 likes57 downloads1y agoHugging Face14kamilakesbi /synthetic_dataset_jpn_2_more_speakersaudio1K<n<10K0 likes46 downloads2y agoHugging Face15MarcoDotIO /jpn-bench JPN-Bench JPN-Bench is a Japanese literacy benchmark for tokenizer evaluation and future Japanese LLM evaluation. This public release contains a small curated tokenizer-literacy dev set plus benchmark-lane source material manifests kept separate from tokenizer training material. This dataset is grouped with the KotodamaLM tokenizer work in the Hugging Face collection "KotodamaLM Japanese Language Infrastructure". Files data/literacy_items.jsonl: 60 tokenizer-literacy… See the full description on the dataset page: https://huggingface.co/datasets/MarcoDotIO/jpn-bench.tabulartext-generation10K<n<100K0 likes39 downloads5mo agoHugging Face16mitsudate /DiffSinger_opencpop_JPN0 likes37 downloads3y agoHugging Face17kamilakesbi /synthetic_dataset_jpnaudio1K<n<10K0 likes35 downloads2y agoHugging Face18kamilakesbi /synthetic_dataset_jpn_volumesaudio1K<n<10K0 likes34 downloads2y agoHugging Face19wmt-pischool /wmt25-jpn-zhotext1M<n<10M1 likes34 downloads1y agoHugging Face20autoprogrammer /nllb_eng_Latn_jpn_Jpan_subset_15ktext10K<n<100K0 likes24 downloads2y agoHugging Face21kamilakesbi /synthetic_dataset_jpn_2audio1K<n<10K0 likes20 downloads2y agoHugging Face22Jpnm89 /MentalChat16K 🗣️ Synthetic Counseling Conversations Dataset 📝 Description Synthetic Data 10K This dataset consists of 9,775 synthetic conversations between a counselor and a client, covering 33 mental health topics such as 💑 Relationships, 😟 Anxiety, 😔 Depression, 🤗 Intimacy, and 👨‍👩‍👧‍👦 Family Conflict. The conversations were generated using the OpenAI GPT-3.5 Turbo model and a customized adaptation of the Airoboros self-generation framework. The Airoboros… See the full description on the dataset page: https://huggingface.co/datasets/Jpnm89/MentalChat16K.textquestion-answering10K<n<100K0 likes19 downloads8mo agoHugging Face23davanstrien /fineweb-c-jpn-previewtextn<1K0 likes18 downloads2y agoHugging Face24open-llm-leaderboard /google__gemma-2-2b-jpn-it-detailsgated Dataset Card for Evaluation run of google/gemma-2-2b-jpn-it Dataset automatically created during the evaluation run of model google/gemma-2-2b-jpn-it The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2-2b-jpn-it-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face25kamilakesbi /synthetic_dataset_jpn_2_bigaudio1K<n<10K0 likes16 downloads2y agoHugging Face26kamilakesbi /mixture_callfriend_jpn_syntheticaudio1K<n<10K0 likes16 downloads2y agoHugging Face27Alsebay /JP-novel-selected-datasettext1K<n<10K3 likes16 downloads2y agoHugging Face28open-llm-leaderboard /ymcki__gemma-2-2b-jpn-it-abliterated-18-detailsgated Dataset Card for Evaluation run of ymcki/gemma-2-2b-jpn-it-abliterated-18 Dataset automatically created during the evaluation run of model ymcki/gemma-2-2b-jpn-it-abliterated-18 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ymcki__gemma-2-2b-jpn-it-abliterated-18-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face29Respair /Style_vectors_JPNtext1K<n<10K0 likes13 downloads2y agoHugging Face30deerfieldgreen /cpi_jpnn<1K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.