datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
joyo-kanji-yomi-benchmark
Joyo Kanji Yomi Benchmark
A kanji-level pronunciation evaluation benchmark for Japanese TTS, covering all 2,136 Joyo kanji and their 4,378 readings with 13,095 native-speaker-verified test sentences.
Dataset Description
Each sample targets a specific kanji-reading pair. The sentence context is designed so that only the target reading is valid. All sentences and annotations have been verified by 35 native Japanese speakers through a three-stage review process.… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/joyo-kanji-yomi-benchmark.J-KGMHQA
J-KGMHQA (Japanese KG-groundedMultiHopQA)
本データセットは以下の2つのファイルを含む.
eval.jsonl: マルチホップQAの評価データ
corpus.jsonl: Wikipeidaダンプ(20241023)に対して,回答根拠を含む合成文書を,対応する記事と置換した文書コーパス.
QAデータ(eval.jsonl)
各QAサンプルは,以下の構造を持つ.
type: マルチホップQAのタイプ
triples_readable: 推論パスのトリプル
question: 質問文
answer: 正答
reasoning_path: 推論パス
derivations: 根拠となるWikipeida記事
category: 初期エンティティのカテゴリ(person, place, facility, organization, work)
hop: 回答に到達するために必要な推論ステップ数
Example
{
"type": "compositional"… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/J-KGMHQA.japanese-multiturn-safety
JMT-Safety
概要
JMT-Safety は、日本語マルチターン対話における安全性評価のためのデータセットです。
このデータセットは、llm-jp/AnswerCarefullyを拡張して構築されたものです。
LLM-jp 様より許諾を得て公開しており、ライセンス(利用規約)は AnswerCarefully から継承しています(日: LICENSE_ja, 英: LICENSE)。
データセットの詳細については、こちらの論文をご覧ください。
本データセットを用いた評価のためのスクリプトを GitHub で公開しています。ご活用ください。
Overview
JMT-Safety is a dataset designed for safety evaluation in Japanese multi-turn conversations.
This dataset was built by extending llm-jp/AnswerCarefully.
It is… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/japanese-multiturn-safety.sbi-small-dataset-qwen3-14Bdefendable-pain-sbir-phase1-pain-v0.1
SBIR Phase I Pain Receipt
"the $250K dream" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 13 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
13 pain… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-sbir-phase1-pain-v0.1.Sbi-Dataset
