datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ryukyuan-okinawan-corpus
琉球語(沖縄語)↔ 日本語 対訳コーパス
概要
国立国語研究所が公開する『沖縄語辞典』索引篇(CC BY 4.0)を構造化した、
日本語 → 琉球語(首里方言) の対訳データセットです。
琉球語はUNESCOが「消滅危機言語」に分類しており、
話者数は現在数百人以下と推定されています。
NLP・機械学習向けのオープンデータとしては世界的にも極めて希少な資源です。
データ統計
項目
数値
総レコード数
9,253件
漢字表記あり
8,249件
品詞・文法注記あり
295件
言語コード(琉球語)
ryu (ISO 639-3)
言語コード(日本語)
jpn (ISO 639-3)
方言
首里方言(沖縄本島中南部)
フォーマット
ryukyuan_japanese_index.csv / .jsonl
列名
説明
例
id
レコードID
ninjal_sakuin_00001
japanese… See the full description on the dataset page: https://huggingface.co/datasets/doraking/ryukyuan-okinawan-corpus.nihongo-legal-finance-autoscientist-data
nihongo-legal-finance-autoscientist
AutoScientist Challenge entry dataset for Japanese expert QA in the language category.
Intended Use
This dataset is designed for supervised fine-tuning of Japanese assistants that explain
legal and financial concepts with uncertainty, source-awareness, and non-advice caveats.
Columns
instruction: user task
context: background information
response: target answer
rubric: quality expectations
category: subdomain… See the full description on the dataset page: https://huggingface.co/datasets/doraking/nihongo-legal-finance-autoscientist-data.
