retrieva-jp/japanese-spoken-language-bert
Model Card for japanese-spoken-language-bert
日本語READMEはこちら
<!-- Provide a quick summary of what the model is/does. [Optional] --> These BERT models are pre-trained on written Japanese (Wikipedia) and fine-tuned on Spoken Japanese. We used CSJ and the Japanese diet record. CSJ (Corpus of Spontaneous Japanese) is provided by NINJAL (https://www.ninjal.ac.jp/). We only provide model parameters. You have to download other config files to use these models.
We provide three models down below:
- 1-6 layer-wise (Folder Name: models/1-6_layer-wise) Fine-Tuned only 1st-6th layers in Encoder on CSJ.
- TAPT512 60k (Folder Name: models/tapt512_60k) Fine-Tuned on CSJ.
- DAPT128-TAPT512 (Folder Name: models/dapt128-tap512) Fine-Tuned on the diet record and CSJ.
Table of Contents
- Model Card for japanese-spoken-language-bert
- Table of Contents
- Model Details
- Model Description
- Training Details
- Training Data
- Training Procedure
- Evaluation
- Testing Data, Factors & Metrics
- Testing Data
- Factors
- Metrics
- Results
- Citation
- More Information
- Model Card Authors
- Model Card Contact
- How to Get Started with the Model
Model Details
Model Description
<!-- Provide a longer summary of what this model is/does. --> These BERT models are pre-trained on written Japanese (Wikipedia) and fine-tuned on Spoken Japanese. We used CSJ and the Japanese diet record. CSJ (Corpus of Spontaneous Japanese) is provided by NINJAL (https://www.ninjal.ac.jp/). We only provide model parameters. You have to download other config files to use these models.
We provide three models down below:
- 1-6 layer-wise (Folder Name: models/1-6_layer-wise) Fine-Tuned only 1st-6th layers in Encoder on CSJ.
- TAPT512 60k (Folder Name: models/tapt512_60k) Fine-Tuned on CSJ.
- DAPT128-TAPT512 (Folder Name: models/dapt128-tap512) Fine-Tuned on the diet record and CSJ.
Model Information
- Model type: Language model
- Language(s) (NLP): ja
- License: Copyright (c) 2021 National Institute for Japanese Language and Linguistics and Retrieva, Inc. Licensed under the Apache License, Version 2.0 (the “License”)
Training Details
Training Data
<!-- This should link to a Data Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
- 1-6 layer-wise: CSJ
- TAPT512 60K: CSJ
- DAPT128-TAPT512: The Japanese diet record and CSJ
Training Procedure
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
We continuously train the pre-trained Japanese BERT model (cl-tohoku/bert-base-japanese-whole-word-masking; written BERT).
In detail, see Japanese blog or Japanese paper.
Evaluation
<!-- This section describes the evaluation protocols and provides the results. -->
Testing Data, Factors & Metrics
Testing Data
<!-- This should link to a Data Card if possible. -->
We use CSJ for the evaluation.
Factors
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
We evaluate the following tasks on CSJ:
- Dependency Parsing
- Sentence Boundary
- Important Sentence Extraction
Metrics
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
- Dependency Parsing: Undirected Unlabeled Attachment Score (UUAS)
- Sentence Boundary: F1 Score
- Important Sentence Extraction: F1 Score
Results
Citation
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
BibTeX:
@inproceedings{csjbert2021,
title = {CSJを用いた日本語話し言葉BERTの作成},
author = {勝又智 and 坂田大直},
booktitle = {言語処理学会第27回年次大会},
year = {2021},
}More Information
https://tech.retrieva.jp/entry/2021/04/01/114943 (In Japanese)
Model Card Authors
<!-- This section provides another layer of transparency and accountability. Whose views is this model card representing? How many voices were included in its construction? Etc. -->
Satoru Katsumata
Model Card Contact
pr@retrieva.jp
How to Get Started with the Model
Use the code below to get started with the model.
<details> <summary> Click to expand </summary>
- Run downloadwikipediabert.py to download BERT model which is trained on Wikipedia.
python download_wikipedia_bert.pyThis script downloads config files and a vocab file provided by Inui Laboratory of Tohoku University from Hugging Face Model Hub. https://github.com/cl-tohoku/bert-japanese
- Run sample_mlm.py to confirm you can use our models.
python sample_mlm.py</details>
