CoolFace
Modelpublic

Ichsan2895/Eval_Indo_LLM

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes
Model Card

We try to evaluate the LLM model performance in Indonesian. There are many ways for calculate it, for example: BLEU, Perplexity, Human Eval, and GPT4 as Judge. However, In our opinion, we use Perplexity as it was the fastest way for big evaluation dataset. If you have any faster and easier ways for calculating BLUE or any other metrics. Feel free to contribute in this repo.