datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.multilingual-llm-evaluation
Multilingual LLM Evaluation
A small evaluation dataset for comparing language models across English, Hindi, and Spanish.
Columns
language: language code (en, hi, or es)
question: question provided to the model
expected_answer: reference answer used for scoring
Intended use
This dataset can be used to compare model accuracy, language adherence, and response speed across languages.
Limitations
This is a small demonstration dataset and… See the full description on the dataset page: https://huggingface.co/datasets/userhuggingface4321/multilingual-llm-evaluation.clean_multilingual_thinking元データ: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/Multilingual-Thinking
データ件数: 197
平均トークン数: 872
最大トークン数: 2,339
合計トークン数: 171,812
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 748.5 KB
加工内容:
フィルタリングによるデータクレンジング
言語フィルタリング: reasoning_languageが「English」のデータのみを抽出します。
文字数フィルタリング: 処理速度の観点から、question(質問)、thought(思考)、answer(回答)の各フィールドで、規定の文字数を超える長大なデータは事前に除外します。
繰り返し表現の除去:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_multilingual_thinking.
