datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.multilingual-thinking-cleaned-chunked-1024yuchenxie__ArlowGPT-3B-Multilingual-details
Dataset Card for Evaluation run of yuchenxie/ArlowGPT-3B-Multilingual
Dataset automatically created during the evaluation run of model yuchenxie/ArlowGPT-3B-Multilingual
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yuchenxie__ArlowGPT-3B-Multilingual-details.lightblue__suzume-llama-3-8B-multilingual-orpo-borda-top25-details
Dataset Card for Evaluation run of lightblue/suzume-llama-3-8B-multilingual-orpo-borda-top25
Dataset automatically created during the evaluation run of model lightblue/suzume-llama-3-8B-multilingual-orpo-borda-top25
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lightblue__suzume-llama-3-8B-multilingual-orpo-borda-top25-details.maywell__Qwen2-7B-Multilingual-RP-details
Dataset Card for Evaluation run of maywell/Qwen2-7B-Multilingual-RP
Dataset automatically created during the evaluation run of model maywell/Qwen2-7B-Multilingual-RP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/maywell__Qwen2-7B-Multilingual-RP-details.lightblue__suzume-llama-3-8B-multilingual-orpo-borda-half-details
Dataset Card for Evaluation run of lightblue/suzume-llama-3-8B-multilingual-orpo-borda-half
Dataset automatically created during the evaluation run of model lightblue/suzume-llama-3-8B-multilingual-orpo-borda-half
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lightblue__suzume-llama-3-8B-multilingual-orpo-borda-half-details.lightblue__suzume-llama-3-8B-multilingual-orpo-borda-full-details
Dataset Card for Evaluation run of lightblue/suzume-llama-3-8B-multilingual-orpo-borda-full
Dataset automatically created during the evaluation run of model lightblue/suzume-llama-3-8B-multilingual-orpo-borda-full
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lightblue__suzume-llama-3-8B-multilingual-orpo-borda-full-details.lightblue__suzume-llama-3-8B-multilingual-orpo-borda-top75-details
Dataset Card for Evaluation run of lightblue/suzume-llama-3-8B-multilingual-orpo-borda-top75
Dataset automatically created during the evaluation run of model lightblue/suzume-llama-3-8B-multilingual-orpo-borda-top75
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lightblue__suzume-llama-3-8B-multilingual-orpo-borda-top75-details.lightblue__suzume-llama-3-8B-multilingual-details
Dataset Card for Evaluation run of lightblue/suzume-llama-3-8B-multilingual
Dataset automatically created during the evaluation run of model lightblue/suzume-llama-3-8B-multilingual
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lightblue__suzume-llama-3-8B-multilingual-details.clean_multilingual_thinking元データ: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/Multilingual-Thinking
データ件数: 197
平均トークン数: 872
最大トークン数: 2,339
合計トークン数: 171,812
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 748.5 KB
加工内容:
フィルタリングによるデータクレンジング
言語フィルタリング: reasoning_languageが「English」のデータのみを抽出します。
文字数フィルタリング: 処理速度の観点から、question(質問)、thought(思考)、answer(回答)の各フィールドで、規定の文字数を超える長大なデータは事前に除外します。
繰り返し表現の除去:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_multilingual_thinking.multi-lingual-llm
Dataset Card for Dataset Name
This data set contains set questions in tamil and possible answers, with the correct answer in the column. This helps to test LLM to see for accuracy.
Dataset Details
Dataset Description
Curated by: Madhumitha Sivalingapandian
Language(s) (NLP): Tamil and English
License: [More Information Needed]
Uses
Used for measuring performance of LLM.
Direct Use
Spoken language accuracy measurement dataset… See the full description on the dataset page: https://huggingface.co/datasets/MithuSi/multi-lingual-llm.
