datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_zeroshot_mixtral_v0.1alpaca-french-mixtral
License & Attribution
MTEB-format derivative of AIffl/Alpaca_french_mixtral (French Alpaca, Mixtral-translated). Query = instruction; corpus = answer. Deterministically subsampled to ~10k. Licensed under Apache-2.0 (same as source).
3_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationmixtraltoken_fineweb_edu_mini_combinedopenhermes-dev__mistralai_Mixtral-8x7B-Instruct-v0.1__1707245027KGQA_Mixtral
Dataset Card for "KGQA_Mixtral"
More Information needed
details_cognitivecomputations__dolphin-2.7-mixtral-8x7b
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.7-mixtral-8x7b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.7-mixtral-8x7b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_cognitivecomputations__dolphin-2.7-mixtral-8x7b.1_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationmistralai__Mixtral-8x22B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x22B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x22B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x22B-v0.1-details.2_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationvicgalle__Merge-Mixtral-Prometheus-8x7B-details
Dataset Card for Evaluation run of vicgalle/Merge-Mixtral-Prometheus-8x7B
Dataset automatically created during the evaluation run of model vicgalle/Merge-Mixtral-Prometheus-8x7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__Merge-Mixtral-Prometheus-8x7B-details.mistralai__Mixtral-8x7B-Instruct-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x7B-Instruct-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x7B-Instruct-v0.1
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x7B-Instruct-v0.1-details.LeroyDyer__Mixtral_AI_SwahiliTron_7b-details
Dataset Card for Evaluation run of LeroyDyer/Mixtral_AI_SwahiliTron_7b
Dataset automatically created during the evaluation run of model LeroyDyer/Mixtral_AI_SwahiliTron_7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__Mixtral_AI_SwahiliTron_7b-details.mistralai__Mixtral-8x7B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x7B-v0.1
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x7B-v0.1-details.NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details.details_cognitivecomputations__dolphin-2.6-mixtral-8x7b
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.6-mixtral-8x7b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.6-mixtral-8x7b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_cognitivecomputations__dolphin-2.6-mixtral-8x7b.mistral-community__mixtral-8x22B-v0.3-details
Dataset Card for Evaluation run of mistral-community/mixtral-8x22B-v0.3
Dataset automatically created during the evaluation run of model mistral-community/mixtral-8x22B-v0.3
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistral-community__mixtral-8x22B-v0.3-details.details_cognitivecomputations__dolphin-2.9.1-mixtral-1x22b
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.1-mixtral-1x22b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.1-mixtral-1x22b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_cognitivecomputations__dolphin-2.9.1-mixtral-1x22b.NousResearch__Nous-Hermes-2-Mixtral-8x7B-SFT-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
The dataset is composed of 78 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-SFT-details.clean_pubmedqa_mixtral_cot元データ: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/PubmedQA-Mixtral-CoT
データ件数: 206,962
平均トークン数: 586
最大トークン数: 1,922
合計トークン数: 121,366,170
ファイル形式: JSONL
ファイル分割数: 3
合計ファイルサイズ: 532.2 MB
加工内容:
文字数によるフィルタリング:
question (質問) 列の文字数が 6,000文字を超える データを削除します。
response (応答) 列の文字数が 80,000文字を超える データを削除します。
応答 (response) の分割:
response 列を、思考プロセスを記述した「thought」部分と、最終的な結論である「answer」部分に分割します。
分割には Answer: や The answer… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_pubmedqa_mixtral_cot.abacusai__Smaug-Mixtral-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-Mixtral-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-Mixtral-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Mixtral-v0.1-details.mistral-community__Mixtral-8x22B-v0.1-details
Dataset Card for Evaluation run of mistral-community/Mixtral-8x22B-v0.1
Dataset automatically created during the evaluation run of model mistral-community/Mixtral-8x22B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistral-community__Mixtral-8x22B-v0.1-details.Translation-deepseek-llama-mixtral-v-deepl
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
This dataset contains ~51k responses from ~11k annotators and compares the translation capabilities of DeepSeek-R1(deepseek-r1-distill-llama-70b-specdec), Llama(llama-3.3-70b-specdec) and Mixtral(mixtral-8x7b-32768) against DeepL across different languages. The comparison involved 100 distinct questions in 4 languages, with each translation being rated by 51 native… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Translation-deepseek-llama-mixtral-v-deepl.cloudyu__Mixtral_34Bx2_MoE_60B-details
Dataset Card for Evaluation run of cloudyu/Mixtral_34Bx2_MoE_60B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_34Bx2_MoE_60B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Mixtral_34Bx2_MoE_60B-details.mistralai__Mixtral-8x7B-Instruct-v0.1research-papers-dataset-mixtral7B-processed2
Research Papers Dataset - Processed with Train/Test/Valid Splits
This dataset contains preprocessed research papers with the following enhancements, split into train/test/validation sets.
Dataset Splits:
Train: 7,328 entries (85.0%)
Test: 431 entries (5.0%)
Valid: 863 entries (10.0%)
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed2.GPT4-Mixtral-MMLU-Preference-Complexity-trainoasst1_for_self-rewarding_AIFT_Mixtral-8x22B-Instructdetails_cloudyu__Mixtral_34Bx2_MoE_60B
Dataset Card for Evaluation run of cloudyu/Mixtral_34Bx2_MoE_60B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_34Bx2_MoE_60B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_cloudyu__Mixtral_34Bx2_MoE_60B.details_NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO.
