datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mistralai-tekken-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
lm-eval-results-BarraHome-Mistroll-7B-v2.2-private
Dataset Card for Evaluation run of BarraHome/Mistroll-7B-v2.2
Dataset automatically created during the evaluation run of model BarraHome/Mistroll-7B-v2.2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-BarraHome-Mistroll-7B-v2.2-private.mistral-675b-eval-logs-and-scoreslm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private
Dataset Card for Evaluation run of chihoonlee10/T3Q-Mistral-Orca-Math-DPO
Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-Mistral-Orca-Math-DPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private.lm-eval-results-chlee10-T3Q-Merge-Mistral7B-private
Dataset Card for Evaluation run of chlee10/T3Q-Merge-Mistral7B
Dataset automatically created during the evaluation run of model chlee10/T3Q-Merge-Mistral7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chlee10-T3Q-Merge-Mistral7B-private.lm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private
Dataset Card for Evaluation run of nbeerbower/bophades-mistral-truthy-DPO-7B
Dataset automatically created during the evaluation run of model nbeerbower/bophades-mistral-truthy-DPO-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private.mistralai__Mistral-7B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1
The dataset is composed of 82 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 34 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-v0.1-details.lm-eval-results-teknium-OpenHermes-2.5-Mistral-7B-private
Dataset Card for Evaluation run of teknium/OpenHermes-2.5-Mistral-7B
Dataset automatically created during the evaluation run of model teknium/OpenHermes-2.5-Mistral-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-teknium-OpenHermes-2.5-Mistral-7B-private.lm-eval-results-nbeerbower-bophades-mistral-math-DPO-7B-private
Dataset Card for Evaluation run of nbeerbower/bophades-mistral-math-DPO-7B
Dataset automatically created during the evaluation run of model nbeerbower/bophades-mistral-math-DPO-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-mistral-math-DPO-7B-private.lm-eval-results-chihoonlee10-T3Q-DPO-Mistral-7B-private
Dataset Card for Evaluation run of chihoonlee10/T3Q-DPO-Mistral-7B
Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-DPO-Mistral-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-DPO-Mistral-7B-private.lm-eval-results-chlee10-T3Q-Mistral-Orca-Math-dpo-v2.0-private
Dataset Card for Evaluation run of chlee10/T3Q-Mistral-Orca-Math-dpo-v2.0
Dataset automatically created during the evaluation run of model chlee10/T3Q-Mistral-Orca-Math-dpo-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chlee10-T3Q-Mistral-Orca-Math-dpo-v2.0-private.lm-eval-results-pkarypis-mistral-lima-private
Dataset Card for Evaluation run of pkarypis/mistral-lima
Dataset automatically created during the evaluation run of model pkarypis/mistral-lima
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-pkarypis-mistral-lima-private.lm-eval-results-nbeerbower-bophades-v2-mistral-7B-private
Dataset Card for Evaluation run of nbeerbower/bophades-v2-mistral-7B
Dataset automatically created during the evaluation run of model nbeerbower/bophades-v2-mistral-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-v2-mistral-7B-private.mistralai__Mixtral-8x22B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x22B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x22B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x22B-v0.1-details.mistralai__Mistral-Large-Instruct-2411-details
Dataset Card for Evaluation run of mistralai/Mistral-Large-Instruct-2411
Dataset automatically created during the evaluation run of model mistralai/Mistral-Large-Instruct-2411
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Large-Instruct-2411-details.lm-eval-results-chihoonlee10-T3Q-EN-DPO-Mistral-7B-private
Dataset Card for Evaluation run of chihoonlee10/T3Q-EN-DPO-Mistral-7B
Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-EN-DPO-Mistral-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-EN-DPO-Mistral-7B-private.lm-eval-results-nbeerbower-slerp-bophades-truthy-math-mistral-7B-private
Dataset Card for Evaluation run of nbeerbower/slerp-bophades-truthy-math-mistral-7B
Dataset automatically created during the evaluation run of model nbeerbower/slerp-bophades-truthy-math-mistral-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-slerp-bophades-truthy-math-mistral-7B-private.mistralai__Mistral-7B-Instruct-v0.3-details
Dataset Card for Evaluation run of mistralai/Mistral-7B-Instruct-v0.3
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-Instruct-v0.3
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-Instruct-v0.3-details.NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mistral-7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mistral-7B-DPO
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details.lm-eval-results-unaidedelf87777-wizard-mistral-v0.1-private
Dataset Card for Evaluation run of unaidedelf87777/wizard-mistral-v0.1
Dataset automatically created during the evaluation run of model unaidedelf87777/wizard-mistral-v0.1
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-unaidedelf87777-wizard-mistral-v0.1-private.lm-eval-results-mistralai-Mistral-7B-v0.3-private
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.3
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.3
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-mistralai-Mistral-7B-v0.3-private.BAAI__Infinity-Instruct-3M-0625-Mistral-7B-details
Dataset Card for Evaluation run of BAAI/Infinity-Instruct-3M-0625-Mistral-7B
Dataset automatically created during the evaluation run of model BAAI/Infinity-Instruct-3M-0625-Mistral-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BAAI__Infinity-Instruct-3M-0625-Mistral-7B-details.mistralai__Mistral-Small-24B-Base-2501-details
Dataset Card for Evaluation run of mistralai/Mistral-Small-24B-Base-2501
Dataset automatically created during the evaluation run of model mistralai/Mistral-Small-24B-Base-2501
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Small-24B-Base-2501-details.NousResearch__DeepHermes-3-Mistral-24B-Preview-details
Dataset Card for Evaluation run of NousResearch/DeepHermes-3-Mistral-24B-Preview
Dataset automatically created during the evaluation run of model NousResearch/DeepHermes-3-Mistral-24B-Preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__DeepHermes-3-Mistral-24B-Preview-details.mistralai__Mistral-7B-v0.3-details
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.3
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.3
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-v0.3-details.BAAI__Infinity-Instruct-7M-Gen-mistral-7B-details
Dataset Card for Evaluation run of BAAI/Infinity-Instruct-7M-Gen-mistral-7B
Dataset automatically created during the evaluation run of model BAAI/Infinity-Instruct-7M-Gen-mistral-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BAAI__Infinity-Instruct-7M-Gen-mistral-7B-details.qualcomm-interactive-cooking-dataset-counterfactual-mistakes
Qualcomm Interactive Cooking Dataset: Ego Counterfactual Mistakes
Description
This synthetic dataset contains mistake-intervention annotations for interactive cooking guidance. Each row contains video segment with instruction/feedback text pairs and their timestamps.
Dataset Details
Files:
annotations.json
Release statistics:
Total rows: 25,087
Unique videos (dataset + video_id): 1,110
Rows by source dataset:
CaptainCook4D: 4,969
Ego4D: 13,847
Ego-Exo4D: 6… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-counterfactual-mistakes.PersonaChat-Qwen-original-Mistral-Small-4-119B-2603
Visual Memory Results: personachat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "mistralai/Mistral-Small-4-119B-2603",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-original-Mistral-Small-4-119B-2603",
"results_jsonl": "results/PersonaChat-Qwen-original-Mistral-Small-4-119B-2603.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-original-Mistral-Small-4-119B-2603.BAAI__Infinity-Instruct-3M-0613-Mistral-7B-details
Dataset Card for Evaluation run of BAAI/Infinity-Instruct-3M-0613-Mistral-7B
Dataset automatically created during the evaluation run of model BAAI/Infinity-Instruct-3M-0613-Mistral-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BAAI__Infinity-Instruct-3M-0613-Mistral-7B-details.BAAI__Infinity-Instruct-7M-0729-mistral-7B-details
Dataset Card for Evaluation run of BAAI/Infinity-Instruct-7M-0729-mistral-7B
Dataset automatically created during the evaluation run of model BAAI/Infinity-Instruct-7M-0729-mistral-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BAAI__Infinity-Instruct-7M-0729-mistral-7B-details.
