datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.full-math-private-n256-Qwen2.5-3B-Instruct-bonfull-math-private-n256-Phi-4-mini-instruct-bonfull-math-private-n256-Llama-3.2-3B-Instruct-bonpreprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonfull-math-private-Qwen2.5-3B-Instruct-bonlm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.stratified-solvable-1k-math-private-Qwen2.5-3B-Instruct-bonlm-eval-results-AurelPx-Pegasus-7b-slerp-private
Dataset Card for Evaluation run of AurelPx/Pegasus-7b-slerp
Dataset automatically created during the evaluation run of model AurelPx/Pegasus-7b-slerp
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AurelPx-Pegasus-7b-slerp-private.lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.full-math-private-Qwen3-4B-Instruct-2507-bondetails_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private
Dataset Card for Evaluation run of hosted_vllm//fsx/anton/deepseek-r1-checkpoint
Dataset automatically created during the evaluation run of model hosted_vllm//fsx/anton/deepseek-r1-checkpoint.
The dataset is composed of 15 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private.lm-eval-results-s3nh-Severusectum-7B-DPO-private
Dataset Card for Evaluation run of s3nh/Severusectum-7B-DPO
Dataset automatically created during the evaluation run of model s3nh/Severusectum-7B-DPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-s3nh-Severusectum-7B-DPO-private.crcis-quranic-eval-leaderboard-results_details_mistralai__Mistral-7B-v0.1_private
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1.
The dataset is composed of 6 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_mistralai__Mistral-7B-v0.1_private.lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private
Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff
Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.lm-eval-results-shyamieee-JARVIS-v2.0-private
Dataset Card for Evaluation run of shyamieee/JARVIS-v2.0
Dataset automatically created during the evaluation run of model shyamieee/JARVIS-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-JARVIS-v2.0-private.crcis-quranic-eval-leaderboard-results_details_PartAI__Dorna-Llama3-8B-Instruct_private
Dataset Card for Evaluation run of PartAI/Dorna-Llama3-8B-Instruct
Dataset automatically created during the evaluation run of model PartAI/Dorna-Llama3-8B-Instruct.
The dataset is composed of 8 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_PartAI__Dorna-Llama3-8B-Instruct_private.ZebraLogicBench-privatelm-eval-results-yleo-EmertonMonarch-7B-slerp-private
Dataset Card for Evaluation run of yleo/EmertonMonarch-7B-slerp
Dataset automatically created during the evaluation run of model yleo/EmertonMonarch-7B-slerp
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yleo-EmertonMonarch-7B-slerp-private.lm-eval-results-shadowml-WestBeagle-7B-private
Dataset Card for Evaluation run of shadowml/WestBeagle-7B
Dataset automatically created during the evaluation run of model shadowml/WestBeagle-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shadowml-WestBeagle-7B-private.lm-eval-results-bunnycore-SmartToxic-7B-private
Dataset Card for Evaluation run of bunnycore/SmartToxic-7B
Dataset automatically created during the evaluation run of model bunnycore/SmartToxic-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bunnycore-SmartToxic-7B-private.lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.lm-eval-results-automerger-Inex12Yamshadow-7B-private
Dataset Card for Evaluation run of automerger/Inex12Yamshadow-7B
Dataset automatically created during the evaluation run of model automerger/Inex12Yamshadow-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-automerger-Inex12Yamshadow-7B-private.lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.preprocessed-full-math-private-Llama-3.2-3B-Instruct-bonlm-eval-results-abideen-AlphaMonarch-daser-private
Dataset Card for Evaluation run of abideen/AlphaMonarch-daser
Dataset automatically created during the evaluation run of model abideen/AlphaMonarch-daser
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-abideen-AlphaMonarch-daser-private.lm-eval-results-yunconglong-DARE_TIES_13B-private
Dataset Card for Evaluation run of yunconglong/DARE_TIES_13B
Dataset automatically created during the evaluation run of model yunconglong/DARE_TIES_13B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-DARE_TIES_13B-private.lm-eval-results-allenai-llama-3-tulu-2-dpo-8b-private
Dataset Card for Evaluation run of allenai/llama-3-tulu-2-dpo-8b
Dataset automatically created during the evaluation run of model allenai/llama-3-tulu-2-dpo-8b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-allenai-llama-3-tulu-2-dpo-8b-private.lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-tulu-human-v2
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-tulu-human-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private.full-gsm8k-private-n256-Qwen2.5-3B-Instruct-bon
