datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_psmathur__model_42_70b
Dataset Card for Evaluation run of psmathur/model_42_70b
Dataset Summary
Dataset automatically created during the evaluation run of model psmathur/model_42_70b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_psmathur__model_42_70b.details_ChuckMcSneed__ArcaneEntanglement-model64-70b
Dataset Card for Evaluation run of ChuckMcSneed/ArcaneEntanglement-model64-70b
Dataset automatically created during the evaluation run of model ChuckMcSneed/ArcaneEntanglement-model64-70b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ChuckMcSneed__ArcaneEntanglement-model64-70b.Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-thre-0Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst
Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst"
More Information needed
Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-temp-v1-v2-thre-1Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-mis-7b-v0.2-thre-1-label-10000Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-lla31-8b-qwen2-7b-inst-0.5Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-qwen2-7b-inst-thre-1-10000Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-qwen2-7b-inst-thre-1Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla31-8b-inst-thre-1-1000Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-yi-1.5-16k-chat-thre-1dsl_icl_eval-2025_01_30_204127_model-deepseek-deepseek-r1-distill-llama-70b_fewshot-5Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-iter1-thre-2Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla31-8b-inst-thre-1-8000Self-J-score-w-ref-ref-lla31-70b-inst-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1dsl_icl_eval-2025_01_28_032427_model-deepseek-deepseek-r1-distill-llama-70b_fewshot-25Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-iter1-thre-1Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-mis-7b-v0.2-thre-1Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-yi-1.5-16k-chat-thre-1-10000dsl_icl_eval-2025_01_26_212421_model-deepseek-deepseek-r1-distill-llama-70b_fewshot-5dsl_icl_eval-2025_01_28_091927_model-deepseek-deepseek-r1-distill-llama-70b_fewshot-5dsl_icl_eval-2025_01_31_012038_model-deepseek-deepseek-r1-distill-llama-70b_fewshot-25Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla31-8b-inst-thre-1-10000Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla31-8b-inst-thre-1-2000dsl_icl_eval-2025_01_27_084453_model-deepseek-deepseek-r1-distill-llama-70b_fewshot-25Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-temp-v1-v2-thre-0.5mem_agent-model_based-llama-3-3-70b-i-infbench-longbook-qa-test-c27000-t4096-10s-agnosticSelf-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-thre-1Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-70b-inst-thre-1Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-temp-v1-v2-thre-0
