judge-model
DRPG_JudgeModel-GGUFjudge-xl-modelbengali_qa_model_AGGRO_banglabertarch-xfam-judgeclean-lpi-260903T0340-qwen-e20-s2-poison-modelarch-xfam-judgeclean-lpi-260903T0340-gemma-e20-s1-ctrl-modelarch-xfam-judgeclean-lpi-260903T0340-qwen-e10-s1-poison-modelarch-xfam-judgeclean-lpi-260903T0340-qwen-e20-s1-ctrl-modelarch-xfam-judgeclean-lpi-260903T0340-qwen-e10-s2-poison-model
BeaverTails-dedupprompt_model-gpt-4o_harmful_cat_judge_clustercat_cot-improvedBeaverTails-dedupprompt_model-gpt-4o_harmful_cat_judge_clustercatphysics-cross-model-judge
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/muller-digital/physics-cross-model-judge.multi-model-judge-comparison-question-viewmulti-model-judge-comparison-20250729-154625multi-model-judge-comparison
LLM-Benchmark-Model-vs-Judgejudge-model-spacerepro-demystifying-llm-as-a-judge-analytically-tractable-model-for-inference-time-scalingrepro-demystifying-llm-as-a-judge-analytically-tractable-model-for-inference-time-scalingrepro-demystifying-llm-as-a-judge-analytically-tractable-model-for-inference-time-scalrepro-demystifying-llm-as-a-judge-analytically-tractable-model-for-inference-time-scaling
