CoolFace
17 results

model-j

livebench /model_judgment Dataset Card for "livebench/model_judgment" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/model_judgment.tabular10K<n<100K2 likes2k downloads1y agoHugging FaceProbeX /Model-J Model-J Dataset This dataset contains the hyperparameters, metadata, and Hugging Face links for all models in the Model-J dataset, introduced in: Learning on Model Weights using Tree Experts (CVPR 2025) by Eliahu Horwitz*, Bar Cavia*, Jonathan Kahana*, Yedid Hoshen 🌐 Project | 📃 Paper | 💻 GitHub | 🤗 Models Overview Model-J is a large-scale dataset of trained neural networks designed for research on learning from model weights. It contains 14,004 models… See the full description on the dataset page: https://huggingface.co/datasets/ProbeX/Model-J.tabular10K<n<100K0 likes108 downloads7mo agoHugging Faceoceanpty /Self-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1tabular10K<n<100K0 likes34 downloads2y agoHugging Faceoceanpty /Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-thre-0tabular10K<n<100K0 likes28 downloads2y agoHugging Faceoceanpty /Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst" More Information needed text10K<n<100K0 likes26 downloads2y agoHugging Faceoceanpty /Self-J-score-wo-ref-skywork-pref-model-yi-1.5-16k-chat-thre-1-10000tabular10K<n<100K0 likes26 downloads2y agoHugging Face