model-j
Datasets
All datasets matching “model-j”model_judgment
Dataset Card for "livebench/model_judgment"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/model_judgment.Model-J
Model-J Dataset
This dataset contains the hyperparameters, metadata, and Hugging Face links for all models in the Model-J dataset, introduced in:
Learning on Model Weights using Tree Experts (CVPR 2025) by Eliahu Horwitz*, Bar Cavia*, Jonathan Kahana*, Yedid Hoshen
🌐 Project | 📃 Paper | 💻 GitHub | 🤗 Models
Overview
Model-J is a large-scale dataset of trained neural networks designed for research on learning from model weights. It contains 14,004 models… See the full description on the dataset page: https://huggingface.co/datasets/ProbeX/Model-J.Self-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1Self-J-score-w-ref-skywork-pref-ref-lla31-70b-inst-model-lla-31-8b-inst-thre-0Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst
Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst"
More Information needed
Self-J-score-wo-ref-skywork-pref-model-yi-1.5-16k-chat-thre-1-10000
