CoolFace
20 results

basemodel

base-model-evals /off-the-shelf-model-evals Babel Tower benchmark archive Filesystem toolkit 1.3.0 provides versioned storage and validation for the proposed multilingual difficulty-calibrated benchmark. It connects a question bank, evaluation protocols, model-family panels, item-level responses, external calibration results, and frozen benchmark releases. Current data status: no real evaluation runs, complete foundation item bank, or fitted calibration parameters have been published here. The toolkit includes separate… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/off-the-shelf-model-evals.1 likes97 downloads1d agoHugging Facebakrihallak /BaseModelPretraintextn<1K0 likes51 downloads7mo agoHugging Facebase-model-evals /global-mmlu-rephrased global_mmlu (rephrased for base-model evaluation) Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.tabularmultiple-choicen<1K0 likes50 downloads8d agoHugging Facebase-model-evals /belebele-rephrased belebele (rephrased for base-model evaluation) Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.tabularmultiple-choicen<1K0 likes49 downloads8d agoHugging FaceKyleyee /eval_data_imdb_with_basemodel_truepreferences0 likes26 downloads2y agoHugging Facelibrarian-bots /base_model_sprint Base Model Metadata Sprint Description Join us in improving the discoverability and understanding of models on the Hugging Face Hub by adding base_model metadata! This sprint aims to enhance the information available for models derived from, fine-tuned on, or quantized versions of existing base models. 🤗 Strong contributions will win prizes!! 🤗 Why It Matters Adding base_model metadata helps users: Easily find models derived from specific architectures… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/base_model_sprint.text1K<n<10K5 likes21 downloads2y agoHugging Face