CoolFace
8 results

frontier-models

Pidoxy /Blind_Spots_of_Frontier_Models Qwen3-0.6B-Base — Blind Spots Dataset Model Tested Qwen/Qwen3-0.6B-Base Type: Causal Language Model (base / pretraining only — not instruction-tuned) Parameters: 0.6B (0.44B non-embedding) Released: April–May 2025 by Alibaba Cloud's Qwen Team Context Length: 32,768 tokens How the Model Was Loaded The model was loaded in a Google Colab T4 GPU notebook using HuggingFace transformers >= 4.51.0(required because the qwen3 architecture key was added in… See the full description on the dataset page: https://huggingface.co/datasets/Pidoxy/Blind_Spots_of_Frontier_Models.texttext-generationn<1K0 likes21 downloads7mo agoHugging Facechayma-rhaiem /Blind-Spots-of-Frontier-Models_Multihop-blindspots-qwen35 Qwen3.5-2B-Base — Multi-Hop Reasoning Blind Spots An evaluation dataset probing 18 Knowledge Graph-style reasoning tasks on Qwen/Qwen3.5-2B-Base, tested in its raw base (pre-training) form with no external graph attached. The dataset covers parametric memory (probes 1–10, no passage provided), standard grounded reasoning (probes 11–15, source passage included), and advanced grounded reasoning (probes 16–18, passage provided but requiring implicit inference or contradiction… See the full description on the dataset page: https://huggingface.co/datasets/chayma-rhaiem/Blind-Spots-of-Frontier-Models_Multihop-blindspots-qwen35.tabularquestion-answeringn<1K1 likes11 downloads7mo agoHugging FaceIzzi1 /blind-spots-for-frontier-models Qwen Model Tasks For Fellowship Note: view 'final_results_hf.json' for model results. For some reason hf is only showing 'test.json' in preview which doesnt contain models response. Model Loading The model was loaded using the transformers library from Hugging Face. The following code snippet was used to load the Qwen model and tokenizer: from transformers import AutoModelForCausalLM, AutoTokenizer # Define the model name model_name = "Qwen/Qwen3-4B-Base" # Load the… See the full description on the dataset page: https://huggingface.co/datasets/Izzi1/blind-spots-for-frontier-models.textn<1K0 likes3 downloads7mo agoHugging FaceTomodovodoo /blindspots-frontier-models-granite-4-0-1b-base Blind Spots of Frontier Models (IBM Granite 4.0 1B Base) Model tested: ibm-granite/granite-4.0-1b-baseModel card: https://huggingface.co/ibm-granite/granite-4.0-1b-base For inference, I ran this model locally, though I also experimented with free models from OpenRouter. This dataset contains 10 evaluation rows with: input expected_output model_output notes is_correct I loaded the model with transformers and evaluated it using strict concise-answer prompts. from transformers… See the full description on the dataset page: https://huggingface.co/datasets/Tomodovodoo/blindspots-frontier-models-granite-4-0-1b-base.texttext-generationn<1K0 likes2 downloads7mo agoHugging Facemostafa21314 /Blind_Spots_of_Frontier_Models0 likes1 downloads7mo agoHugging Facecristhianchr /Blind-Spots-of-Frontier-Models-Dataset0 likes1 downloads6mo agoHugging Face