datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.qwen3.5-4b-base-blindspot-samples
Blind Spots of Frontier Models: Qwen3.5-4B-Base
Fatima Fellowship 2026 -- Technical Challenge Report
GitHub: GrantorShadow/Fatima-Fellowship-2026
1. Model Selection
Model: Qwen/Qwen3.5-4B-Base
Property
Value
Parameters
4B
Architecture
Gated DeltaNet hybrid: 8x(3xDeltaNet + FFN + 1xAttention + FFN)
Context Window
262K tokens (claimed)
Type
Base model (no instruction tuning)
Modality
Multimodal (text + vision)
Release
Within last 6 months on Hugging… See the full description on the dataset page: https://huggingface.co/datasets/EtherealGlorious/qwen3.5-4b-base-blindspot-samples.
