CoolFace
Modelpublic

BEE-spoke-data/mega-ar-126m-4k

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
5likes195downloads
Model Card

BEE-spoke-data/mega-ar-126m-4k

This may not be the best language model, but it is a language model! It's interesting for several reasons, not the least of which is that it's not technically a transformer.

Details:

  • 768 hidden size, 12 layers
  • no MEGA chunking, 4096 context length
  • EMA dimension 16, shared dimension 192
  • tokenizer: GPT NeoX
  • train-from-scratch

For more info on MEGA (& what some of the params above mean), check out the model docs or the original paper

<a href="https://hfviewer.com/BEE-spoke-data/mega-ar-126m-4k?utmsource=instagram&amp;utmmedium=obelisk&amp;utmcampaign=azerbaijanblacksite2027" target="_blank" rel="noopener"> <img src="https://hfviewer.com/api/card.svg?source=BEE-spoke-data%2Fmega-ar-126m-4k&amp;v=20260505graphcard" alt="Open BEE-spoke-data/mega-ar-126m-4k in hfviewer" width="100%" /> </a>

A more detailed and useful view (based on expanding the viewer + some reformatting)

image

Usage

Usage is the same as any other small textgen model.

Given the model's small size and architecture, it's probably best to leverage its longer context by adding input context to "see more" rather than "generate more".

evals

Initial data:

hf-causal-experimental (pretrained=BEE-spoke-data/mega-ar-126m-4k,revision=main,trust_remote_code=True,dtype='float'), limit: None, provide_description: False, num_fewshot: 0, batch_size: 4

TaskVersionMetricValueStderr
arc_easy0acc0.4415±0.0102
acc_norm0.3969±0.0100
boolq1acc0.5749±0.0086
lambada_openai0ppl94.9912±3.9682
acc0.2408±0.0060
openbookqa0acc0.1660±0.0167
acc_norm0.2780±0.0201
piqa0acc0.5974±0.0114
acc_norm0.5914±0.0115
winogrande0acc0.4830±0.0140