CoolFace
Modelpublic

Rubin-Wei/MemoryDecoder-Qwen3-1.7B-finance

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes16downloads
Model Card

MemoryDecoder-Qwen3-1.7B-finance

Resources

This repository contains the 1.7B finance Memory Decoder released with Memory Decoder at Scale. It is a pretrained parametric long-term memory that can be swapped into a compatible frozen language-model backbone.

This 1.7B memory was trained on the finance CPT corpus using retrieval-derived sparse target distributions. The released dataset contains the aligned CPT text, Qwen3-preprocessed data, and KNN distributions. In the paper, the 1.7B Qwen3-vocabulary domain memories are evaluated with frozen Qwen3 Base backbones from 0.6B to 14B.

This checkpoint is a memory component, not a standalone chat- or instruction-tuned model. The memory and backbone must use compatible token IDs and vocabularies.

Model details

FieldValue
Memory size1.7B class
Architecture/tokenizer familyQwen3ForCausalLM memory; Qwen3 tokenizer/vocabulary
DomainFinance
Evaluation benchmarkFinEval
Intended backboneFrozen, tokenizer-compatible Qwen3 Base model
Release contentsInference weights, configuration, and tokenizer files

Usage

Install the matching environment from LUMIA-Group/MemoryDecoder-at-Scale, then set MODEL_PATH to a compatible frozen backbone and MEMDEC_PATH to this repository:

bash
MODEL_PATH=/path/to/compatible-base-model \
MEMDEC_PATH=Rubin-Wei/MemoryDecoder-Qwen3-1.7B-finance \
bash eval/lm-evaluation-harness/scripts/domain/evaluate_fineval.sh

See the repository documentation and launcher for benchmark-specific options, including interpolation weights and batch settings.

Intended use and limitations

This checkpoint is intended for research and evaluation in the finance domain. Its outputs depend on the backbone, prompt, and interpolation settings. Domain specialization does not guarantee factual correctness or safety, and the model may inherit biases and errors from its training sources.

Citation

If you use this checkpoint, please cite:

bibtex
@misc{wei2026memorydecoderscalepretrained,
      title={Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory},
      author={Rubin Wei and Jiaqi Cao and Jiarui Wang and Junming Zhang and Qipeng Guo and Bowen Zhou and Zhouhan Lin},
      year={2026},
      eprint={2607.27919},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2607.27919},
}