CoolFace
Modelpublic

uzabase/LLM2Vec-Llama-2-7b-hf-wikipedia-jp-mntp

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes8downloads
Model Card

Model Info

This is a model that applies LLM2Vec to Llama2. Only the PEFT Adapter is distributed. LLM2Vec fine-tunes on two tasks: MNTP and SimCSE, but this repository contains the results of applying only the MNTP task.

Model Details

Model Description

<!-- Provide a longer summary of what this model is. -->

  • —Model type: PEFT
  • —Language(s) (NLP): Japanese
  • —License: Apache2.0
  • —Finetuned from model: Llama-2-7b-hf

Sources

  • —Repository: https://github.com/McGill-NLP/llm2vec
  • —Paper: https://arxiv.org/abs/2404.05961

Usage

Training Details

Training Data

Training Hyperparameter
  • —batch_size: 64,
  • —gradientaccumulationsteps: 1
  • —maxseqlength": 512,
  • —masktokentype: "blank"
  • —mlm_probability: 0.2
  • —lora_r: 16
  • —torch_dtype "bfloat16"
  • —attnimplementation "flashattention_2"
  • —bf16: true
  • —gradient_checkpointing: true
Accelerator Settings
  • —deepspeed_config:
  • —gradientaccumulationsteps: 1
  • —gradient_clipping: 1.0
  • —offloadoptimizerdevice: nvme
  • —offloadoptimizernvme_path: /nvme
  • —zero3save16bit_model: true
  • —zero_stage: 2
  • —distributed_type: DEEPSPEED
  • —downcast_bf16: 'no'
  • —dynamo_config:
  • —dynamo_backend: INDUCTOR
  • —dynamo_mode: default
  • —dynamousedynamic: true
  • —dynamousefullgraph: true
  • —enablecpuaffinity: false
  • —machine_rank: 0
  • —maintrainingfunction: main
  • —mixed_precision: bf16
  • —num_machines: 1
  • —num_processes: 2
  • —rdzv_backend: static
  • —same_network: true
  • —quse_cpu: false

Framework versions

  • —Python: 3.12.3
  • —PEFT 0.11.1
  • —Sentence Transformers: 3.0.1
  • —Transformers: 4.41.0
  • —PyTorch: 2.3.0
  • —Accelerate: 0.30.1
  • —Datasets: 2.20.0
  • —Tokenizers: 0.19.1
  • —MTEB: 1.13.0