CoolFace
Modelpublic

Pradheep1647/run_gqa_rope-meetingbank-bs8-e20-fp32-19

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes10downloads
Model Card

Run Gqa Rope

Custom PyTorch Transformer checkpoint trained on MeetingBank for meeting summarization research. This repository is part of the `transformer-lab` collection.

Model Details

FieldValue
RepositoryPradheep1647/run_gqa_rope-meetingbank-bs8-e20-fp32-19
Attentiongqa_rope
Datasetmeetingbank
Layers6
Hidden size512
Heads8
Batch size8
Epochs20
Precisionfp32
Checkpointmeeting_model19.pt

Architecture

[image]

Static architecture diagram generated from this run's config.json, including model width, depth, sequence dimensions, and attention-specific settings.

Training Loss

[image]

Raw curve data is available in `loss_curve.csv`.

Available Models

Files

FilePurpose
meeting_model19.ptPyTorch checkpoint containing model_state_dict, optimizer states, epoch, and global step.
config.jsonTraining and architecture config converted from the Hydra run config.
architecture.pngArchitecture diagram generated from the saved model config, with block shapes and dimensions.
tokenizer.jsonMeetingBank transcript tokenizer alias for source inputs.
transcript_tokenizer.jsonExplicit MeetingBank transcript tokenizer.
summary_tokenizer.jsonMeetingBank summary tokenizer for target text.
loss_curve.csvTensorBoard train/loss scalar export.
loss_curve.svgStatic training-loss plot generated from loss_curve.csv.

Usage

These checkpoints are from a custom PyTorch codebase, not a transformers.AutoModel checkpoint. Use the repo-native builder to instantiate the architecture, then load the checkpoint state dict.

python
from pathlib import Path

import torch
from huggingface_hub import hf_hub_download
from omegaconf import OmegaConf

import src  # registers components
from src.model.builder import build_transformer

repo_id = "Pradheep1647/run_gqa_rope-meetingbank-bs8-e20-fp32-19"

config_path = hf_hub_download(repo_id=repo_id, filename="config.json")
checkpoint_path = hf_hub_download(repo_id=repo_id, filename="meeting_model19.pt")

cfg = OmegaConf.load(config_path)
model = build_transformer(cfg)

state = torch.load(checkpoint_path, map_location="cpu")
model.load_state_dict(state["model_state_dict"])
model.eval()

print(f"Loaded {repo_id} from {Path(checkpoint_path).name}")

Notes

  • —This is a research checkpoint for comparing attention variants under the same MeetingBank setup.
  • —The config and tokenizers are included so future runs can reproduce the architecture and preprocessing assumptions.
  • —Use config.json as the source of truth for architecture parameters.