CoolFace
Modelpublic

okestro-ai-lab/FastSLM

sourceHugging Faceapache-2.0updated 23d agoView on Hugging Face
1likes91downloads
Model Card
  • โ€”FastSLM is an designed for efficient and accurate speech-to-text transcription.

Paper: FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Code: GitHub - Lee-junseok1025/FastSLM

<!-- ๐Ÿ”Š HFQ-Former: Hierarchically compresses high-frame-rate audio features while preserving the audio's local and global contextual information. --> <!-- ๐Ÿ”Š Adaptor: --> <!-- ๐Ÿง  LLM Adaptation*: Effectively adapts pre-trained Large Language Models (LLMs) to the audio modality. -->

๐Ÿ“– Model Architecture

๐Ÿš€ FastSLM is a designed for efficient speech-to-text transcription.

๐ŸŽ‰ Accepted at EMNLP 2026 Findings

๐Ÿ“Œ Key Features (Safe Version)

  • โ€”โšก Efficient long-form speech processing (Processes up to 8 hours of audio on a 40GB GPU)
  • โ€”๐Ÿง  Adaptation of pre-trained LLMs to audio
  • โ€”โœ… Evaluation results on standard ASR benchmarks (WERs listed above)
  • โ€”๐Ÿ“ Text Representation Preservation: Maintains the original LLM's capabilities for text-only tasks via the disable LoRA feature.

<p align="center"> <img src="HTA.png" width="1024" alt="HTA architecture"> </p>

๐Ÿš€ Getting Started

1. Installation

First, install the required libraries.

bash
sudo apt install ffmpeg
# pip
torch==2.3.1
peft==0.14.0
librosa==0.11.0
transformers>=4.53.1
accelerate==0.34.2
einops==0.8.1
torchaudio==2.3.1
openai-whisper
soundfile

2. Load Model and Tokenizer

You can easily load the model using AutoModelForCausalLM.from_pretrained. This model includes custom code, so the trust_remote_code=True option is required.

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig

# โฌ…๏ธ Enter your Hugging Face repository ID here.
repo_id = "okestro-ai-lab/FastSLM" 

model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(repo_id)
generation_config = GenerationConfig.from_pretrained(repo_id)

model.eval()

3. Sample Inference (Automatic Speech Recognition)

This example shows how to load an audio file and transcribe it to text.

python
import torch
import librosa

# 1. Load and resample the audio file
# โฌ…๏ธ Path to the audio file to be transcribed
wav_path = "sample_audio/English_audio.wav" 
wav,sample_rate  = librosa.load(wav_path)

# FastSLM requires 16kHz audio.
if sample_rate != 16000:
    audio = librosa.resample(wav,orig_sr=sample_rate,target_sr=16000)
else:
    audio = wav

# 2. Prepare the prompt and tokenize the prompt
# Automatic Speech Recognition (ASR) task
# Addiational Tasks: please refer to Supported Tasks
# A task token is not required, but it is recommended for achieving a more appropriate task.
TASK_TOKEN = "<|ASR|>" 
AUDIO_TOKEN = "<|audio_bos|><|AUDIO|><|audio_eos|>"
user_prompt = f"{TASK_TOKEN}{AUDIO_TOKEN}
Transcribe the audio clip into text."

prompt = [{"role": "user", "content": user_prompt}]
input_ids = tokenizer.apply_chat_template(
    prompt,
    add_generation_prompt=True,
    tokenize=True,
    return_tensors='pt'
).to(model.device)

# 3. Perform inference
# The model's generate function expects the audio input as a list.
audio_tensor = torch.tensor((audio,),dtype=torch.float32).cuda()

with torch.no_grad():
    with torch.cuda.amp.autocast(dtype=torch.bfloat16):
        output_ids = model.generate(
            input_ids=input_ids,
            audio=audio_tensor,
            generation_config=generation_config,
            max_new_tokens=256
        )

# 5. Decode the result
transcription = tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0]

print("--- Transcription Result ---")
print(transcription)

๐Ÿ“Œ Supported Tasks

You can perform different tasks by using the following special tokens in your prompt:

  • โ€”<|ASR|>: Automatic Speech Recognition - Transcribes audio into text.
  • โ€”<|AST|>: Automatic Speech Translation - Translates audio into text of another language.
  • โ€”<|SSUM|>: Speech Summarization - Summarizes the content of an audio clip.
  • โ€”<|SQQA|>: Spoken Query-based Question Answering - Answers questions based on the content of an audio clip.

โšก GPU Requirements

FastSLM inference requires a GPU with sufficient memory.

TaskRecommended GPUMinimum VRAM
InferenceNVIDIA A100 / H100โ‰ฅ 11.8 GB
๐Ÿ’ก Using mixed precision (bfloat16 or fp16) is recommended to reduce memory usage.

๐Ÿ“– Citation

bibtex
@inproceedings{lee2026fastslm,
  title     = {FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation},
  author    = {Lee, Junseok and Chun, Chang-Jae},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}