CoolFace
Modelpublic

khanhld/chunkformer-large-en-libri-960h

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
2likes24downloads
Model Card

ChunkFormer-Large-En-Libri-960h: Pretrained ChunkFormer-Large on 960 hours of LibriSpeech dataset

<style> img { display: inline; } </style> ![License: CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) ![GitHub](https://github.com/khanld/chunkformer) ![Paper](https://arxiv.org/abs/2502.14673) ![Model size](#description)


Table of contents

  1. 1.Model Description
  2. 2.Documentation and Implementation
  3. 3.Benchmark Results
  4. 4.Usage
  5. 5.Citation
  6. 6.Contact

<a name = "description" ></a>

Model Description

ChunkFormer-Large-En-Libri-960h is an English Automatic Speech Recognition (ASR) model based on the ChunkFormer architecture, introduced at ICASSP 2025. The model has been fine-tuned on 960 hours of LibriSpeech, a widely-used dataset for ASR research.


<a name = "implementation" ></a>

Documentation and Implementation

The [Documentation]() and Implementation of ChunkFormer are publicly available.


<a name = "benchmark" ></a>

Benchmark Results

We evaluate the models using Word Error Rate (WER). To ensure a fair comparison, all models are trained exclusively with the **WENET** framework.

STTModelTest-CleanTest-OtherAvg.
1ChunkFormer2.696.914.80
2Efficient Conformer2.716.954.83
3Conformer2.776.934.85
4Squeezeformer2.877.165.02

<a name = "usage" ></a>

Quick Usage

To use the ChunkFormer model for English Automatic Speech Recognition, follow these steps:

Option 1: Install from PyPI (Recommended)

bash
pip install chunkformer

Option 2: Install from source

bash
git clone https://github.com/khanld/chunkformer.git
cd chunkformer
pip install -e .

Python API Usage

python
from chunkformer import ChunkFormerModel

# Load the English model from Hugging Face
model = ChunkFormerModel.from_pretrained("khanhld/chunkformer-large-en-libri-960h")

# For single long-form audio transcription
transcription = model.endless_decode(
    audio_path="path/to/long_audio.wav",
    chunk_size=64,
    left_context_size=128,
    right_context_size=128,
    total_batch_duration=14400,  # in seconds
    return_timestamps=True
)
print(transcription)

# For batch processing of multiple audio files
audio_files = ["audio1.wav", "audio2.wav", "audio3.wav"]
transcriptions = model.batch_decode(
    audio_paths=audio_files,
    chunk_size=64,
    left_context_size=128,
    right_context_size=128,
    total_batch_duration=1800  # Total batch duration in seconds
)

for i, transcription in enumerate(transcriptions):
    print(f"Audio {i+1}: {transcription}")

Command Line Usage

After installation, you can use the command line interface:

bash
chunkformer-decode \
    --model_checkpoint khanhld/chunkformer-large-en-libri-960h \
    --long_form_audio path/to/audio.wav \
    --total_batch_duration 14400 \
    --chunk_size 64 \
    --left_context_size 128 \
    --right_context_size 128

Example Output:

[00:00:01.200] - [00:00:02.400]: this is a transcription example
[00:00:02.500] - [00:00:03.700]: testing the long-form audio

Advanced Usage can be found HERE


<a name = "citation" ></a>

Citation

If you use this work in your research, please cite:

bibtex
@INPROCEEDINGS{10888640,
  author={Le, Khanh and Ho, Tuan Vu and Tran, Dung and Chau, Duc Thanh},
  booktitle={ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, 
  title={ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription}, 
  year={2025},
  volume={},
  number={},
  pages={1-5},
  keywords={Scalability;Memory management;Graphics processing units;Signal processing;Performance gain;Hardware;Resource management;Speech processing;Standards;Context modeling;chunkformer;masked batch;long-form transcription},
  doi={10.1109/ICASSP49660.2025.10888640}}
}

<a name = "contact"></a>

Contact

  • khanhld218@gmail.com
  • ![GitHub](https://github.com/khanld)
  • ![LinkedIn](https://www.linkedin.com/in/khanhld257/)