Nathan9/xcodec_mini_infer
0
1---2language:3- en4tags:5- audio6- music7- codec8- neural-audio9- audio-compression10- transformers11pipeline_tag: audio-to-audio12library_name: transformers13inference: true14---15 16 17# XCodec Mini - Neural Audio Codec18 19## Model Description20 21XCodec Mini is a state-of-the-art neural audio codec designed for high-quality music compression and reconstruction. It combines semantic and acoustic encoding approaches to achieve efficient compression while maintaining audio quality.22 23### Key Features24 25- **Dual Encoding Architecture**26 - Semantic encoder for high-level musical features27 - Acoustic encoder for detailed sound information28 - Multi-scale processing for efficient compression29 30- **Advanced Compression**31 - Multiple codebooks for flexible quality/size tradeoff32 - Support for 44.1kHz high-fidelity audio33 - Separate processing paths for vocals and instrumentals34 35- **Technical Specifications**36 - Input: Raw audio at 44.1kHz37 - Output: Compressed representations and reconstructed audio38 - Model Size: [Add total size]39 - Compression Ratio: [Add typical ratio]40 41## Intended Uses42 43- High-quality music compression44- Audio archival and storage45- Music streaming applications46- Audio processing pipelines47 48## Training Data49 50The model was trained on a diverse dataset of music, including:51- Various genres and styles52- Vocal and instrumental tracks53- High-quality studio recordings54 55## Performance and Limitations56 57### Strengths58- High-quality audio reconstruction59- Efficient compression ratios60- Separate handling of vocals and instrumentals61- Support for high sample rates62 63### Limitations64- Computationally intensive for real-time applications65- Requires significant GPU memory66- Best suited for offline processing67- May introduce artifacts in extreme compression settings68 69## Technical Specifications70 71### Model Architecture721. **Semantic Encoder**73 - Based on HuBERT architecture74 - Captures high-level musical features75 - Outputs semantic tokens76 772. **Acoustic Encoder**78 - Multi-scale convolutional architecture79 - Processes detailed sound information80 - Generates acoustic tokens81 823. **Dual Decoders**83 - Separate decoders for vocals and instrumentals84 - Multi-stage reconstruction process85 - Quality-focused design86 87### Input Requirements88- Audio Format: WAV/MP389- Sample Rate: 44.1kHz90- Channels: Mono/Stereo91- Bit Depth: 16-bit92 93### Output Format94- Reconstructed Audio: 44.1kHz WAV95- Intermediate Representations: Compressed tokens96 97## Usage Guidelines98 99### Hardware Requirements100- GPU: NVIDIA GPU with 8GB+ VRAM101- RAM: 16GB+ recommended102- Storage: SSD recommended for faster processing103 104### Software Requirements105- Python 3.8+106- PyTorch 2.0+107- CUDA 11.0+108- Additional dependencies listed in installation guide109 110## Ethical Considerations111 112- **Copyright**: Users should ensure they have proper rights to process copyrighted material113- **Attribution**: Proper attribution should be given when using this model114- **Data Privacy**: Consider data privacy implications when processing sensitive audio115 116 117## Additional Information118 119### Model Weights120The model requires several checkpoint files:121- Semantic Encoder: `semantic_ckpts/hf_1_325000/pytorch_model.bin`122- Vocal Decoder: `decoders/decoder_131000.pth`123- Instrumental Decoder: `decoders/decoder_151000.pth`124- Final Checkpoint: `final_ckpt/ckpt_00360000.pth`125 126### Contact127For issues and questions, please use the GitHub repository's issue tracker. 