CoolFace
Modelpublic

Nathan9/xcodec_mini_infer

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes
README.md127 linesDownload Raw Back to root
1---2language:3- en4tags:5- audio6- music7- codec8- neural-audio9- audio-compression10- transformers11pipeline_tag: audio-to-audio12library_name: transformers13inference: true14---15 16 17# XCodec Mini - Neural Audio Codec18 19## Model Description20 21XCodec Mini is a state-of-the-art neural audio codec designed for high-quality music compression and reconstruction. It combines semantic and acoustic encoding approaches to achieve efficient compression while maintaining audio quality.22 23### Key Features24 25- **Dual Encoding Architecture**26  - Semantic encoder for high-level musical features27  - Acoustic encoder for detailed sound information28  - Multi-scale processing for efficient compression29 30- **Advanced Compression**31  - Multiple codebooks for flexible quality/size tradeoff32  - Support for 44.1kHz high-fidelity audio33  - Separate processing paths for vocals and instrumentals34 35- **Technical Specifications**36  - Input: Raw audio at 44.1kHz37  - Output: Compressed representations and reconstructed audio38  - Model Size: [Add total size]39  - Compression Ratio: [Add typical ratio]40 41## Intended Uses42 43- High-quality music compression44- Audio archival and storage45- Music streaming applications46- Audio processing pipelines47 48## Training Data49 50The model was trained on a diverse dataset of music, including:51- Various genres and styles52- Vocal and instrumental tracks53- High-quality studio recordings54 55## Performance and Limitations56 57### Strengths58- High-quality audio reconstruction59- Efficient compression ratios60- Separate handling of vocals and instrumentals61- Support for high sample rates62 63### Limitations64- Computationally intensive for real-time applications65- Requires significant GPU memory66- Best suited for offline processing67- May introduce artifacts in extreme compression settings68 69## Technical Specifications70 71### Model Architecture721. **Semantic Encoder**73   - Based on HuBERT architecture74   - Captures high-level musical features75   - Outputs semantic tokens76 772. **Acoustic Encoder**78   - Multi-scale convolutional architecture79   - Processes detailed sound information80   - Generates acoustic tokens81 823. **Dual Decoders**83   - Separate decoders for vocals and instrumentals84   - Multi-stage reconstruction process85   - Quality-focused design86 87### Input Requirements88- Audio Format: WAV/MP389- Sample Rate: 44.1kHz90- Channels: Mono/Stereo91- Bit Depth: 16-bit92 93### Output Format94- Reconstructed Audio: 44.1kHz WAV95- Intermediate Representations: Compressed tokens96 97## Usage Guidelines98 99### Hardware Requirements100- GPU: NVIDIA GPU with 8GB+ VRAM101- RAM: 16GB+ recommended102- Storage: SSD recommended for faster processing103 104### Software Requirements105- Python 3.8+106- PyTorch 2.0+107- CUDA 11.0+108- Additional dependencies listed in installation guide109 110## Ethical Considerations111 112- **Copyright**: Users should ensure they have proper rights to process copyrighted material113- **Attribution**: Proper attribution should be given when using this model114- **Data Privacy**: Consider data privacy implications when processing sensitive audio115 116 117## Additional Information118 119### Model Weights120The model requires several checkpoint files:121- Semantic Encoder: `semantic_ckpts/hf_1_325000/pytorch_model.bin`122- Vocal Decoder: `decoders/decoder_131000.pth`123- Instrumental Decoder: `decoders/decoder_151000.pth`124- Final Checkpoint: `final_ckpt/ckpt_00360000.pth`125 126### Contact127For issues and questions, please use the GitHub repository's issue tracker.