ACE-Step/acestep-captioner
731.8k
<a href="https://arxiv.org/abs/2602.00744">Tech Report</a>
ACE-Step Captioner
Description
ACE-Step Captioner is the annotation model used by ACE-Step v1.5 for training data labeling. It is a professional-grade music captioning model that generates detailed, structured descriptions of audio content.
Performance
๐ Accuracy surpasses Gemini Pro 2.5 in music description tasks
Key Features
- ๐ผ Musical Style Analysis - Identifies genres, sub-genres, and stylistic influences
- ๐ธ Instrument Recognition - Detects and describes 1000+ instrument types and combinations
- ๐ญ Structure & Progression - Analyzes musical arrangement including intro, verse, chorus, bridge, climax, and outro
- ๐ Timbre Description - Captures tonal qualities, textures, and sonic characteristics
- ๐ Rich Vocabulary - Supports 1000+ descriptive terms for comprehensive music annotation
Usage
The usage is the same as Qwen2.5 Omni-7B.
Prompt Format
Use the following prompt to caption audio:
*Task* Describe this audio in detail
<audio>Output Format
The model generates natural language descriptions covering multiple aspects of the music.
Example Output
A melancholic indie folk track featuring fingerpicked acoustic guitar
as the primary instrument. The song opens with a sparse, contemplative
intro before the vocals enter with a breathy, intimate delivery.
The arrangement gradually builds through the verse, adding subtle
string pads and a gentle kick drum. The chorus lifts with layered
harmonies and a warmer, fuller texture. The bridge introduces a
key change and emotional climax before returning to the stripped-down
acoustic arrangement for the outro.Descriptive Capabilities
Musical Styles (Examples)
Instruments (1000+ Supported)
Structure Analysis
- Intro / Outro - Opening and closing sections
- Verse / Pre-Chorus / Chorus - Main song structure
- Bridge / Break - Transitional sections
- Build-up / Drop / Climax - Dynamic progression
- Interlude / Solo - Instrumental passages
Timbre Descriptions
Use Cases
- Music AI Training - Generate high-quality captions for music generation models
- Music Information Retrieval - Create searchable metadata for audio databases
- Content Moderation - Analyze and categorize music content
- Music Education - Provide detailed analysis for learning purposes
- Audio Production - Document and describe sound design elements
