Darveht/zenvion-voice-detector-v0.4
ποΈ Zenvion Voice Detector v0.4
Multi-task voice analysis model β 8 tasks in a single forward pass. Detects speech activity, gender, emotion, language, age, acoustic noise type, conversational intent, and accent simultaneously from raw audio in real time.
π [Try the live demo β](https://huggingface.co/spaces/Darveht/zenvion-voice-detector-demo)
π Table of Contents
- Overview
- Tasks & Labels
- Performance Benchmarks
- Quick Start
- Installation
- Usage Examples
- API Reference
- Architecture
- Training Details
- Limitations & Bias
- Changelog
- Citation
- License
Overview
Zenvion Voice Detector v0.4 is a multi-task speech analysis system built on top of [facebook/wav2vec2-base-960h](https://huggingface.co/facebook/wav2vec2-base-960h). A single forward pass produces 8 independent classification outputs, making it efficient for audio pipelines.
Tasks & Labels
Full label indexes are in label_mapping.json.
Performance Benchmarks
β οΈ The metrics below are the numbers published by the model author. They are markedverified: falsein the model card and have not been independently verified. The checkpoint currently shipped contains a pretrained backbone with freshly initialised task heads β runevaluation.pyon your own data before relying on these figures.
Voice Activity Detection (VAD)
Emotion Recognition
Language Identification (50 languages)
Gender Classification
Quick Start
Using Hugging Face Inference API
import requests
API_URL = "https://api-inference.huggingface.co/models/Darveht/zenvion-voice-detector-v0.4"
headers = {"Authorization": "Bearer YOUR_HF_TOKEN"}
with open("audio.wav", "rb") as f:
data = f.read()
response = requests.post(API_URL, headers=headers, data=data)
print(response.json())Using transformers pipeline
from transformers import pipeline
pipe = pipeline(
"audio-classification",
model="Darveht/zenvion-voice-detector-v0.4",
trust_remote_code=True,
)
result = pipe("audio.wav")
print(result)Using the bundled ZenvionPipeline
from inference import ZenvionPipeline
pipe = ZenvionPipeline(
model_id="Darveht/zenvion-voice-detector-v0.4",
device="cpu", # or "cuda"
half_precision=False, # True for FP16 on CUDA
)
result = pipe("path/to/audio.wav")
print(result)
# {
# "vad": {"label": "SPEECH", "score": 0.982},
# "gender": {"label": "MALE", "score": 0.871},
# "emotion": {"label": "NEUTRAL", "score": 0.763},
# "language": {"label": "en", "score": 0.941},
# "age": {"label": "ADULT", "score": 0.802},
# "noise": {"label": "CLEAN", "score": 0.913},
# "intent": {"label": "STATEMENT","score": 0.688},
# "accent": {"label": "AMERICAN", "score": 0.754},
# }Installation
pip install -r requirements.txtMinimum dependencies:
torch>=2.1.0
transformers>=4.37.0
torchaudio>=2.1.0
librosa>=0.10.0
soundfile>=0.12.1
numpy>=1.24.0Usage Examples
Batch Processing
from inference import ZenvionPipeline
from pathlib import Path
pipe = ZenvionPipeline()
audio_files = list(Path("audio_dir").glob("*.wav"))
for f in audio_files:
res = pipe(str(f))
print(f"{f.name}: {res['vad']['label']} | {res['emotion']['label']} | {res['language']['label']}")GPU Inference with FP16
from inference import ZenvionPipeline
pipe = ZenvionPipeline(device="cuda", half_precision=True)
result = pipe("audio.wav")Streaming / Real-time (chunk-based)
import numpy as np
from inference import ZenvionPipeline
pipe = ZenvionPipeline()
SAMPLE_RATE = 16000
CHUNK_S = 2 # 2-second windows
# numpy/tensor inputs are expected at 16 kHz mono
chunk = np.random.randn(SAMPLE_RATE * CHUNK_S).astype(np.float32)
result = pipe(chunk)
print(result)Run only specific tasks
from inference import ZenvionPipeline
pipe = ZenvionPipeline(tasks=["vad", "emotion", "language"])
result = pipe("audio.wav")Using with soundfile
import soundfile as sf
import numpy as np
from inference import ZenvionPipeline
pipe = ZenvionPipeline()
audio, sr = sf.read("audio.wav")
if audio.ndim > 1:
audio = audio.mean(axis=1)
# NOTE: array inputs are expected at 16 kHz mono; resample first if needed,
# e.g. with librosa.resample(audio, orig_sr=sr, target_sr=16000)
result = pipe(audio.astype(np.float32))API Reference
ZenvionPipeline
ZenvionPipeline(
model_id: str = "Darveht/zenvion-voice-detector-v0.4",
device: Optional[str] = None, # auto-detects cuda/cpu
tasks: Optional[List[str]] = None, # subset of the 8 tasks; None = all
half_precision: bool = False, # FP16 (CUDA only)
allow_random_init: bool = False, # True: run with random weights if the checkpoint is missing
)__call__(audio, return_all_scores=False)
Returns: dict[str, dict] β one entry per task with label (str) and score (float 0β1).
ZenvionConfig
from config_class import ZenvionConfig
cfg = ZenvionConfig.from_pretrained("Darveht/zenvion-voice-detector-v0.4")
print(cfg.num_labels_emotion) # 8
print(cfg.num_labels_language) # 50
print(cfg.num_labels) # 109 (total across all tasks)Architecture
Input: raw audio (16 kHz mono)
|
v
[Wav2Vec2 Encoder] β 12 transformer layers, 768-dim hidden
|
v (learned weighted sum of all 13 hidden states, softmax-normalised)
[AttentivePooling] β attention-weighted temporal pooling -> 768-dim
|
+---+--------------------------------------------+
v v v v v v v v
VAD Gender Emotion Lang Age Noise Int Acc
2 3 8 50 6 5 15 20 <- output classes (109 total)- Total parameters: ~95 M (wav2vec2-base) + ~1.6 M (layer weights, pooling, 8 heads)
- Inference speed (CPU): ~180 ms per 2-second clip
- Inference speed (A100 GPU): ~12 ms per 2-second clip
Training Details
Data
Hyperparameters
Augmentations
- Speed perturbation (0.9Γ β 1.1Γ)
- Additive noise (SNR 5 β 30 dB)
- Room impulse response (RIR) convolution
- Codec simulation (telephone, mp3)
- Random gain (β6 to +6 dB)
- Time masking (SpecAugment-style)
Limitations & Bias
- Accent detection is English-centric; accuracy drops significantly for non-English accents.
- Emotion models trained on acted speech (RAVDESS, CREMA-D) may underperform on spontaneous conversational emotion.
- Age estimation is coarse (6 buckets) and may be biased toward training demographics.
- Gender outputs only MALE / FEMALE / UNKNOWN; does not capture the full spectrum of gender expression.
- Language ID accuracy varies: high for European languages (>95%), lower for low-resource languages such as Tagalog or Malay (~85%).
- Min duration: clips shorter than 0.5 s may produce unreliable outputs.
- Music / non-speech: the model may output unpredictable emotion or language labels on pure music β check VAD output first.
Changelog
v0.4.1 β 2026-09-09 (code & config fixes)
- Critical β weight wipe fixed:
post_init()βinit_weights()was silently re-initialising the pretrained wav2vec2 backbone on every model construction, destroying the pretrained weights. The backbone is now protected via an_init_weights()override (transformers 4.x) and_is_hf_initializedmarking (transformers 5.x); verified byte-identical againstfacebook/wav2vec2-base-960h. - First real checkpoint: the repo previously shipped no weights at all (
model.safetensors/pytorch_model.binwere missing), sofrom_pretrained()could not load a working model. Addedmodel.safetensorswith the pretrained backbone + freshly initialised task heads (heads still need training β seetrain.py). - Crash fix:
forward()passedlogits=toZenvionOutput, which had no such field (TypeError). Field added;logitsis the [B, 109] concat. - Label maps fixed:
config.jsonhad three different classes all namedOTHER(noise/intent/accent), collapsinglabel2idfrom 109 to 107 entries. Renamed toNOISE_OTHER/INTENT_OTHER/ACCENT_OTHER; 109/109 consistent.ZenvionConfigno longer overwritesnum_labels=109with 2, and normalisesid2labelkeys to ints. - NaN guards:
AttentivePoolingno longer returns NaN on fully-masked rows; multi-task loss skips tasks whose labels are all-100(was NaN). - `predict()` now decodes all 8 tasks (was 5) and honours
threshold. - `inference.py`: removed the silent fallback to random weights (now raises unless
allow_random_init=True); validates task names, empty audio, shapes, and min/max duration; the final partial VAD window is analysed instead of dropped. - `train.py`:
--amp/--no-ampflags; scheduler is rebuilt when the optimizer is recreated at unfreeze (was bound to the discarded optimizer);global_stepnow increments; leftover gradient-accumulation steps are applied at epoch end; AMP is CUDA-only; gradient loss is scaled by the accumulation factor; resuming a post-unfreeze checkpoint unfreezes first so optimizer parameter groups line up, and training continues at the next epoch. - `masked_spec_embed` NaN fix (transformersβ₯5): the 960h checkpoint does not contain this pre-training-only parameter, and
from_pretrainedmaterialised it as uninitialised memory (NaN). It is now explicitly uniform-initialised like wav2vec2 pre-training does; the backbone stays byte-identical otherwise. - `num_labels` on transformersβ₯5: it is a read-only property derived from
len(id2label)β assigning it invoked the property setter and regeneratedid2labelas genericLABEL_Xentries. The assignment was removed; 109 is derived from the real label map.predict()now derives every task's label names fromconfig.id2label(the hardcodednoiselist still saidOTHER). - `evaluation.py`: EER is now the threshold minimising |FAR β FRR| (was the misleading min of (FAR+FRR)/2); AUC is reported raw (was
max(auc, 1-auc)). - `dataset_loader.py`: Speech Commands
_silence_/_background_noise_correctly map toNO_SPEECH(incl. int ClassLabel indices); fixedmd5("")filename collisions in Common Voice and therandint(0, 1e9)float crash in FLEURS. - Docs: corrected class counts (noise 5, intent 15, accent 20),
ZenvionPipelinesignatures/examples, and the architecture description. Benchmark figures remain the author's unverified numbers (verified: false).
v0.4 β 2025-07-24
- Added
preprocessor_config.jsonβ required for transformerspipeline()to work out of the box - Added
tokenizer_config.jsonandspecial_tokens_map.jsonfor AutoTokenizer compatibility - Added
vocab.jsonfor tokenizer - Added
label_mapping.jsonwith all 8 task labels, thresholds, and metadata - Added
CITATION.cfffor academic references - Full README rewrite: benchmarks table, architecture diagram, training details, API reference, limitations
- Added live Gradio demo Space: Darveht/zenvion-voice-detector-demo
- Improved noise robustness: +2.1% VAD accuracy on telephone-codec audio
- Fixed language head misclassification of Malayalam (ml) as Hindi
v0.3 β 2025-06-10
- First public release
- 8-head multi-task model: VAD, gender, emotion, language (50), age, noise, intent, accent
- Base: facebook/wav2vec2-base-960h
Citation
@misc{darveht2025zenvion,
author = {Darveht},
title = {Zenvion Voice Detector v0.4: Multi-task Speech Analysis},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Darveht/zenvion-voice-detector-v0.4}},
note = {Apache 2.0 License}
}License
Released under the Apache 2.0 License. Base model (facebook/wav2vec2-base-960h) is also Apache 2.0.
[Live Demo](https://huggingface.co/spaces/Darveht/zenvion-voice-detector-demo) Β· [Report an Issue](https://huggingface.co/Darveht/zenvion-voice-detector-v0.4/discussions)
