01Yassine/AudioLLM-Deepfake-Detection
0
AudioLLM-Deepfake-Detection — Results Hub
Best-run checkpoints, evaluation CSVs/JSONs, and aggregated metrics for the DeepFense AudioLLM Deepfake Detection project.
Hub repo: 01Yassine/AudioLLM-Deepfake-Detection
Contents (~283 GB)
Each run folder includes: best_run_meta.json, per-dataset eval CSVs, metrics JSON (with EER), and checkpoints (lora_best/, checkpoint_best.pt, etc.).
Aggregated metrics (machine-readable)
Best overall model
OpenSmile-After / Lora-256 / unfrozen / Whisper / Qwen-0.5B / α=256
- Avg Macro F1: 94.42%
- Avg Accuracy: 95.11%
- Avg EER: 5.36%
- Path:
OpenSmile-After/Lora-256/unfrozen/whisper/Qwen-0.5B/
Results Summary (local documentation)
Unified table of best runs across all experiment families.
Metrics
Datasets
- asv19_test — ASVspoof 2019 LA eval
- itw — In-The-Wild
- la21 — ASVspoof 2021 LA eval
- mlaad_en — MLAAD English
Averages (avg_*) are computed over evaluated datasets for each run (typically 4/4).
Experiment Families
Best Overall Models
Highest average Macro F1
- OpenSmile-After / Lora-256 / unfrozen / whisper / Qwen-0.5B / α=256 — 94.42% macro F1, 95.11% accuracy, 5.36% EER
- Path:
results/OpenSmile-After/Lora-256/unfrozen/whisper/Qwen-0.5B
Lowest average EER
- OpenSmile-After / Lora-256 / unfrozen / whisper / Qwen-0.5B / α=256 — 5.36% EER, 94.42% macro F1, 95.11% accuracy
- Path:
results/OpenSmile-After/Lora-256/unfrozen/whisper/Qwen-0.5B
Files
Notes
- Some runs borrow missing eval splits (documented in JSON
notes/borrowed_or_approximate). - OpenSmile before stage: LoRA α=16 and Lora-128 only for Qwen-0.5B; NoLoRA all sizes.
- OpenSmile after stage: full LoRA α sweep (0.5B) + α=16 for 3B/7B.
- EER requires score columns in eval CSV; if metrics JSON missing, EER computed from CSV.
