Tachyeon/audio-fingerprint-indian-bench
Audio Fingerprinting Benchmark on Indian Classical Music A reproducible, pre-registered benchmark of five audio-fingerprinting systems on the Saraga 1.5 corpus (Hindustani + Carnatic), plus a pre-registered training-recipe improvement to the NAFP baseline that achieves Bonferroni-significant gains on 1-second queries. v0.7 · closed-world retrieval · 5 systems · 6 528 evaluation cells · pooled McNemar p = 3.18 × 10⁻⁶… See the full description on the dataset page: https://huggingface.co/datasets/Tachyeon/audio-fingerprint-indian-bench.
Remove tabular configs from YAML; fetch via hf_hub_download instead
Add explicit dataset_info schema to override audio auto-detection
Fix dataset preview: rename split 'library' → 'train' for 5 tabular configs
Trim changelog to current version only
v0.7: production cleanup
README: lead with error-rate reduction (~68% across benchmark)
v0.6: 5 systems × 4 lengths × main+ablation + recipe v3 (3 seeds, Bonferroni-significant on 1s queries) + 2 pre-registered negative results
v0.5: post-mortem audit fixes — twin flag MBID-only, Dejavu ref_stop, score.py no_match denominator, NAFP stable tie-break
v0.4: add NAFP (4th system); new metrics (MRR, top1_near); Panako retune; refreshed baseline table
v0.3: refreshed ablation results against 632-query manifest + length-degradation table for main queries
v0.2: expanded ablation coverage (NFKD-normalised section bucketing); ablation 624→632 (H 200→207, C 424→425); refreshed README + language metadata
README: switch queries / queries_ablation configs to AudioFolder glob so file_name auto-resolves to Audio()
v0.1: 1624 query clips (1000 main + 624 ablation) + refs metadata + 3-system baseline results (Olaf, Dejavu, Panako) + inspection tables
initial commit
