AGmind/strizh-ru-retriever
card: state throughput as ~4000, matching the article
card: unify -ngl with the published logs
card: fix distillation wording, align throughput with the article, mark rerank stage
docs: ctx-per-slot serving caveat (verified live), align exposure attribution with article
docs: align metric naming with article (filtered / opus pair / short-query req/s, EN approximate)
docs: layer selection + warm-start + contrastive (not distillation) wording
docs: correct data license (ru_stackoverflow CC BY-SA), soften cross-lingual negatives claim
docs: soften dogfood corpus claim (not fresh/leak-free; synthetic single-gold)
fix: also raise stale max_length 256->8192 (some clients read it, not model_max_length)
docs: correct eval (passage-exposure-clean RU 0.75, split RU-recall/EN-nDCG), drop every-axis/smallest-fastest overclaims, note CPU/RU-first
docs: rewrite model card (honest positioning, verified eval grid, retrieval design)
fix: raise model_max_length 256->8192 (was silently truncating; base supports 8192)
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
Upload tokenizer_config.json with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
initial commit
