MD2204/multi_modality
1
๐ค Multimodal RAG Assistant
A powerful RAG system capable of "seeing" and analyzing images, charts, and tables within PDF documents using Gemini 2.0 Flash for vision and GPT-OSS 20B for chat.
๐ Features
- Visual Transcription: Automatically extracts and describes images using Gemini Vision.
- Smart Caching: Reuses previous transcriptions to save API costs.
- Precision Retrieval: Optimized chunking and K-retrieval for data-heavy documents.
- Modern UI: Clean Gradio interface for chatting and document ingestion.
