Rohitto-ai/ev-multimodal-rag
EV Multimodal RAG System — TATA Motors Electric Vehicle Knowledge Base
    
BITS WILP — Multimodal RAG Bootcamp | Individual Assignment Student: Rohit Bhardwaj | Email: 2024tm05006@wilp.bits-pilani.ac.in
Problem Statement
Domain: Electric Vehicle Engineering at TATA Motors
As a member of the Electric Vehicle (EV) product engineering team at TATA Motors, I work daily with a large corpus of heterogeneous technical documents — battery datasheets, homologation reports, ARAI certification submissions, thermal management white papers, motor controller specifications, and charging infrastructure compliance records. These documents are dense, multimodal artefacts that combine structured data (cell chemistry tables, performance matrices, regulatory checklists), visual information (battery architecture schematics, range-versus-speed curves, charging time bar charts), and dense engineering prose that ties them together.
The Problem
Engineers and product managers at TATA Motors currently struggle to extract precise, cross-referenced insights from this document corpus for two core reasons:
First, document heterogeneity. A single question — "What is the maximum DC fast-charging power supported by the Nexon EV Max, and under what temperature conditions does the BMS throttle charging?" — may require simultaneously reading a table of charging port specifications on page 5, a paragraph about BMS thermal control strategy on page 12, and interpreting a performance chart on page 9. No keyword search tool can reason across these three modalities in a single pass. Traditional enterprise search returns pages; it does not synthesise answers.
Second, specialised domain terminology. Indian EV regulatory documents reference standards like AIS-038, AIS-156, and IS-17017 that generic language models have not been trained on in depth. Acronyms like SoH (State of Health), PMSM, CCS2, UN GTR 20, and NMC 622 carry very specific technical meaning. A general-purpose chatbot without grounded context from the actual documents frequently halluccinates specifications or confuses similar models.
Why This Problem Is Unique
Unlike a generic document Q&A system, the EV engineering domain presents a specific challenge set: tables contain battery electrochemical data with units, tolerance ranges, and standard references that must be read together; diagrams show BMS architecture, thermal loops, and power electronics topology that require engineering interpretation, not just caption reading; and performance charts (range vs. speed curves, charging time plots) encode quantitative insight that text alone cannot convey. A system that treats all content as plain text loses the structured meaning of a table row or the trend encoded in a chart. Additionally, homologation documents are updated quarterly as regulatory standards evolve, making a static fine-tuned model impractical — the knowledge must be updatable by re-ingesting new PDFs.
Why RAG Is the Right Approach
Fine-tuning a large language model on this corpus would require curating thousands of question-answer pairs from restricted internal documents, retraining every quarter when standards are updated, and accepting a model that cannot cite its sources — a hard compliance requirement in automotive engineering. Keyword search returns documents, not answers. RAG uniquely solves all three problems: it retrieves only the relevant chunks (text, table, or image description) at query time, grounds the LLM response in those specific chunks, provides verifiable citations (filename + page number), and updates the knowledge base simply by re-running POST /ingest with a new PDF — no model retraining required.
Expected Outcomes
A successful system should enable an engineer to ask: "Which Indian safety standards govern the Nexon EV battery pack, and what is the cell-level IP rating?", and receive a grounded answer that cites the exact table row from the homologation document and the page it appears on. It should answer queries that require combining a table (e.g., charging specs) with an image description (e.g., a charging time chart) — demonstrating true cross-modal retrieval. It should support the product manager asking "what is the 0–100 km/h time for the Nexon EV Max?" from a specification sheet without knowing which page it is on or what document it lives in.
Architecture Overview
System Architecture Diagram
flowchart TD
subgraph Ingestion Pipeline
A[PDF Upload\n POST /ingest] --> B[PDFParser\nPyMuPDF + pdfplumber]
B --> C[Text Chunks]
B --> D[Table Chunks\nMarkdown format]
B --> E[Raw Images]
E --> F[VisionModel\nGemini 1.5 Flash]
F --> G[Image Description\nChunks]
C & D & G --> H[Embedder\nsentence-transformers\nall-MiniLM-L6-v2]
H --> I[(ChromaDB\nPersistent Vector Store\nCosine Similarity)]
end
subgraph Query Pipeline
J[Question\n POST /query] --> K[Embedder\nembed question]
K --> L[ChromaDB\ntop-K retrieval\nacross all chunk types]
L --> M[Retrieved Chunks\ntext + table + image]
M --> N[LanguageModel\nGemini 1.5 Flash\nRAG Prompt]
N --> O[Grounded Answer\n+ Source References]
end
subgraph FastAPI Endpoints
P[GET /health]
Q[POST /ingest]
R[POST /query]
S[GET /documents]
T[DELETE /documents/filename]
U[GET /docs\nSwagger UI]
end
A -.-> Q
J -.-> RIngestion Flow
- Client uploads a PDF via
POST /ingest PDFParseruses PyMuPDF to extract text and raw images page-by-page- pdfplumber detects table regions and converts them to Markdown
- Each extracted image is sent to Gemini 1.5 Flash Vision for a text description
- All chunk types (text, table, image-description) are embedded via sentence-transformers
- Chunks are upserted into ChromaDB with metadata:
{type, source, page}
Query Flow
- Client POSTs a question to
/query - The question is embedded using the same sentence-transformer model
- ChromaDB performs cosine-similarity search returning top-K chunks across all modalities
- Chunks are assembled into a structured context string
- Gemini 1.5 Flash generates a grounded answer using a strict RAG prompt template
- The response includes the answer and a list of source references with page numbers
Technology Choices
Setup Instructions
Prerequisites
- Python 3.10 or higher
- A free Google Gemini API key from Google AI Studio
- ~2 GB free disk space (for model weights and ChromaDB)
Step 1: Clone the Repository
git clone https://github.com/YOUR_USERNAME/ev-multimodal-rag.git
cd ev-multimodal-ragStep 2: Create a Virtual Environment
python -m venv venv
# macOS / Linux
source venv/bin/activate
# Windows
venv\Scripts\activateStep 3: Install Dependencies
pip install -r requirements.txtNote: The first run downloads the all-MiniLM-L6-v2 sentence-transformer model (~90 MB) from HuggingFace. Subsequent runs use the local cache.Step 4: Configure Environment Variables
cp .env.example .envOpen .env and set your Gemini API key:
GEMINI_API_KEY=your_actual_api_key_hereAll other settings have sensible defaults and do not need to be changed for a basic setup.
Step 5: Start the Server
uvicorn main:app --host 0.0.0.0 --port 8000 --reloadThe server starts at http://localhost:8000. The Swagger UI is at http://localhost:8000/docs.
Step 6: Ingest the Sample Document
curl -X POST "http://localhost:8000/ingest" \
-F "file=@sample_documents/tata_nexon_ev_technical_report.pdf"Step 7: Run a Query
curl -X POST "http://localhost:8000/query" \
-H "Content-Type: application/json" \
-d '{"question": "What is the battery capacity and ARAI certified range of the Nexon EV Max?"}'API Documentation
GET /health
Returns system readiness, model information, and vector index statistics.
Sample Response:
{
"status": "healthy",
"gemini_model": "gemini-1.5-flash",
"embedding_model": "all-MiniLM-L6-v2",
"indexed_documents": 1,
"total_chunks": 47,
"indexed_filenames": ["tata_nexon_ev_technical_report.pdf"],
"uptime_seconds": 142.5
}POST /ingest
Upload a multimodal PDF to parse, summarise, embed, and index.
Request: multipart/form-data with field file (PDF only, max 50 MB)
Sample Response (201 Created):
{
"message": "Document 'tata_nexon_ev_technical_report.pdf' ingested successfully.",
"filename": "tata_nexon_ev_technical_report.pdf",
"chunks": {
"text": 28,
"table": 12,
"image": 7,
"total": 47
},
"processing_time_seconds": 14.3
}Error Responses:
400— Non-PDF file uploaded413— File exceeds 50 MB limit422— PDF has no extractable content500— Internal parsing or embedding error
POST /query
Ask a natural language question against all indexed documents.
Request Body:
{
"question": "What safety standards govern the Nexon EV battery pack?"
}Sample Response (200 OK):
{
"question": "What safety standards govern the Nexon EV battery pack?",
"answer": "The TATA Nexon EV battery pack is governed by several Indian and international standards including AIS-038 (Rev 1) for electric powertrain safety, AIS-156 (Phase 1 & 2) for battery pack safety, AIS-049 for high-voltage safety, and IS 17017 for charging systems. The pack has also achieved IP67 certification per IEC 60529 and passed UN GTR 20 for thermal runaway propagation prevention. [Source: tata_nexon_ev_technical_report.pdf, Page 6, Table 4]",
"sources": [
{
"filename": "tata_nexon_ev_technical_report.pdf",
"page": 6,
"chunk_type": "table",
"excerpt": "| AIS-038 (Rev 1) | Electric power train — safety requirements | Certified | ARAI |...",
"relevance_score": 0.9341
},
{
"filename": "tata_nexon_ev_technical_report.pdf",
"page": 6,
"chunk_type": "text",
"excerpt": "The Nexon EV platform has been validated against all mandatory Indian automotive safety standards applicable to Battery Electric Vehicles...",
"relevance_score": 0.8876
}
],
"retrieval_count": 5
}Error Responses:
400— Empty or too-short question404— No documents have been indexed yet502— Gemini API generation failure
GET /documents
List all currently indexed documents with chunk counts.
Sample Response:
{
"total_documents": 2,
"documents": [
{"filename": "tata_nexon_ev_technical_report.pdf", "chunk_count": 47},
{"filename": "nexon_ev_service_manual_2024.pdf", "chunk_count": 83}
]
}DELETE /documents/{filename}
Remove a document and all its chunks from the vector index.
Sample Response:
{
"message": "Document 'tata_nexon_ev_technical_report.pdf' removed from the index.",
"filename": "tata_nexon_ev_technical_report.pdf",
"chunks_deleted": 47
}Screenshots
Screenshots are captured after running the system end-to-end with the sample PDF. All screenshots are in the screenshots/ folder and embedded below.
1. Swagger UI — All Endpoints
2. POST /ingest — Successful Ingestion
3. POST /query — Text-Based Answer
4. POST /query — Table-Based Answer
5. POST /query — Image Summary Answer
6. GET /health — Index Status
Note: Screenshots are added after the system is running locally. To reproduce: start the server, ingest the sample PDF, and run the queries shown in Setup Instructions.
Limitations & Future Work
Current Limitations
PDF parsing quality: PyMuPDF and pdfplumber work well for digitally-created PDFs but may produce poor results on scanned documents without embedded text layers. OCR integration (e.g., Tesseract or Amazon Textract) would be needed for older TATA Motors homologation documents that exist only as scans.
Image filtering: The system uses minimum dimension thresholds to skip decorative images, but may occasionally include page logos or watermarks and send them to the VLM, adding latency. A more sophisticated image classifier (e.g., filtering out images with near-uniform colour histograms) would reduce unnecessary VLM calls.
Chunk boundaries: The character-based chunker with sentence-boundary heuristics can split at sub-optimal points for highly structured text (e.g., numbered lists in regulatory documents). A semantic chunker using embeddings to detect topic shifts would preserve coherence better.
Multilingual support: TATA Motors documents include Hindi text in some regulatory filings. The all-MiniLM-L6-v2 embedding model is English-centric and degrades on non-English input.
Gemini API rate limits: The free tier allows 15 requests per minute, which can become a bottleneck when ingesting PDFs with many images (each image = one VLM API call). A job queue with exponential backoff would make this more robust.
Future Work
- Semantic chunking using embedding-based topic segmentation for higher-quality chunk boundaries
- OCR pipeline using Tesseract or Docling's built-in OCR for scanned PDF support
- Re-ranking using a cross-encoder model (e.g.,
ms-marco-MiniLM-L-6-v2) to reorder retrieved chunks before LLM generation - Multi-document reasoning with explicit citation tracking per sentence in the generated answer
- Streaming responses via FastAPI's
StreamingResponsefor long answers to improve perceived latency - Authentication middleware with API key validation for production deployment
- Docker Compose setup for reproducible deployment across TATA Motors engineering workstations
Project Structure
ev-multimodal-rag/
├── README.md # This file
├── main.py # FastAPI app + lifespan management
├── config.py # Pydantic settings (all env vars)
├── requirements.txt # Pinned Python dependencies
├── .env.example # API key template (copy to .env)
├── .gitignore # Excludes .env, chroma_db/, cache/
│
├── src/
│ ├── ingestion/
│ │ ├── parser.py # PDF → text/table/image chunks (PyMuPDF + pdfplumber)
│ │ └── embedder.py # sentence-transformers wrapper
│ ├── retrieval/
│ │ ├── vector_store.py # ChromaDB wrapper (upsert, query, delete)
│ │ └── retriever.py # RAG pipeline orchestrator
│ ├── models/
│ │ ├── vlm.py # Gemini Vision (image → text summary)
│ │ └── llm.py # Gemini LLM (context → grounded answer)
│ └── api/
│ ├── schemas.py # Pydantic request/response models
│ └── routes.py # FastAPI route handlers
│
├── sample_documents/
│ └── tata_nexon_ev_technical_report.pdf # Multimodal EV sample PDF
│
└── screenshots/
├── 01_swagger_ui.png
├── 02_ingest_response.png
├── 03_query_text.png
├── 04_query_table.png
├── 05_query_image.png
└── 06_health_check.pngLicense
MIT License — see LICENSE for details.
BITS Pilani • Work Integrated Learning Programmes • Multimodal RAG Bootcamp • 2024
