krishna0506/Document_Intelligence
Document Intelligence RAG System
A Retrieval-Augmented Generation (RAG) based document question-answering system built using FastAPI, LangChain, ChromaDB, and Groq LLM.
This project allows users to ingest documents, create embeddings, retrieve relevant context, and generate AI-powered answers.
Features
- Document ingestion and chunking
- Vector embeddings using ChromaDB
- Semantic search and retrieval
- LLM-based answer generation
- FastAPI REST API
- Modular project structure
Project Structure
documentintelligence/ │ ├── app/ │ ├── api.py │ ├── chunking.py │ ├── embeddings.py │ ├── generator.py │ ├── guardrails.py │ ├── ingest.py │ ├── main.py │ ├── retriever.py │ ├── schemas.py │ └── vectorstore.py │ ├── data/ │ ├── clientxrequirements.txt │ ├── screeningchecklistpython.txt │ ├── compliancepolicy.txt │ ├── ratecard2026.txt │ └── placementchecklist.txt │ ├── requirements.txt └── README.md
Setup Instructions
1. Create Virtual Environment
Windows: python -m venv .venv .venv\Scripts\activate
Mac/Linux: python -m venv .venv source .venv/bin/activate
2. Install Dependencies
pip install -r requirements.txt
3. Environment Variables
Create a file named .env in the project root:
GROQAPIKEY=yourapikey_here
4. Run the Application
uvicorn document_intelligence.app.main:app --reload
Open in browser: http://127.0.0.1:8000/docs
How It Works
- Add documents to the data/ folder.
- Run ingestion to create embeddings.
- Ask questions through API endpoints.
- The system retrieves relevant content and generates answers.
Technologies
- Python
- FastAPI
- LangChain
- ChromaDB
- Groq LLM
Notes
- Keep documents organized inside the data folder.
- Update paths in ingest.py if you change folder structure.
- Store API keys securely using .env file.
License
Educational and internal project use.
