Adityax-07/CodeSage
<div align="center">
<!-- Animated Banner --> <img width="100%" src="https://capsule-render.vercel.app/api?type=waving&color=gradient&customColorList=6,11,20&height=200§ion=header&text=CodeSage%20๐ง&fontSize=60&fontColor=fff&animation=twinkling&fontAlignY=35&desc=LLM%20vs%20RAG%20vs%20Fine-Tuning%20โ%20Live%20Trade-off%20Platform&descAlignY=60&descSize=18" />
<!-- Typing SVG --> <img src="https://readme-typing-svg.demolab.com?font=Fira+Code&weight=700&size=22&pause=1200&color=A78BFA¢er=true&vCenter=true&multiline=true&repeat=true&width=700&height=60&lines=Same+Question.+Three+Architectures.+Real+Numbers." alt="Typing SVG" />
<br/>
<!-- Primary Badges --> <p align="center"> <a href="https://github.com/Adityax-07/LLM-vs-RAG-vs-Fine-Tuning-/stargazers"> <img src="https://img.shields.io/github/stars/Adityax-07/LLM-vs-RAG-vs-Fine-Tuning-?style=for-the-badge&logo=starship&color=f59e0b&logoColor=white&labelColor=1a1a2e" /> </a> <a href="https://github.com/Adityax-07/LLM-vs-RAG-vs-Fine-Tuning-/network/members"> <img src="https://img.shields.io/github/forks/Adityax-07/LLM-vs-RAG-vs-Fine-Tuning-?style=for-the-badge&logo=git&color=8b5cf6&logoColor=white&labelColor=1a1a2e" /> </a> <img src="https://img.shields.io/badge/License-MIT-22c55e?style=for-the-badge&logo=opensourceinitiative&logoColor=white&labelColor=1a1a2e" /> <img src="https://img.shields.io/badge/Status-Production%20Ready-22c55e?style=for-the-badge&logo=checkmarx&logoColor=white&labelColor=1a1a2e" /> <img src="https://img.shields.io/badge/Python-3.10+-3776AB?style=for-the-badge&logo=python&logoColor=white&labelColor=1a1a2e" /> </p>
<!-- Live Demo Button --> <p align="center"> <a href="https://huggingface.co/spaces/Adityax-07/CodeSage"> <img src="https://img.shields.io/badge/๐ค%20Live%20Demo-Try%20CodeSage%20on%20HF%20Spaces-FFD21E?style=for-the-badge&labelColor=1a1a2e" /> </a> </p>
<!-- Tech Stack Badges --> <p align="center"> <img src="https://img.shields.io/badge/Streamlit-FF4B4B?style=for-the-badge&logo=streamlit&logoColor=white" /> <img src="https://img.shields.io/badge/LangChain-1C3C3C?style=for-the-badge&logo=langchain&logoColor=white" /> <img src="https://img.shields.io/badge/GroqAPI-F55036?style=for-the-badge&logo=groq&logoColor=white" /> <img src="https://img.shields.io/badge/HuggingFace-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black" /> <img src="https://img.shields.io/badge/FAISS-0467DF?style=for-the-badge&logo=meta&logoColor=white" /> <img src="https://img.shields.io/badge/LoRA%2FPEFT-EF4444?style=for-the-badge&logo=pytorch&logoColor=white" /> <img src="https://img.shields.io/badge/Plotly-3F4F75?style=for-the-badge&logo=plotly&logoColor=white" /> <img src="https://img.shields.io/badge/GoogleColab-F9AB00?style=for-the-badge&logo=googlecolab&logoColor=black" /> </p>
<!-- Skill Icons --> <p align="center"> <img src="https://skillicons.dev/icons?i=python,pytorch,tensorflow,git,github,vscode&theme=dark" /> </p>
<br/>
<blockquote> ๐งช <strong>CodeSage</strong> is a live, side-by-side AI research platform that fires the same programming question at three fundamentally different architectures โ <strong>Baseline LLM</strong>, <strong>RAG</strong>, and <strong>Fine-Tuning</strong> โ then auto-scores every answer on accuracy, hallucination, groundedness, relevance, and cost.<br/><br/> No cherry-picking. No manual grading. <strong>Real numbers, real trade-offs.</strong> </blockquote>
<br/>
<!-- Quick stats strip --> <p align="center"> <img src="https://img.shields.io/badge/50-Benchmark%20Questions-8b5cf6?style=flat-square" /> <img src="https://img.shields.io/badge/3-AI%20Systems%20Compared-06b6d4?style=flat-square" /> <img src="https://img.shields.io/badge/8-Auto%20Eval%20Metrics-f59e0b?style=flat-square" /> <img src="https://img.shields.io/badge/85.3%25-Fine--Tune%20Accuracy-22c55e?style=flat-square" /> <img src="https://img.shields.io/badge/0%25-Hallucination%20Rate-ef4444?style=flat-square" /> </p>
</div>
๐ Table of Contents
โก Benchmark Results
Full evaluation:3 systemsร50 Q&A pairsร8 metricsโ fully automated, zero manual grading
๐ Key Findings
๐ง What is CodeSage?
CodeSage is a decision-making tool for AI engineers. When building a domain-specific assistant, you always hit the same three-way fork:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Domain-Specific AI Assistant โ
โโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโ
โผ โผ โผ
โโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโ
โ BASELINE LLM โ โ RAG PIPELINE โ โ FINE-TUNING โ
โ โ โ โ โ โ
โ + Zero setup โ โ + Always fresh โ โ + 10x cheaper โ
โ + Broad topics โ โ + Grounded โ โ + 0% hallucin. โ
โ - Hallucinates โ โ - Retrieval lag โ โ - Hard to update โ
โโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโCodeSage makes this trade-off visible and measurable โ same question, same moment, real output from all three.
โจ Features
<div align="center"> <table> <tr> <td align="center" width="220"> <strong>๐ Side-by-Side Compare</strong><br/><br/> Three answers to one question,<br/>simultaneously, in one view </td> <td align="center" width="220"> <strong>๐ Auto Evaluation</strong><br/><br/> 8-metric LLM-as-Judge scores<br/>every response automatically </td> <td align="center" width="220"> <strong>๐ Winner Badge</strong><br/><br/> Best answer highlighted;<br/>hallucination flag raised on low-confidence </td> </tr> <tr> <td align="center"> <strong>๐ Analytics Dashboard</strong><br/><br/> Plotly charts + paper-style TABLE II<br/>aggregated over 50 benchmarks </td> <td align="center"> <strong>๐พ Persistent Cache</strong><br/><br/> Results stored in <code>benchmark_cache.json</code><br/>โ instant reload, no re-running </td> <td align="center"> <strong>๐ PDF Ingestion</strong><br/><br/> Drop any PDF into <code>data/pdfs/</code><br/>โ RAG ingests it automatically </td> </tr> </table> </div>
๐๏ธ Architecture
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๐ฅ๏ธ Streamlit UI โ
โ โโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโ โ
โ โ โก System 1 โ โ ๐ System 2 โ โ ๐ง System 3 โ โ
โ โ Baseline LLM โ โ RAG Pipeline โ โ Fine-Tuned โ โ
โ โโโโโโโโโโฌโโโโโโโโโโ โโโโโโโโโโโโฌโโโโโโโโโโโโโ โโโโโโโโโโฌโโโโโโโโโโ โ
โโโโโโโโโโโโโชโโโโโโโโโโโโโโโโโโโโโโโโชโโโโโโโโโโโโโโโโโโโโโโโโโชโโโโโโโโโโโโโโ
โ โ โ
โผ โผ โผ
โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโ
โ Groq API โ โ FAISS Index โ โ Qwen2.5-1.5B โ
โ Llama-3.1-8Bโ โ all-MiniLM-L6-v2 โ โ + LoRA Adapters โ
โ (zero-shot) โ โ (top-3 chunks) โ โ (PEFT, local) โ
โโโโโโโโโโโโโโโ โโโโโโโโโโฌโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโ
โ
Groq API (with context)
โ
โโโโโโโโโโโผโโโโโโโโโโโ
โ ๐๏ธ LLM-as-Judge โ
โ 8 metrics, auto โ
โโโโโโโโโโโโโโโโโโโโโโโก System 1 โ Baseline LLM
Sends the question directly to Llama-3.1-8B via Groq with a minimal system prompt. No extra knowledge. Represents what an off-the-shelf LLM can do โ the floor every other system must beat.
๐ System 2 โ RAG Pipeline
- Question โ
all-MiniLM-L6-v2embedding - Top-3 chunks retrieved from FAISS vector store (17 documents)
- Chunks injected as context into Llama-3.1-8B via Groq
- Groundedness scored โ answers must be traceable to retrieved text
๐ง System 3 โ Fine-Tuned Model
Qwen2.5-1.5B fine-tuned with LoRA (r=8, ฮฑ=32) on curated CS Q&A pairs via Google Colab T4 GPU. Adapters loaded locally via peft โ zero cloud inference cost, sub-second latency.
๐ Evaluation Pipeline
Each answer is auto-scored by an LLM judge across 8 dimensions:
๐ Quick Start
Step 1 โ Clone & Install
git clone https://github.com/Adityax-07/LLM-vs-RAG-vs-Fine-Tuning-.git
cd LLM-vs-RAG-vs-Fine-Tuning-
pip install -r requirements.txtStep 2 โ Configure API Key
echo "GROQ_API_KEY=your_key_here" > .env๐ Free key at console.groq.com
Step 3 โ Launch
streamlit run demo.pyFAISS vector store builds automatically on first launch. Systems 1 & 2 are ready instantly.
Step 4 โ (Optional) Activate Fine-Tuned Model
python -c "
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('Qwen/Qwen2.5-1.5B-Instruct')
model = PeftModel.from_pretrained(base, 'checkpoint-25')
model.merge_and_unload().save_pretrained('finetuned_model')
AutoTokenizer.from_pretrained('checkpoint-25').save_pretrained('finetuned_model')
"Or open system3_finetune_colab.ipynb in Google Colab to train from scratch on a free T4 GPU (~10 min).Step 5 โ Regenerate Benchmark (optional)
# Pre-computed results already included in data/benchmark_cache.json
python run_benchmark.py๐ Knowledge Base
The RAG system retrieves from 17 hand-crafted topic documents in data/docs/:
<div align="center"> <table> <tr> <td valign="top" width="33%"> <strong>๐งฎ Algorithms & DSA</strong><br/><br/> <code>binarysearch</code><br/> <code>sortingalgorithms</code><br/> <code>dynamicprogramming</code><br/> <code>graphalgorithms</code><br/> <code>trees</code><br/> <code>linkedlist</code><br/> <code>stackqueue</code><br/> <code>recursion</code><br/> <code>backtracking</code> </td> <td valign="top" width="33%"> <strong>๐ More DSA</strong><br/><br/> <code>greedyalgorithms</code><br/> <code>hashing</code><br/> <code>stringalgorithms</code><br/> <code>twopointers</code><br/> <code>bigonotation</code><br/> <code>heaps</code> </td> <td valign="top" width="33%"> <strong>๐ Web & Tooling</strong><br/><br/> <code>reacthooks</code><br/> <code>restapi</code><br/> <code>javascriptpromises</code><br/> <code>cssflexbox</code><br/> <code>typescriptbasics</code><br/> <code>sqlbasics</code><br/> <code>gitbasics</code> </td> </tr> </table> </div>
๐ก Decision Guide
๐ ๏ธ Tech Stack
<div align="center"> <table> <tr> <th>Layer</th> <th>Technology</th> <th>Purpose</th> </tr> <tr> <td>๐ <strong>UI</strong></td> <td> <img src="https://img.shields.io/badge/Streamlit-FF4B4B?style=flat-square&logo=streamlit&logoColor=white" /> <img src="https://img.shields.io/badge/Plotly-3F4F75?style=flat-square&logo=plotly&logoColor=white" /> </td> <td>3-way comparison dashboard + analytics charts</td> </tr> <tr> <td>โก <strong>LLM</strong></td> <td><img src="https://img.shields.io/badge/GroqAPI-F55036?style=flat-square&logo=groq&logoColor=white" /></td> <td>Llama-3.1-8B โ Baseline + RAG generation</td> </tr> <tr> <td>๐ค <strong>Embeddings</strong></td> <td><img src="https://img.shields.io/badge/sentence--transformers-FFD21E?style=flat-square&logo=huggingface&logoColor=black" /></td> <td><code>all-MiniLM-L6-v2</code> โ RAG semantic retrieval</td> </tr> <tr> <td>๐ <strong>Vector DB</strong></td> <td><img src="https://img.shields.io/badge/FAISS-0467DF?style=flat-square&logo=meta&logoColor=white" /></td> <td>CPU-based semantic search over knowledge base</td> </tr> <tr> <td>๐ง <strong>Fine-Tuning</strong></td> <td> <img src="https://img.shields.io/badge/PEFT%2FLoRA-EF4444?style=flat-square&logo=pytorch&logoColor=white" /> <img src="https://img.shields.io/badge/Transformers-FFD21E?style=flat-square&logo=huggingface&logoColor=black" /> </td> <td>LoRA adapter (r=8, ฮฑ=32) on Qwen2.5-1.5B</td> </tr> <tr> <td>๐๏ธ <strong>Base Model</strong></td> <td><img src="https://img.shields.io/badge/Qwen2.5--1.5B-FFD21E?style=flat-square&logo=huggingface&logoColor=black" /></td> <td>Alibaba's compact LLM โ LoRA fine-tuned locally</td> </tr> <tr> <td>โ๏ธ <strong>Training</strong></td> <td><img src="https://img.shields.io/badge/GoogleColab-F9AB00?style=flat-square&logo=googlecolab&logoColor=black" /></td> <td>Free T4 GPU โ LoRA training in ~10 minutes</td> </tr> <tr> <td>๐ <strong>Orchestration</strong></td> <td><img src="https://img.shields.io/badge/LangChain-1C3C3C?style=flat-square&logo=langchain&logoColor=white" /></td> <td>RAG pipeline, FAISS integration, PDF ingestion</td> </tr> <tr> <td>๐ <strong>Metrics</strong></td> <td><img src="https://img.shields.io/badge/rouge--score-EF4444?style=flat-square&logo=python&logoColor=white" /></td> <td>ROUGE-L + cosine similarity for auto-evaluation</td> </tr> </table> </div>
๐๏ธ Project Structure
๐ฆ LLM-vs-RAG-vs-Fine-Tuning/
โ
โโโ ๐ demo.py โ Streamlit app (main entry point)
โโโ ๐ system1_baseline.py โ Baseline LLM via Groq API
โโโ ๐ system2_rag.py โ RAG pipeline: FAISS + LangChain + Groq
โโโ ๐ system3_inference.py โ Fine-tuned model inference (PEFT)
โโโ ๐ system3_finetune_colab.ipynb โ LoRA training notebook (Colab T4)
โโโ ๐ evaluate.py โ Standalone evaluation script
โโโ ๐ run_benchmark.py โ Regenerates benchmark_cache.json
โ
โโโ ๐ checkpoint-25/ โ Trained LoRA weights (included)
โ โโโ adapter_model.safetensors
โ โโโ adapter_config.json โ r=8, alpha=32
โ โโโ tokenizer.json
โ
โโโ ๐ finetuned_model/ โ Merged model (after merge step)
โ
โโโ ๐ data/
โ โโโ ๐ docs/ โ 17 knowledge base .txt files
โ โโโ ๐ faiss_index/ โ FAISS vector store (auto-built)
โ โโโ ๐ pdfs/ โ Drop PDFs here for RAG ingestion
โ โโโ benchmark_cache.json โ Pre-computed 50Q benchmark results
โ โโโ reference_answers.json โ Ground-truth Q&A pairs
โ โโโ finetune_data.jsonl โ LoRA training data (ChatML format)
โ
โโโ ๐ requirements.txt๐ฎ Roadmap
<!-- Wave footer --> <img width="100%" src="https://capsule-render.vercel.app/api?type=waving&color=gradient&customColorList=6,11,20&height=120§ion=footer" />
<div align="center">
<strong>Built with ๐ง by <a href="https://github.com/Adityax-07">Adityax-07</a></strong>
<br/>
<em>Powered by Groq ยท HuggingFace ยท FAISS ยท LangChain ยท Streamlit</em>
<br/><br/>
<a href="https://github.com/Adityax-07"> <img src="https://img.shields.io/badge/GitHub-Adityax--07-181717?style=for-the-badge&logo=github&logoColor=white" /> </a>
<br/><br/>
โญ <strong>If CodeSage helped you understand the LLM trade-off space, drop a star!</strong>
</div>
