CoolFace
Apppublic

Vignesh0503/sentence-similarity-tool

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes
App README

<div align="center">

๐Ÿ“˜ AI Similarity Assist Tool ๐Ÿ”

<div align="center"> <div align="center">

Python Streamlit HuggingFace License Status

An intelligent dual-analysis system combining FAISS vector similarity with LLM-powered semantic understanding

๐Ÿš€ Live Demo โ€ข ๐Ÿ“– Documentation โ€ข ๐ŸŽฏ Use Cases โ€ข โšก Quick Start

</div>


๐ŸŒŸ Overview

The AI Similarity Assist Tool is a production-ready Streamlit application that revolutionizes text similarity analysis by combining two powerful approaches:

  1. 1.โšก FAISS Vector Search: Lightning-fast embedding-based similarity using BGE-large-en-v1.5
  2. 2.๐Ÿง  LLM Semantic Analysis: Deep relationship understanding powered by Groq's GPT-OSS-20B

This dual-layer approach ensures both speed and accuracy, making it perfect for document comparison, duplicate detection, semantic search, and content analysis tasks.


โœจ Key Features

๐ŸŽฏ Dual Analysis Engine

  • โ€”FAISS Similarity: Efficient vector-based matching with cosine similarity scores
  • โ€”LLM Analysis: Contextual relationship classification (Exact Match, Equivalent, Related, Contradictory)
  • โ€”Hybrid Results: Get both quantitative scores and qualitative insights

๐Ÿš€ Performance Optimized

  • โ€”Smart Caching: User-specific session management with automatic cleanup
  • โ€”Exact Match Detection: Bypasses embedding for identical strings (saves 50%+ compute time)
  • โ€”Batch Processing: Efficient LLM calls with token-aware batching
  • โ€”Progress Tracking: Real-time unified progress indicators across all phases

๐Ÿ“Š Rich Visualizations

  • โ€”3D Embedding Plot: Interactive Plotly visualization of semantic space
  • โ€”Comparative Charts: Distribution analysis for both FAISS and LLM results
  • โ€”Token Usage Metrics: Track LLM consumption for cost optimization
  • โ€”Detailed Statistics: Comprehensive breakdown by relationship types

๐Ÿ’พ Export & Integration

  • โ€”Excel Export: Highlighted differences with color-coded relationships
  • โ€”JSON Export: Structured data for downstream processing
  • โ€”Metadata Support: Preserve additional columns from source files
  • โ€”Truncation Indicators: Visual warnings for processed long texts

๐Ÿ”ง Enterprise Ready

  • โ€”Multi-User Support: Isolated sessions for concurrent users (up to 100)
  • โ€”Automatic Cleanup: 24-hour session timeout with garbage collection
  • โ€”Error Handling: Graceful degradation with fallback mechanisms
  • โ€”Logging: Comprehensive tracking for debugging and monitoring

๐ŸŽฏ Use Cases

DomainApplication
๐Ÿ“„ Legal & ComplianceContract comparison, clause matching, regulatory alignment
๐Ÿข Enterprise DataDuplicate detection, data deduplication, record linkage
๐Ÿ“š Content ManagementPlagiarism detection, content similarity, version tracking
๐ŸŽ“ Research & AcademiaLiterature review, citation analysis, semantic clustering
๐Ÿ›’ E-commerceProduct matching, review analysis, catalog normalization
๐Ÿ’ฌ Customer SupportTicket categorization, FAQ matching, response suggestions

๐Ÿ—๏ธ Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                     Streamlit Web Interface                 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  File Upload  โ†’  Column Selection  โ†’  Run Analysis          โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                      โ”‚
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ”‚   Pipeline Manager        โ”‚
        โ”‚   (progress_manager.py)   โ”‚
        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                      โ”‚
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ”‚                                             โ”‚
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  FAISS Search  โ”‚                          โ”‚  LLM Analysis    โ”‚
โ”‚  (core.py)     โ”‚                          โ”‚  (llm_service.py)โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค                          โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ โ€ข BGE-large    โ”‚                          โ”‚ โ€ข Groq GPT-OSS   โ”‚
โ”‚ โ€ข Embeddings   โ”‚                          โ”‚ โ€ข Batch Process  โ”‚
โ”‚ โ€ข Index Cache  โ”‚                          โ”‚ โ€ข Token Tracking โ”‚
โ”‚ โ€ข Exact Match  โ”‚                          โ”‚ โ€ข Result Cache   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
        โ”‚                                             โ”‚
        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                      โ”‚
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ”‚   Results Processing      โ”‚
        โ”‚   (postprocess.py)        โ”‚
        โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
        โ”‚ โ€ข Visualization           โ”‚
        โ”‚ โ€ข Excel/JSON Export       โ”‚
        โ”‚ โ€ข Summary Statistics      โ”‚
        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿš€ Quick Start

Prerequisites

bash
Python 3.9+
pip (Python package manager)

Installation

  1. 1.Clone the Repository
bash
git clone https://github.com/Vignesh-Manivasakam/sentence-similarity-tool.git
cd sentence-similarity-tool
  1. 1.Install Dependencies
bash
pip install -r requirements.txt
  1. 1.Set Up Environment Variables
bash
# Create .env file or set environment variable
export GROQ_API_KEY="your_groq_api_key_here"
  1. 1.Run the Application
bash
streamlit run app.py
  1. 1.Access the Tool
Open your browser and navigate to: http://localhost:8501

๐ŸŽฎ Usage Guide

Step 1: Upload Files

  • โ€”Base File: Your reference dataset (Excel format)
  • โ€”Check File: The dataset you want to compare against base

Step 2: Configure Columns

  • โ€”Select Identifier Column: Unique ID for each entry
  • โ€”Select Text Column: The content to analyze
  • โ€”(Optional) Select Additional Columns: Metadata to preserve

Step 3: Set Parameters

  • โ€”Top K Matches: Number of similar entries to return per query (1-10)

Step 4: Run Analysis

  • โ€”Click ๐Ÿš€ Run Similarity Search
  • โ€”Watch the unified progress bar track all phases:
  • โ€”Phase 1: Preprocessing & Embedding Generation
  • โ€”Phase 2: FAISS Similarity Search
  • โ€”Phase 3: LLM Relationship Analysis

Step 5: Explore Results

  • โ€”Summary View: High-level statistics and metrics
  • โ€”Visualization: 3D embedding space with interactive exploration
  • โ€”Results Table: Detailed matches with highlighted differences
  • โ€”Download: Export as Excel (with formatting) or JSON

๐Ÿ“Š Sample Results

LLM Relationship Classifications

RelationshipScore RangeDescriptionExample
๐ŸŸข Exact Match1.0Identical strings"The quick brown fox" โ†” "The quick brown fox"
๐ŸŸข Equivalent0.95-1.0Same meaning, different words"automobile" โ†” "car"
๐ŸŸก Related0.50-0.94Partial overlap or related concepts"sedan" โ†” "vehicle"
๐Ÿ”ด Contradictory0.00-0.20Opposite meanings"hot" โ†” "cold"

๐Ÿ› ๏ธ Technical Stack

Core Technologies

  • โ€”Framework: Streamlit 1.28+
  • โ€”Vector Search: FAISS (Facebook AI Similarity Search)
  • โ€”Embeddings: BAAI/bge-large-en-v1.5 (1024 dimensions)
  • โ€”LLM: Groq GPT-OSS-20B via Groq API
  • โ€”Visualization: Plotly, Matplotlib
  • โ€”Data Processing: Pandas, NumPy, SciKit-Learn

Key Libraries

streamlit>=1.28.0
sentence-transformers>=2.2.2
faiss-cpu>=1.7.4
groq>=0.4.0
plotly>=5.14.0
openpyxl>=3.1.0
pandas>=2.0.0

โš™๏ธ Configuration

Environment Variables

VariableDescriptionDefault
GROQ_API_KEYGroq API key for LLM analysisRequired
LOG_LEVELLogging verbosityINFO
CUDA_AVAILABLEEnable GPU accelerationfalse

Configurable Parameters (app/config.py)

python
# Embedding Configuration
EMBEDDING_MODEL_NAME = 'BAAI/bge-large-en-v1.5'
EMBEDDING_DIMENSION = 1024
EMBEDDING_BATCH_SIZE = 32

# Text Processing
MAX_TOKENS_FOR_TRUNCATION = 512

# LLM Configuration
GROQ_MODEL = "openai/gpt-oss-20b"
GROQ_TEMPERATURE = 0.0
LLM_BATCH_TOKEN_LIMIT = 100000

# Session Management
SESSION_TIMEOUT_HOURS = 24
MAX_CONCURRENT_USERS = 100

๐Ÿ“ Project Structure

sentence-similarity-tool/
โ”‚
โ”œโ”€โ”€ app.py                      # Main Streamlit application
โ”œโ”€โ”€ requirements.txt            # Python dependencies
โ”œโ”€โ”€ Dockerfile                  # Container configuration
โ”‚
โ”œโ”€โ”€ app/
โ”‚   โ”œโ”€โ”€ config.py              # Configuration & environment setup
โ”‚   โ”œโ”€โ”€ core.py                # FAISS indexing & similarity search
โ”‚   โ”œโ”€โ”€ llm_service.py         # Groq LLM integration & batching
โ”‚   โ”œโ”€โ”€ pipeline.py            # Orchestration of analysis workflow
โ”‚   โ”œโ”€โ”€ preprocess.py          # Text cleaning & normalization
โ”‚   โ”œโ”€โ”€ postprocess.py         # Results formatting & export
โ”‚   โ”œโ”€โ”€ progress_manager.py    # Unified progress tracking
โ”‚   โ””โ”€โ”€ utils.py               # Helper functions & utilities
โ”‚
โ”œโ”€โ”€ app/prompts/
โ”‚   โ””โ”€โ”€ system_prompt.txt      # LLM prompt template
โ”‚
โ””โ”€โ”€ static/
    โ””โ”€โ”€ css/
        โ””โ”€โ”€ custom.css         # UI styling & themes

๐Ÿ”ฌ Advanced Features

Hierarchy-Based Tie Breaking

When multiple entries have identical similarity scores, the tool uses hierarchical identifiers (e.g., "1.2.3") to select the most structurally relevant match.

Smart Caching Strategy

  • โ€”User-Specific Cache: Each session maintains isolated embeddings and indices
  • โ€”Hash-Based Validation: Automatic cache invalidation on file changes
  • โ€”LLM Response Cache: Avoid redundant API calls for identical pairs

Token Optimization

  • โ€”Batch Aggregation: Groups LLM calls to maximize tokens per request
  • โ€”Dynamic Batching: Respects token limits while minimizing API calls
  • โ€”Usage Tracking: Real-time display of prompt/completion token consumption

๐ŸŽจ Screenshots

Main Interface

Main Interface Upload files, select columns, and configure analysis parameters

Processing Pipeline

Processing Real-time progress across preprocessing, FAISS search, and LLM analysis

Results Dashboard

Results Comprehensive summary with FAISS and LLM insights

Results Table

Table Color-coded relationships with highlighted text differences


๐Ÿงช Testing

Run tests with:

bash
# Unit tests
pytest tests/

# Integration tests
pytest tests/integration/

# Coverage report
pytest --cov=app tests/

๐Ÿค Contributing

Contributions are welcome! Please follow these steps:

  1. 1.Fork the repository
  2. 2.Create a feature branch (git checkout -b feature/AmazingFeature)
  3. 3.Commit your changes (git commit -m 'Add some AmazingFeature')
  4. 4.Push to the branch (git push origin feature/AmazingFeature)
  5. 5.Open a Pull Request

๐Ÿ“ License

This project is licensed under the MIT License - see the LICENSE file for details.


๐Ÿ™ Acknowledgments

  • โ€”BAAI for the BGE-large-en-v1.5 embedding model
  • โ€”Groq for ultra-fast LLM inference
  • โ€”Facebook Research for FAISS vector search library
  • โ€”Streamlit for the amazing web framework
  • โ€”HuggingFace for model hosting and deployment

๐Ÿ“ง Contact

Vignesh Manivasakam


<div align="center">

โญ Star this repository if you find it helpful! โญ

Made with โค๏ธ by Vignesh Manivasakam

</div>