ruggsea/Llama3-philosophy-demo
Reduce GPU reservation from 120s to 60s per call
Add /chat_with_history API endpoint for multi-turn conversations
Fix prefix stripping: handle at chunk level before accumulation
Strip 'assistant' prefix artifact from model responses
fix: revert to working minimal version (no additional_inputs)
feat: restore additional inputs (system prompt, sliders) with working API
fix: simplify to minimal ChatInterface to debug zero endpoints issue
fix: add explicit type='messages' for Gradio 5 compatibility
fix: match official HF ZeroGPU pattern exactly
fix: use .to('cuda') instead of device_map='auto' for ZeroGPU compatibility
fix: migrate to ZeroGPU (free H200), remove broken bnb 4-bit quantization
req
push
redesign
Fixing the chat history
roba
fixing
fix
fixing
fix
fixing
fixed structure
style
update requirements
update
Fixing css
fix scrolling
reverted
fix
requirements
changed model, added system prompt
changing ui
Updated to use ruggsea/Llama3.1-8B-SEP-Chat with multi-turn support
Update app.py
Update README.md
Update app.py
Update app.py
Update app.py
Update app.py
Adapting space to philosophy model
Update
Update
Upgrade Gradio to version 4.26.0 (#52)
Update
Update
Update
Update
Update
Update
Update
