SahilCodevally/codevally-vision-language-action
0
๐ค Vision Language Action (VLA) Demo
Built by [Codevally](https://www.codevally.com) โ AI-Powered Software Development & AI Solutions
Upload an image and enter a natural-language command. The system detects objects using Grounding DINO (open-vocabulary), interprets your instruction with an LLM, generates a structured action plan, and visualizes the result.
Try it: "Pick up the defective board and place it in the red bin"
How it works
- ๐ธ Vision โ Grounding DINO detects objects based on your command text
- ๐ง Language โ LLM (OpenAI / Groq) interprets the command
- ๐ Planning โ Structured action steps are generated
- ๐ฌ Action โ Annotated visualization is produced
๐ข About Codevally
Codevally is a forward-thinking technology company at the intersection of AI and software development. We specialize in building intelligent systems, AI-powered applications, and custom AI solutions that transform businesses โ from traditional software to cutting-edge machine learning implementations.
๐ฌ Our AI Specializations
โ Why Codevally
- AI-First Approach โ Deep expertise in AI/ML and modern LLM technologies
- End-to-End Solutions โ From concept to deployment and monitoring (MLOps)
- Practical AI โ Focused on real business value, not just technology
- Responsible AI โ Ethics, bias detection, and governance built-in
This demo is an open-vocabulary Vision Language Action pipeline built by the Codevally AI team. [Learn more โ](https://www.codevally.com)
