SyncShift/sql-correction-env
SQL Correction RL Environment
An OpenEnv-compliant reinforcement learning environment where an AI agent learns to fix broken SQL queries — a real task that developers face every day.
Description & Motivation
SQL errors are one of the most common and costly mistakes in software development. This environment trains agents to identify and correct SQL syntax and logical errors, ranging from simple typos to complex multi-join query reconstruction including column name mismatches.
The environment provides partial progress signals at every step; the agent receives graded feedback even for near-correct answers, enabling meaningful learning across the full trajectory rather than sparse end-of-episode rewards. A stagnation penalty further discourages agents from repeating the same wrong answer across steps.
Observation Space
Action Space
Tasks
Reward Function
A stagnation penalty of −0.1 is applied when the agent submits the same reward-equivalent answer for two or more consecutive steps, encouraging active correction rather than looping.
Episodes terminate when reward reaches 0.99 (success) or max steps is reached.
Setup & Usage
Local Development
# Clone and install
git clone https://huggingface.co/spaces/YOUR_USERNAME/sql-correction-env
cd sql-correction-env
pip install -r requirements.txt
# Start the server
uvicorn server:app --host 0.0.0.0 --port 7860
# Test endpoints
curl -X POST http://localhost:7860/reset \
-H "Content-Type: application/json" -d '{"difficulty": "easy"}'
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{"action": {"corrected_query": "SELECT * FROM users WHERE id = 1"}}'
curl http://localhost:7860/tasksRun Tests
pip install pytest
pytest tests/ -vDocker
docker build -t sql-correction-env .
docker run -p 7860:7860 sql-correction-envRunning Inference
export HF_TOKEN=your_token
export API_BASE_URL=https://router.huggingface.co/v1
export MODEL_NAME=Qwen/Qwen2.5-72B-Instruct
export ENV_URL=http://localhost:7860
# Run all tasks
python inference.pyBaseline Scores
Run `inference.py` against the live Space to reproduce these scores.
