CoolFace
Modelpublic

center-of-excellence/extract-prompt-quality-criteria

sourceHugging Facemitupdated 12d agoView on Hugging Face
1likes15downloads
Model Card

Prompt Quality Scoring — IBM Bob

Automatic prompt quality scoring for IBM Bob, IBM's AI coding assistant. Scores user prompts on a 1–10 scale validated against manually annotated IBM Bob prompts and confirmed on additional unseen prompts.

Primary scorer (v3): agentlans/bge-small-en-v1.5-prompt-quality + regex (90/10 blend). MIT license confirmed.


Performance

Evaluated on unseen IBM Bob prompts (honest zero-shot generalization):

ModelSpearman rMAEWithin 2pt %ms/prompt
agentlans+regex (recommended)0.7531.9253.9%~12.8ms
agentlans-only0.7381.9555.0%12.7ms
Regex+SetFit (v2)0.6791.7669.6%35.1ms
Regex-only0.6492.1763.4%0.264ms
TinyLlama v1 (original)0.3462.1662.5%1,680ms

The agentlans+regex blend improvement over agentlans-only is statistically significant (p=0.003, Bonferroni-corrected permutation test, n=191).


Quick Start

agentlans+regex scorer (recommended)

python
from src.agentlans_scorer import AgentlansScorer

scorer = AgentlansScorer(blend_regex=True)  # 90% agentlans + 10% regex
result = scorer.score("Fix the null pointer in login.py.")
print(result['score'])        # float 1.0–10.0
print(result['blend_ratio'])  # '90% agentlans + 10% regex'
print(result['agentlans_raw'])  # raw model output before calibration

Regex-only scorer (fastest, zero dependencies)

python
from src.regex_detector import score

result = score("Fix the null pointer in login.py.")
print(result['score'])        # float 1.0–10.0
print(result['explanation'])  # human-readable breakdown
print(result['criteria'])     # {'clear_task': True, 'examples_provided': False, ...}

sklearn transformer (data-bob integration)

python
from data_bob.prompt_quality_transformer import PromptQualityTransformer
import pandas as pd

df = pd.DataFrame({'text': ['Fix the bug.', '## Task\nImplement retry...']})
transformer = PromptQualityTransformer(prompt_column='text')
df_scored = transformer.fit_transform(df)
print(df_scored[['text', 'prompt_score', 'prompt_criteria', 'prompt_explanation']])

What It Scores

Six binary quality criteria:

CriterionDescription
clear_taskExplicit action verb and clear goal
examples_providedFile paths, code refs, before/after, sample data
structure_and_clarityHeaders, numbered lists, markdown sections
output_format_specifiedSpecifies JSON, code, list, specific file, language
constraints_definedTechnical or business constraints, versions, limits
edge_cases_handledError handling, null checks, fallback behavior

Score formula: score = 1 + (criteria_met / 6) × 9

Possible values: 1.0, 2.5, 4.0, 5.5, 7.0, 8.5, 10.0


Repository Contents

├── src/
│   ├── agentlans_scorer.py   # agentlans+regex blend scorer (v3, recommended)
│   ├── regex_detector.py     # standalone regex scorer, zero dependencies
│   ├── hybrid_scorer.py      # Regex+SetFit (v2, kept for reference)
│   └── evaluate.py           # metrics, permutation test, comparison table
│
├── setfit_models/            # v2 SetFit classifiers (kept for reference)
│   ├── output_format_specified/
│   ├── constraints_defined/
│   └── edge_cases_handled/
│
└── prompt_quality_transformer.py   # sklearn transformer (v3)

Licensing

  • —agentlans/bge-small-en-v1.5-prompt-quality — MIT license confirmed. Safe for IBM commercial use.
  • —BAAI/bge-small-en-v1.5 (base model) — MIT license.
  • —All code in this repository — MIT license.

Citation

bibtex
@misc{extract-prompt-quality-criteria,
  author = {Gregoire Cattan and Elif Uzun},
  title = {Extract Prompt Quality Criteria},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/center-of-excellence/extract-prompt-quality-criteria}
}