# Prompt Versioning and Cost Control: Running LLMs Responsibly in Production

The monthly bill was ten times the estimate.

Investigation took two hours. Three things were wrong.

A retry loop was calling the model on every validation error — including schema mismatches that would never succeed regardless of how many times you retried. Each failed extraction attempt cost tokens. The loop retried five times before giving up. Five model calls per bad document.

Prompts were three times longer than necessary. They'd been iterated during development — adding examples, clarifying instructions, handling edge cases — and nobody had trimmed them after the outputs stabilised. The final prompts were carrying three paragraphs of instructions that could have been one.

One new feature was calling GPT-4o for a classification task that returned one of four labels. GPT-4o-mini handles classification reliably at a fraction of the cost. Nobody had checked.

No alerting. No per-feature visibility. No limits. The first signal was the bill.

* * *

## Why LLM Costs Spiral

The three failure modes above share a root cause: LLM API costs are invisible until they're not. Unlike compute or storage, where you can see utilisation in real time, token costs accumulate silently until the billing cycle ends. By then the damage is done.

Three things create runaway costs:

**Retries without discrimination.** Retrying transient errors (rate limits, 500s) is correct. Retrying validation failures — where the model returned structurally wrong output — is not. The model doesn't know it returned bad JSON. Retrying it with the same prompt produces the same bad JSON. You've just multiplied your costs by your retry count.

**Prompt bloat.** Prompts grow during development and rarely shrink. Every added example, every clarification, every edge case instruction adds tokens to every call. A prompt that's 2,000 tokens instead of 600 costs 3x more on every invocation.

**Wrong model tier.** GPT-4o is significantly more capable than GPT-4o-mini — and significantly more expensive. For tasks that don't need the capability gap (classification, extraction from structured input, summarisation with clear constraints), the cheaper model is the right choice.

* * *

## Counting Tokens Before You Send

Don't estimate token costs from output — measure them before the call. `tiktoken` gives you exact token counts for OpenAI models:

```python
# services/token_counter.py
import tiktoken
from typing import list

PRICING = {
    "gpt-4o":       {"input": 2.50,  "output": 10.00},  # per 1M tokens
    "gpt-4o-mini":  {"input": 0.15,  "output": 0.60},
    "gemini-1.5-flash": {"input": 0.075, "output": 0.30},
}

def count_tokens(text: str, model: str = "gpt-4o-mini") -> int:
    enc = tiktoken.encoding_for_model(model)
    return len(enc.encode(text))

def estimate_cost(prompt_tokens: int, completion_tokens: int, model: str) -> float:
    pricing = PRICING.get(model, {"input": 0, "output": 0})
    return (
        (prompt_tokens / 1_000_000) * pricing["input"] +
        (completion_tokens / 1_000_000) * pricing["output"]
    )
```

Log actual usage on every call — the SDK returns it in the response:

```python
# services/llm_client.py
async def call_model(prompt: str, system: str, model: str = "gpt-4o-mini", feature: str = "unknown") -> str:
    response = await client.chat.completions.create(
        model=model,
        messages=[
            {"role": "system", "content": system},
            {"role": "user", "content": prompt},
        ],
        max_tokens=1000,
    )

    usage = response.usage
    cost = estimate_cost(usage.prompt_tokens, usage.completion_tokens, model)

    logger.info("llm_call", extra={
        "feature":           feature,
        "model":             model,
        "prompt_tokens":     usage.prompt_tokens,
        "completion_tokens": usage.completion_tokens,
        "total_tokens":      usage.total_tokens,
        "estimated_cost_usd": round(cost, 6),
    })

    return response.choices[0].message.content
```

With structured logging to Cloud Logging, you can query total cost per feature:

```plaintext
jsonPayload.feature="document_extraction" | sum(jsonPayload.estimated_cost_usd)
```

* * *

## Hard Limits

A token count check before sending prevents runaway calls from ever reaching the API:

```python
# services/llm_client.py
MAX_PROMPT_TOKENS = {
    "gpt-4o":      8_000,
    "gpt-4o-mini": 4_000,
}

async def call_model_safe(prompt: str, system: str, model: str = "gpt-4o-mini", feature: str = "unknown") -> str:
    combined = system + prompt
    token_count = count_tokens(combined, model)
    limit = MAX_PROMPT_TOKENS.get(model, 4_000)

    if token_count > limit:
        raise ValueError(f"Prompt too long: {token_count} tokens exceeds limit of {limit} for {model}")

    return await call_model(prompt, system, model, feature)
```

For a daily budget cap, use Redis as a counter:

```python
# services/budget.py
import aioredis
from config import settings

redis = aioredis.from_url(settings.redis_url)
DAILY_BUDGET_USD = 50.0

async def check_and_record_cost(feature: str, cost_usd: float) -> None:
    today = datetime.utcnow().strftime("%Y-%m-%d")
    key = f"llm_cost:{today}"

    current = await redis.get(key)
    current_total = float(current or 0)

    if current_total + cost_usd > DAILY_BUDGET_USD:
        raise RuntimeError(f"Daily LLM budget of ${DAILY_BUDGET_USD} exceeded. Current: ${current_total:.2f}")

    pipe = redis.pipeline()
    pipe.incrbyfloat(key, cost_usd)
    pipe.expire(key, 86400 * 2)   # keep for 2 days for debugging
    await pipe.execute()
```

Call `check_and_record_cost` after each successful model call, not before — you don't know the actual cost until the response arrives.

* * *

## Prompt Versioning in Config

Prompts that live as string constants in service files can't be changed without a code deployment. They can't be A/B tested. They can't be rolled back independently when a model update changes how they behave.

Store prompts in versioned config files, loaded at startup:

```json
// prompts/prod.json
{
  "document_extraction": {
    "v1": {
      "model": "gpt-4o",
      "system": "You are a document extraction assistant...",
      "user_template": "Extract: {fields}\n\nDocument: {document}",
      "max_tokens": 2000,
      "notes": "Original — verbose, high accuracy"
    },
    "v2": {
      "model": "gpt-4o-mini",
      "system": "Extract structured data. Return JSON only. If a field is absent return null.",
      "user_template": "Fields: {fields}\n\n{document}",
      "max_tokens": 1000,
      "notes": "Trimmed prompt — 60% cheaper, same accuracy on structured docs"
    }
  },
  "classification": {
    "v1": {
      "model": "gpt-4o-mini",
      "system": "Classify the input into one of: positive, negative, neutral, unclear. Return the label only.",
      "user_template": "{text}",
      "max_tokens": 10
    }
  }
}
```

The `model` field lives in the prompt config — not hardcoded in service code. Switching a feature from GPT-4o to GPT-4o-mini is a config change, not a code change. It's reviewable, rollbackable, and doesn't require a deployment.

To A/B test cost vs quality:

```python
import random

def get_prompt_for_request(key: str) -> dict:
    versions = prompts[key]
    # 10% of traffic on v1 (GPT-4o) to monitor quality, 90% on v2 (cheaper)
    if random.random() < 0.1:
        return versions["v1"]
    return versions["v2"]
```

Log which version was used on every call. After a week of data, compare `estimated_cost_usd` and output quality (however you measure it) between versions. If quality holds, promote v2 to 100%.

* * *

## The Checklist

Before shipping any LLM feature:

*   \[ \] Prompt tokens counted and logged on every call
    
*   \[ \] Retry logic distinguishes transient errors from validation failures
    
*   \[ \] Model tier chosen based on task complexity, not default
    
*   \[ \] Prompts stored as versioned config, not string constants
    
*   \[ \] Token limit check before sending
    
*   \[ \] Daily cost budget with alerting
    

None of this is complex. It's just easier to skip when you're moving fast. The bill is how you find out you skipped it.
