Skip to main content

Command Palette

Search for a command to run...

Testing AI-Dependent Code: Mocking LLMs, Evaluating Outputs, and Avoiding Flaky Tests

Updated
•6 min read•View as Markdown
Testing AI-Dependent Code: Mocking LLMs, Evaluating Outputs, and Avoiding Flaky Tests

The test checked whether the extracted invoice total matched the expected value. It passed nine times out of ten. On the tenth run, the model formatted the number differently — "1,234.56" instead of "1234.56". Exact string equality, both technically correct, test failed.

Developers learned to re-run CI when it went red. Nobody fixed the test. Six months later the suite had twelve flaky tests. People had stopped trusting CI entirely — a genuinely broken change could slip through unnoticed because the red runs were noise.

Flaky AI tests aren't a testing framework problem. They're an assertion strategy problem.


Why AI Tests Are Different

Three things make testing AI-dependent code harder than testing regular service code:

Non-determinism. The same prompt with temperature > 0 produces different outputs on every call. Exact value assertions will fail intermittently. This isn't a bug — it's the model working as designed.

Cost and latency. Every real API call costs money and adds 5–30 seconds to your test suite. A test suite that calls OpenAI on every run is slow, expensive, and breaks when you're offline.

Output structure vs output value. For extraction and classification tasks, the structure of the output (valid JSON, correct field names, right data types) is deterministic even when the values aren't. Asserting on structure is reliable. Asserting on values often isn't.


Mock at the Client Level

The right place to mock LLM calls is at the client — not at the HTTP layer, not with unittest.mock.patch scattered across test files. A pytest fixture that replaces the client once, cleanly:

# tests/conftest.py
import pytest
from unittest.mock import AsyncMock, MagicMock
from openai.types.chat import ChatCompletion, ChatCompletionMessage, Choice
from openai.types import CompletionUsage

def make_completion(content: str, model: str = "gpt-4o-mini") -> ChatCompletion:
    """Helper to build a realistic ChatCompletion response."""
    return ChatCompletion(
        id="test-completion-id",
        choices=[Choice(
            finish_reason="stop",
            index=0,
            message=ChatCompletionMessage(role="assistant", content=content),
        )],
        created=1700000000,
        model=model,
        object="chat.completion",
        usage=CompletionUsage(prompt_tokens=100, completion_tokens=50, total_tokens=150),
    )

@pytest.fixture
def mock_openai(monkeypatch):
    """Replace the AsyncOpenAI client with a controllable mock."""
    mock_client = AsyncMock()
    monkeypatch.setattr("services.llm_client.client", mock_client)
    return mock_client

Usage in tests:

# tests/integration/test_extraction.py
import pytest
import json
from services.extraction import extract_document

@pytest.mark.anyio
async def test_extraction_valid_output(mock_openai):
    expected_output = {
        "invoice_number": "INV-001",
        "total_amount": 1234.56,
        "vendor_name": "Acme Corp",
        "date": "2025-09-01",
        "line_items": []
    }
    mock_openai.chat.completions.create.return_value = make_completion(
        json.dumps(expected_output)
    )

    result = await extract_document("test document text")

    assert result.invoice_number == "INV-001"
    assert result.total_amount == 1234.56
    mock_openai.chat.completions.create.assert_called_once()

Zero API calls. Deterministic. Runs in milliseconds. The fixture controls exactly what the model returns, so you can test every code path — including the ones where the model returns malformed output.


Testing Structured Output — Schema, Not Value

For extraction tasks, the most important assertion is that the output matches the expected schema — not that it contains specific values. Schema validity is deterministic even when output values aren't:

# tests/integration/test_extraction.py
@pytest.mark.anyio
async def test_extraction_schema_validity(mock_openai):
    # Simulate a realistic but variable model output
    mock_openai.chat.completions.create.return_value = make_completion(
        '{"invoice_number": "INV-002", "total_amount": 500.0, "vendor_name": "Vendor", "date": "2025-09-02", "line_items": []}'
    )

    result = await extract_document("any document")

    # Assert schema — not specific values
    assert isinstance(result.invoice_number, str)
    assert len(result.invoice_number) > 0
    assert isinstance(result.total_amount, float)
    assert result.total_amount > 0
    assert isinstance(result.line_items, list)

For Pydantic models, the validation itself is the schema test — if extract_document returns an ExtractionResult without raising, the schema is valid. Your assertions can focus on business logic rather than type checking.


Snapshot Testing for Prompts

Prompts change — during development, after model updates, when requirements shift. Unexpected prompt changes can silently degrade output quality. Snapshot tests catch them:

# tests/unit/test_prompts.py
from prompts.store import get_prompt
import json

def test_extraction_prompt_unchanged(snapshot):
    prompt_config = get_prompt("document_extraction", version="v2")
    # Snapshot saves the first run, asserts equality on subsequent runs
    assert snapshot == json.dumps(prompt_config, sort_keys=True, indent=2)

Use pytest-snapshot for this. On the first run, it saves the output as a .txt file in a snapshots/ directory. On subsequent runs, it asserts equality. When a prompt legitimately changes, you update the snapshot intentionally — the change is visible in the PR diff.

This is how you catch "I just tweaked the wording" prompt changes that alter model behaviour without anyone realising.


What to Assert on Non-Deterministic Output

For classification tasks where the model returns one of a fixed set of labels, exact assertion is safe:

@pytest.mark.anyio
async def test_classification_returns_valid_label(mock_openai):
    mock_openai.chat.completions.create.return_value = make_completion("positive")

    result = await classify_sentiment("Great product!")

    assert result in {"positive", "negative", "neutral", "unclear"}

For generation tasks where the output is free-form, assert on properties rather than values:

@pytest.mark.anyio
async def test_summary_reasonable_length(mock_openai):
    mock_openai.chat.completions.create.return_value = make_completion(
        "This is a summary of the document covering the main points."
    )

    result = await summarise_document("long document text...")

    assert len(result) > 20       # not empty
    assert len(result) < 2000     # not truncated mid-sentence
    assert not result.startswith("```")  # no markdown fences

The assertions test that your code handled the output correctly — stripped fences, checked length, returned a string — without asserting on content that would vary with each real model call.


Integration Tests That Call the Real API

Unit tests with mocks cover the code paths. Integration tests against the real API catch prompt regressions and model behaviour changes. Run them separately — not on every push.

# tests/integration/test_real_api.py
import pytest

@pytest.mark.real_api   # custom marker — skipped unless explicitly enabled
@pytest.mark.anyio
async def test_extraction_real_api():
    result = await extract_document(SAMPLE_INVOICE_TEXT)
    assert result.total_amount > 0
    assert len(result.invoice_number) > 0
# pytest.ini
[pytest]
markers =
    real_api: marks tests that call the real LLM API (deselect with '-m "not real_api"')
# .github/workflows/test.yml
- name: Run tests (no real API)
  run: pytest -m "not real_api"

# Run weekly or manually
- name: Run real API tests
  if: github.event_name == 'schedule'
  run: pytest -m real_api
  env:
    OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

Weekly scheduled runs catch model behaviour drift. Manual runs before a prompt change goes to prod. Never in the default push pipeline.


The Rule

Mock the client, assert on schema not values, snapshot the prompts, gate real API tests behind a flag. The twelve flaky tests from the opening story all had the same fix: stop asserting on the exact string the model returned and start asserting on what your code did with it.

AI-Powered Backend Services

Part 5 of 5

LLM APIs don't fail the way REST APIs do. A 200 response with a well-formed body can still contain output you can't use, costs you didn't expect, or behaviour that changed silently after a model update. This series covers building AI-powered backend services that hold up in production — prompt management, output validation, async pipelines, cost control, testing non-deterministic code, and deployment patterns for variable-latency workloads. Written from real experience integrating OpenAI and Gemini into production FastAPI services on GCP.

Start from the beginning

Building AI-Powered Backend Services: Why Most Integrations Break in Production

The integration looked solid. In testing, the AI-powered document extraction service worked exactly as designed — pull a PDF, send it to the model, parse the structured output, write to the database.