Knowing What's Broken: Prometheus Metrics and an 88% Test Coverage Sprint
Part of the AI Software House series.
In short: Once the pipeline started running unattended, logs were no longer enough. I added Prometheus metrics so failures were visible quickly, then spent seven sprints taking test coverage from 37% to 88%.
Once I started leaving the pipeline to run overnight, two weaknesses became hard to ignore. I usually found failures by reading the logs the next morning, and every refactor still needed too much manual checking.
For v0.6.0 I tackled both: a small Prometheus metrics server, and a much less glamorous seven-sprint push on test coverage.
The observability gap
The watcher already had an event bus. Every significant thing that happened β a circuit breaker opening, a DLQ retry, a degradation policy activating β emitted an event. The events were well-structured dataclasses. Nobody was consuming them.
# core/events.py β events already existed and were already emitted
@dataclass
class CircuitBreakerEvent:
event_type: str = "circuit_breaker"
name: str = ""
state: str = "" # "open" | "closed" | "half_open"
timestamp: float = 0.0
So I didn't need another event system. I only needed somewhere to send the events and expose them to Prometheus.
The metrics server
metrics_server.py is a standalone HTTP server. It receives events via POST /event and exposes counters at GET /metrics for Prometheus scraping.
Three counters cover the failure modes that matter most in a long-running pipeline system:
| Counter | Labels | What it tracks |
|---|---|---|
aisw_circuit_breaker_events_total |
name, state |
Circuit breaker state transitions (open/closed/half-open) |
aisw_dlq_events_total |
action, backend |
Dead-letter-queue operations (enqueue, retry, discard) |
aisw_degradation_events_total |
trigger |
Degradation policy activations |
Starting it is one command:
METRICS_PORT=9091 python3 metrics_server.py
Wiring it to the watcher is one config line:
# watchers.yml
settings:
metrics_url: http://localhost:9091
Metrics must never block the pipeline
The watcher already depends on enough outside services. I did not want a broken metrics server to become one more reason for a pipeline to stall.
The sink is designed around this constraint. Every event is posted from a daemon thread:
def build_callback(metrics_url: str) -> Callable[[AnyEvent], None]:
endpoint = metrics_url.rstrip("/") + "/event"
def _callback(event: AnyEvent) -> None:
threading.Thread(target=_send, args=(event,), daemon=True).start()
def _send(event: AnyEvent) -> None:
try:
body = json.dumps(dataclasses.asdict(event)).encode()
req = urllib.request.Request(endpoint, data=body,
headers={"Content-Type": "application/json"},
method="POST")
with urllib.request.urlopen(req, timeout=1):
pass
except Exception as exc:
_log.debug("metrics_sink: failed to post %s: %s", event.event_type, exc)
return _callback
The timeout is 1 second. Any exception is swallowed and logged at DEBUG. The watcher never waits for metrics. If your Prometheus instance goes down, you lose metric data β you don't lose pipeline runs.
Grafana in 10 minutes
With the server running, a basic Prometheus scrape config:
# prometheus.yml
scrape_configs:
- job_name: 'ai-software-house'
scrape_interval: 15s
static_configs:
- targets: ['localhost:9091']
I then added a Grafana dashboard with three panels, one for each counter and grouped by label. It is simple, but it tells me within one scrape interval whether an overnight run is healthy.
The test coverage problem
The starting point was not good.
The system was built iteratively over several months. The core pipeline worked. The agents produced good output. The feature set had grown to 30+ capabilities listed in the README. And test coverage sat at 37%.
The codebase had simply grown much faster than its test suite. Too many new features were still tested by running a real pipeline and seeing what happened. Every refactor felt riskier than it should have.
Closing that gap took seven sprints. There wasn't a clever shortcut.
The sprint structure
Each sprint targeted a specific layer, running a consistent process:
- Measure β coverage report showing exact uncovered lines
- Plan β which paths matter most, written as a spec
- Implement β tests written to the spec using TDD
- Two-stage review β spec compliance first, then code quality
- Fix loops β reviewers found issues, implementer fixed them, reviewers re-approved
The two-stage review was important. Spec compliance asks: did the tests actually cover what we said they would? Code quality asks: are the assertions strong enough to catch real regressions, or do they just pass trivially?
A test that calls assert result is not None is worse than no test β it creates false confidence. The code quality review caught these and pushed for specificity.
What got covered
T8-T9: Test fixes β The oldest tests had accumulated against interfaces that had since changed. Before adding new coverage, the existing suite was made green.
T10: Critical paths β The paths that matter most when the system runs live: pipeline stage sequencing, checkpoint save/resume, error propagation and recovery.
T11-A: Event bus + metrics β The Prometheus metrics sink itself, tested against the actual event types it handles.
T11-B: Watcher polling β check_waiting_issues() and _process_resume_queue() β the two functions that run every hour and determine what work gets dispatched. These had been run but never unit-tested.
T12-A: Agents β Every agent's run_with_github() path: what gets posted as a comment, which emoji appears for which verdict, what happens when the LLM returns an unexpected format.
T12-B: Infrastructure β GitHub client API methods (PR creation, tree fetching, merge), CLI command coverage, orchestrator stage ordering and context passing.
T13: Low-priority paths β RepoAutoIndexer, refactor agent, watcher dispatch integration, DLQ retry flow end-to-end.
T14-T15: Final push β The remaining 87% β 88% gap: run_with_github methods on smaller agents, BaseAgent backend routing (Ollama, Anthropic, Nvidia NIM), truncate_files, __repr__, skills loader edge cases.
What 88% actually means
I don't think 88% is interesting by itself. The useful part is which paths are now covered.
The test suite now covers:
- Every agent's
run_with_github()path, including all verdict emojis - Every LLM backend construction path (6 backends, including all kwargs)
- Every watcher dispatch path β normal, resume, DLQ retry
- The checkpoint save/resume cycle, including atomic write and corruption recovery
- The pipeline stage sequencer, including loop blocks and conditional stages
- All three circuit breaker state transitions
- The skills loader, including marketplace fetch, cache, invalid YAML, and
for_role()filtering - The Prometheus metrics sink, including fire-and-forget failure handling
2,029 tests. 0 failures. Running in under 2 minutes.
The infrastructure cost
One practical lesson came out of all this: tests must construct agents without opening real LLM connections.
The agents have an __init__ that wires up LLM connections. Calling it in tests means live LLM calls. The right approach is bypassing __init__ entirely:
from unittest.mock import MagicMock
from agents.base_agent import BaseAgent
def _make_agent():
agent = SomeAgent.__new__(SomeAgent) # bypass __init__
agent._llm = MagicMock()
agent.system_prompt = ""
agent._history = []
agent.model = "gpt-4.1"
agent._backend = "github_models"
return agent
This pattern β __new__ + manual attribute injection β became standard across the entire test suite. It gives complete control over what the agent "knows" without any network activity.
What changed in practice
Before the coverage sprint, a refactor of BaseAgent would require manual testing across three or four agent types to verify nothing broke. Now it requires running the test suite β which takes under two minutes and tells you exactly which paths regressed.
The metrics server means that when a circuit breaker opens in a production run, a counter increments. A Grafana alert fires. You know within 15 seconds that your Ollama backend is struggling, rather than finding out in the morning when you read the logs.
Neither change makes a good demo. Together, though, they make the system much easier for me to trust when nobody is watching it.
Related reading: Pluggable Deploy Backends: Docker, VMs, and Nothing at All
Related reading: One Watcher, Any Pipeline: Label-Based Dispatch