From Pipeline to API: The REST and MCP Integration Layer
Part of the AI Software House series.
In short: GitHub labels were the pipeline's only front door. I added a FastAPI and MCP layer so I can start and monitor a run from curl, a web UI, or an AI coding assistant without first opening GitHub.
For a long time, the only way to start a pipeline was with a GitHub label.
Add ai-feature to an issue. The watcher polls, finds it, and dispatches the pipeline. The result β a PR, a comment, a status update β comes back via GitHub.
That was a reasonable first interface. The watcher already understood GitHub, labels are easy to audit, and the whole team can see them.
Eventually it became awkward. Starting a run from a terminal, another CI system, or a chat assistant still meant detouring through the GitHub UI.
So I gave the existing pipeline another front door.
What the integration layer is
aisw_server.py starts a single process that exposes two interfaces on the same port:
- A REST API β standard HTTP endpoints for submitting requirements, tracking job state, streaming logs, and cancelling runs
- An MCP server β auto-generated from the REST routes, so any MCP-compatible tool (Copilot CLI, Claude Code, OpenCode) can discover and call the same operations as tool calls
Both interfaces live on port 8765 by default. There's no separate MCP process. The MCP server is a bridge over the REST API.
The REST API
Six endpoints:
| Method | Path | What it does |
|---|---|---|
GET |
/health |
Liveness check β no auth required |
POST |
/runs |
Submit a new pipeline requirement |
GET |
/runs |
List recent runs |
GET |
/runs/{id} |
Get full detail for a single run |
DELETE |
/runs/{id} |
Cancel a queued or running job |
GET |
/runs/{id}/stream |
Stream live log output as SSE |
Submitting a requirement:
curl -X POST http://localhost:8765/runs \
-H "X-API-Key: your-key" \
-H "Content-Type: application/json" \
-d '{"requirement": "Build a bookmark manager REST API", "repo": "me/my-repo"}'
Response:
{
"run_id": "4a7c3e91-...",
"status": "queued",
"stream_url": "/runs/4a7c3e91-.../stream"
}
The job is queued immediately. The pipeline runs asynchronously in a background thread. The client gets back a run_id and knows where to watch the output.
Live log streaming
The streaming endpoint is Server-Sent Events:
curl -N http://localhost:8765/runs/4a7c3e91-.../stream \
-H "X-API-Key: your-key"
event: log
data: [PM] Reading requirement...
event: log
data: [Architect] Designing module structure...
event: log
data: [Engineer] Implementing auth module...
event: done
data: {"verdict": "approved", "pr_url": "https://github.com/me/my-repo/pull/47", ...}
The stream replays everything written since the job started, then tails live output until the job reaches a terminal state. If you connect after the job has finished, you get the full log replay followed by the terminal event β same format, no special case on the client side.
The job store
Job state is persisted to SQLite. This means:
- The server can restart without losing track of running or completed jobs
- Clients can query past runs hours or days later
- The
interruptedstatus marks jobs that were running when the server stopped β so they don't show as permanently "running"
On startup, init_db() finds any running jobs from the previous process and marks them interrupted:
def init_db(self) -> None:
# ... create table if not exists ...
# Mark any running jobs from before restart as interrupted
conn.execute(
"UPDATE jobs SET status = 'interrupted', updated_at = ? WHERE status = 'running'",
(_now(),)
)
Clean restart state without requiring a separate migration step.
Concurrency: wrapping a synchronous orchestrator
The existing Orchestrator is synchronous β it was built to run from the command line or from a GitHub Actions runner where blocking the process is fine. The API server needs to serve multiple concurrent requests without blocking.
I kept the orchestrator synchronous and put it behind a ThreadPoolExecutor:
class JobRunner:
def __init__(self, ..., max_workers: int = 4):
self._executor = ThreadPoolExecutor(max_workers=max_workers)
def submit(self, req: RunRequest) -> str:
run_id = str(uuid.uuid4())
self.store.insert_job(...)
self._executor.submit(self._run_job, run_id, job)
return run_id
Each job gets a thread from the pool. _run_job runs the synchronous orchestrator, captures stdout/stderr to a per-job log file, and updates the job store when done.
The log capture is thread-safe: a _ThreadLocalWriter proxy sits on sys.stdout and routes each write to the calling thread's log file. Concurrent jobs write to separate files without interfering.
class _ThreadLocalWriter(io.TextIOBase):
def __init__(self, fallback):
self._fallback = fallback # main-thread output (uvicorn logs etc.)
def write(self, s: str) -> int:
fh = getattr(_tls, "log_fh", None)
if fh is not None:
return fh.write(s)
return self._fallback.write(s) # don't discard main-thread output
The fallback matters: uvicorn's access logs, startup messages, and any main-thread print statements go to the original stdout instead of disappearing.
MCP: the same API as tool calls
The MCP bridge uses fastapi_mcp:
from fastapi_mcp import FastApiMCP
mcp = FastApiMCP(app)
mcp.mount()
This reads the FastAPI app's OpenAPI schema and turns each route into an MCP tool. /mcp lives in the same process, on the same port, with the same authentication, so there is no second API to keep in sync.
From a Copilot CLI ~/.copilot/config.yaml:
mcp_servers:
- name: ai-software-house
url: http://localhost:8765/mcp
headers:
X-API-Key: "your-key"
From that point, Copilot can call submit_run, list_runs, get_run, cancel_run as tool calls. The same operations available via curl are available to any MCP client β without writing a separate MCP server, and without duplicating any route logic.
Authentication
All routes except /health require an X-API-Key header. The key is set once at startup, either from aisw_server.yaml or the AISW_API_KEY environment variable:
api_key = os.environ.get("AISW_API_KEY") or srv.get("api_key", "")
auth_mod.set_api_key(api_key)
The comparison uses hmac.compare_digest to avoid timing attacks:
def _check(api_key: str | None = Security(_api_key_header)):
if not _configured_key:
return # dev mode: no key configured
if not hmac.compare_digest(api_key or "", _configured_key):
raise HTTPException(status_code=401, headers={"WWW-Authenticate": "ApiKey"})
Empty key means open access β useful during local development. Non-empty key requires exact match on every protected request.
Starting the server
python aisw_server.py
With an override:
AISW_API_KEY=secret python aisw_server.py --port 9000
Output:
MCP server mounted at http://0.0.0.0:8765/mcp
AISW server starting on http://0.0.0.0:8765
The full configuration lives in aisw_server.yaml:
server:
host: 0.0.0.0
port: 8765
api_key: "change-me"
defaults:
repo: "owner/default-repo"
pipeline: "ai-feature"
engineers: 2
config_yaml: "config.yaml"
Default repo, pipeline, and engineer count are applied to every submitted job unless the request overrides them. This means the common case β "run the standard feature pipeline on my main repo" β requires nothing more than:
curl -X POST http://localhost:8765/runs \
-H "X-API-Key: your-key" \
-d '{"requirement": "Add dark mode to the settings page"}'
The workflow now
Before the integration layer, using the system from a terminal meant:
- Open a GitHub issue
- Add a label
- Wait for the watcher to poll
- Check back in the GitHub UI
Now it means:
curl -X POST http://localhost:8765/runs \
-H "X-API-Key: your-key" \
-d '{"requirement": "Add dark mode to the settings page"}' \
| jq .run_id
And then watch it run:
curl -N http://localhost:8765/runs/{run_id}/stream -H "X-API-Key: your-key"
Nothing inside the pipeline changed. I only added a quicker way to reach it.
The MCP angle
The MCP side came from a workflow I wanted for myself: while working in Copilot CLI, Claude Code, or OpenCode, I wanted to hand a larger requirement to the pipeline without leaving the session.
With the MCP server running, a Copilot session can do this:
"Submit this requirement to the ai-software-house pipeline: Add rate limiting to the API endpoints."
The assistant calls submit_run as a tool, gets back a run_id, and can check status or stream logs β all within the same session. The pipeline runs in the background. The coding session continues.
It now feels less like a script tied to GitHub and more like a service other tools can use. GitHub labels still work; they just aren't the only option anymore.
Related reading: Pluggable Deploy Backends: Docker, VMs, and Nothing at All