Building an AI Press Team with Human Editorial Review
Part of the AI Software House series.
In short: I reused the software pipeline to run a small press team. RSS stories become GitHub issues, agents research and edit an article, and the finished Markdown arrives as a PR for a human to approve before publication.
This describes the initial press workflow. The articles on editorial triage and The News Reviewer β A Quality Gate Before the PR cover the checks added afterward.
The use case
Until this point, almost every pipeline was about software delivery: PM, architect, engineers, QA, deployment. I wanted to find out whether the orchestration was genuinely reusable or only looked reusable inside one domain.
News was a useful test. It still needs specialist roles, research, handoffs, review, and a final human decision, but the output is an article rather than a codebase.
The result watches feeds about AI, Linux, open source, and security, then researches and drafts selected stories. Nothing goes live automatically: publication still waits for a human to merge the PR.
The architecture
Three repos. Three responsibilities.
ai-software-house β the engine. Agents, orchestrator, tools, pipeline runner. No press-specific content beyond the agent implementations and pipeline stages.
ai-it-press β the press workspace. Holds the pipeline definition, the incoming articles (as GitHub Issues), and the published articles (as merged PRs). This is where the editorial workflow lives.
hklug-sitegen β the publishing target. A static site generator. When an article PR merges to ai-it-press, a GitHub Actions workflow converts the markdown and pushes it to hklug-sitegen.
I kept these separate because I did not want press-specific decisions leaking into the engine. The workspace can change its editorial process, or the publishing target can be replaced, without changing the orchestrator.
From RSS to GitHub Issue: the watcher
rss_watcher.py runs as a cron job (every 15 minutes). It reads a list of RSS feeds from config, fetches each one with a 30-second timeout, deduplicates against a local SQLite database of seen URLs, and creates a GitHub Issue for each new story.
# config.local.yaml
rss_watcher:
press_repo: wanleung/ai-it-press
label: news-article
max_age_hours: 48
feeds:
- url: https://feeds.feedburner.com/oreilly/radar
source: O'Reilly Radar
- url: https://www.linux.com/feed/
source: Linux.com
- url: https://feeds.feedburner.com/TheHackersNews
source: The Hacker News
The Issue body contains the article URL, title, summary, and source feed. That's the brief the pipeline works from.
One small detail caused a surprisingly practical problem: feedparser is good at parsing feeds, but fetching a URL through it gives me no timeout control. The watcher therefore downloads the bytes with requests.get(timeout=30) and only then hands them to feedparser.parse().
The fetch_url tool
The news writer agent needs to read the source article, not just the RSS summary. fetch_url is a new tool registered in the LocalToolRegistry that agents can call during tool-use rounds.
It fetches a URL, strips HTML tags (removing scripts, nav, footer, head), and returns plain text up to a configurable character limit. Two guards:
SSRF protection. The tool blocks requests to private and loopback addresses before the HTTP call is made β 127.x, 169.254.x, 10.x, 192.168.x, 172.16-31.x, ::1. An LLM-called tool that can reach internal services is a meaningful attack surface.
Download cap. Streaming with iter_content stops at 5 MB. Some URLs redirect to very large files. Without a cap, a single bad URL can exhaust memory.
# fetch_url returns clean plain text from any public URL
fetch_url("https://example.com/article")
# β "Linux kernel 6.9 released today with significant improvements to..."
Two new agents
NewsWriter β the journalist. Given an issue brief (title, URL, summary) and optionally a discussion synthesis from a pre-write research discussion, it calls fetch_url to read the source article, then writes a first draft with a headline, lede, body, and context section.
NewsEditor β the copy editor. Reads the draft, checks for factual alignment with the source, improves structure and clarity, and produces the publication-ready version. Like the NewsWriter, it receives any synthesis from a post-draft discussion.
Both agents are implemented as standard LLM agents with tool registries. The NewsWriter gets fetch_url. Both follow personas defined in roles/news_writer.md and roles/news_editor.md.
The 5-stage pipeline
# ai-it-press/pipelines/news-article.yaml
stages:
- discuss_news_analysis # writer + editor research the story together
- news_writer # drafts the article
- discuss_news_draft # writer + editor critique the draft
- news_editor # edits to publication standard
- news_article_pr # opens a PR with the finished article
Stage 1 β discuss_news_analysis. NewsWriter and NewsEditor discuss the source article before anything is written. What's the actual story? What context is missing? What angle is most relevant to the target audience? The synthesis is injected into the writer's prompt.
Stage 2 β news_writer. The writer drafts the article, informed by the discussion. It can call fetch_url to re-read specific sections of the source.
Stage 3 β discuss_news_draft. Writer and editor critique the draft. Is the lede strong? Does the body cover what was identified in the analysis? Are there factual gaps? The synthesis is injected into the editor's prompt.
Stage 4 β news_editor. The editor revises the draft. It knows exactly what the writer intended and what the pre-edit critique flagged.
Stage 5 β news_article_pr. Creates a branch in ai-it-press, commits the article as articles/YYYYMMDD-{issue_num}-{slug}.md with YAML frontmatter, and opens a PR. The PR description includes the source URL and a link to the original issue.
pipeline_file: β fetching pipeline config at runtime
One new watcher feature made this possible: pipeline_file:.
Previously, pipeline stages were defined in the ai-software-house repo config or hardcoded per-repo. The problem with a press team: the pipeline might need to evolve (add a fact-checker stage, change discussion rounds) without touching the engine repo.
# repos-available/ai-it-press.yaml
repo: wanleung/ai-it-press
pipeline_file: pipelines/news-article.yaml # β fetched from ai-it-press at runtime
label: news-article
When the watcher picks up an issue labelled news-article, it fetches pipelines/news-article.yaml from ai-it-press via the GitHub API, validates the stages, and uses them for this run. The pipeline definition lives with the editorial workspace, not the engine. A press editor can change the pipeline without a deploy.
The publish step
When a news article PR merges to ai-it-press, a GitHub Actions workflow fires:
- Detects new
articles/*.mdfiles in the merge - Runs
scripts/convert_articles.pyto convert YAML frontmatter to hklug-sitegen's.txtformat - Commits the converted file to hklug-sitegen via a PAT with
contents:write
The static site rebuilds automatically. Article goes live.
I kept the PR gate because generated journalism can still contain a wrong detail, a bad source fetch, or an embarrassing headline. The automation prepares the work; a person still decides whether it should be published.
What I learned from it
The experiment convinced me that the reusable part really is the pipeline, not the software-development roles.
The pipeline stages are declarative. The agents are replaceable. The discussion presets are reusable. discuss_news_analysis follows exactly the same protocol as discuss_spec_brief β participants, rounds, synthesis injection β just with different roles and a different prompt context.
I only needed three new pieces of infrastructure: fetch_url, rss_watcher.py, and pipeline_file:. The rest was existing machinery with different roles and prompts.
Everything else β LLM agents, discussion stages, PR creation, GitHub integration β was already there. The press team is a configuration of existing primitives.
What's still manual
Feed curation. The list of RSS feeds in config is hand-picked. There's no automatic feed discovery.
Issue triage. Every new RSS entry creates an issue. There's no relevance filter β a low-quality story gets an issue the same as a major one. A classifier stage (before the writer) could filter by relevance score, but it isn't implemented yet.
SITEGEN_PAT. The publish step requires a GitHub personal access token with write access to hklug-sitegen, added as a secret to ai-it-press. This is a one-time manual setup that can't be automated.
DNS rebinding. The SSRF guard in fetch_url blocks known private IP ranges at URL-parse time, but doesn't re-check after DNS resolution. A controlled DNS server could resolve a public hostname to a private IP after the check passes. Full protection requires resolving the hostname and checking the resulting IP β not implemented, documented as a known limitation.
Related reading: Debate Before You Write: Pre-Spec Discussion