Multilingual by Default — Translating the Press Team to Cantonese and Traditional Chinese

📁 Computer

Part of the AI Software House series.

In short: “Traditional Chinese” was too vague for the audience I had in mind. The press pipeline now produces informal Written Cantonese for Hong Kong and formal Traditional Chinese for Taiwan, using one translator agent with two carefully separated registers.


Why two languages, not one

My first thought was to translate each English article once and label it Traditional Chinese. That would have been technically correct and still wrong for many readers. A Taiwanese broadsheet and a Hong Kong tabloid use the same character set but do not sound alike.

Hong Kong tech professionals constantly switch between English and 廣東話書面語. Written Cantonese has its own particles—係、唔、嘅、喺、咗、咁—its own vocabulary, and a recognisable rhythm. A Mandarin-register translation can use flawless Traditional characters and still feel as though it was written for somebody else.

The zh-tw output serves a different purpose: shareable content that reads naturally to Taiwanese readers and formal broadsheet audiences across the Sinosphere. It uses Taiwan-standard vocabulary — 軟體 not 软件, 網路 not 网络, 影片 not 视频 — and no Cantonese colloquialisms at all.

These aren't stylistic preferences. They're different registers that require different knowledge to produce correctly.


One agent class, two stages

The code is the easy part: one TranslatorAgent, called with a target_language parameter.

class TranslatorAgent(BaseAgent):
    role_name = "translator"

    def run(
        self,
        article: str,
        target_language: Literal["cantonese", "traditional_chinese"],
        reviewer_notes: str = "",
    ) -> dict:
        label = _LANGUAGE_LABELS.get(target_language, target_language)
        prompt = (
            f"Translate the following news article to {label}.\n\n"
            f"Follow your role instructions exactly.\n\n"
            f"<ARTICLE>\n{article}\n</ARTICLE>\n\n"
            f"Output the translated article only."
        )
        translated_article = self.call(prompt)
        return {"translated_article": translated_article}

The _LANGUAGE_LABELS dict maps the internal key to the full description the LLM sees in the prompt:

_LANGUAGE_LABELS = {
    "cantonese": "Written Cantonese (zh-hk)",
    "traditional_chinese": "Formal Traditional Chinese (zh-tw)",
}

This keeps the orchestrator side clean. Two pipeline stages, same function, different parameter:

"translate_cantonese": PipelineStage(
    fn=lambda r: self._stage_translate(r, "cantonese", "article_zh_hk"),
),
"translate_zh_traditional": PipelineStage(
    fn=lambda r: self._stage_translate(r, "traditional_chinese", "article_zh_tw"),
),

The stages run sequentially after news_editor and before news_article_pr. Keeping them as separate checkpointable stages means a failed Cantonese translation doesn't prevent the Traditional Chinese stage from running, and a retry of one stage doesn't require re-running the other.


What the role prompt needs to do

The role file (roles/translator.md) carries the linguistic knowledge the LLM prompt alone can't carry. The critical parts:

For cantonese: the instruction to use vocabulary and particles like 係、唔、咁、嘅、喺、而家 — these are the linguistic fingerprints of Written Cantonese. Without explicit instruction, most LLMs default to a Mandarin-influenced formal register even when asked to write "Cantonese".

For traditional_chinese: the instruction to follow Taiwan/HK broadsheet style, citing 台灣蘋果日報 and 香港明報 as reference points, with no colloquialisms and formal written Chinese register throughout.

Both sections share a set of absolute rules: translate the YAML frontmatter (title, tags) and the article body, preserve all markdown structure, keep source_url and author unchanged, and output only the translated article — no meta-commentary or agent remarks.

I added the last rule after seeing translations arrive with “Here is the translation:” at the top or an explanatory note at the bottom. Without the rule, that helpful commentary would be committed as part of the article.


The three-file PR

The PR stage (_stage_news_article_pr) was updated to commit all three files when translations are present:

extra_files = {}
if result.article_zh_hk.strip():
    extra_files[filename.replace(".md", ".zh-hk.md")] = result.article_zh_hk
if result.article_zh_tw.strip():
    extra_files[filename.replace(".md", ".zh-tw.md")] = result.article_zh_tw
result.all_files = {filename: article, **extra_files}

A single PR review now shows the human reviewer the English original alongside both translations. They can check whether the Cantonese version reads naturally or whether the Traditional Chinese version used any Mainland vocabulary before approving. One merge, three files published simultaneously.

Translation failure is non-fatal. If one stage times out or returns an empty response, the PR contains whichever versions completed. I would rather review an English-only article than lose the whole run because one translation call failed.


What goes wrong in practice

Simplified characters bleeding in. An LLM generating Traditional Chinese will occasionally emit a Simplified character — particularly for compound words where the Simplified form is more common in its training data. 国 instead of 國, 软 instead of 軟, 网 instead of 網. These are exactly the errors the news reviewer stage (added later) was built to catch.

Mandarin patterns in the Cantonese output. The zh-hk translation looks superficially correct — Traditional characters — but reads like translated Putonghua rather than Cantonese. 現在 instead of 而家, 很多 instead of 好多, 這個 instead of 呢個. The vocabulary particles are absent. This happens when the LLM doesn't have strong enough Cantonese training signal and the role prompt isn't specific enough to pull it out of the Mandarin-writing mode.

Taiwan vs HK vocabulary in zh-tw. 軟體 (Taiwan) vs 軟件 (HK) is the canonical example, but 網路/網絡, 影片/視頻, 程式/程序 all have regional variants. The role prompt asks for Taiwan-standard vocabulary. It doesn't always get it right.

Frontmatter translation inconsistency. The LLM sometimes leaves English tags in the YAML front matter while translating the body. The role prompt addresses this with an explicit "Translate EVERYTHING" instruction at the top.

The model choice matters here too. The translator is configured to use opencode/opencode-go/qwen3.5-plus with an ollama/thinker fallback — Qwen has significantly stronger CJK language capability than typical English-dominant models, which shows up most clearly in the Cantonese register quality.


The pipeline as it stands

After this change, news-article.yaml reads:

stages:
  - discuss_news_analysis
  - news_writer
  - discuss_news_draft
  - news_editor
  - translate_cantonese
  - translate_zh_traditional
  - news_article_pr

A human still checks whether the translations actually sound right before merging. Later I added an automated reviewer for the more mechanical errors, but I do not treat it as a replacement for a native reader.


Related reading: Building an AI Press Team with Human Editorial Review