# How to Run a Blog Content Engine with Coding Agents

How I run this blog with coding agents: MDX in Git, deterministic validation, human review, measured publishing.

## Give agents an operating system, not just a prompt

> Don't ask an AI agent to become your content team. Build it an operating system.

An agent is only as reliable as the environment around it. Prompts are temporary context. Repository files, schemas, validators, and Git history are persistent controls. So this site runs its blog like code: content lives in versioned MDX, voice rules live in docs agents must read, deterministic checks reject what humans should never have to review twice, and nothing publishes without a pull request carrying evidence.

This is the system currently running this site's content workflow — not a hypothetical architecture. This post explains that operating model end to end, using the real implementation on pantaleone.net as the example. The architecture is generic enough to reproduce on any MDX-based blog.

```text
CONTENT ENGINE AT A GLANCE

Strategy → Inventory → Issue Queue → Coding Agent → Research + Draft
  → Deterministic QA → PR + Evidence → Preview → Human Review
  → Deploy → Measure → Refresh
```

The agent is not trusted to decide everything. It operates inside a system where persistent files provide context, deterministic code provides enforcement, and Git provides accountability. That sentence is the shortest possible explanation of this entire article.

## The Content Engine in 30 Seconds

Content is code. Every article on this site is an MDX file under `features/blog/content/`, reviewed and deployed through the same Git workflow as application code. There is no CMS detour, no draft lost in a chat thread, no publish button that skips review.

Agents operate against persistent rules, not prompts. Strategy, voice, content inventory, and linking rules live in files the agent reads at the start of every task. When the rules change, the files change, and every future draft follows the new version automatically.

Validation is automated and publishing is gated. Scripts check metadata, titles, prohibited phrases, and prose quality before a human looks at the draft. The pull request is the publishing gate: it must carry editorial results, quality-gate results, a preview URL, and the internal links used. Merge deploys through Vercel.

## Why "Just Give the Agent a Prompt" Fails

A long prompt feels like control. It is not. The context window resets, the agent improvises around whatever the prompt forgot, and two drafts written a week apart follow two different standards. Six failure modes repeat everywhere I have seen this tried.

Temporary context is the first. A prompt carries instructions for one session. It does not persist, version, or diff. When output quality drifts, there is nothing to inspect except the prompt you already overwrote.

Inconsistent voice follows. Without a written voice spec with budgets and banned phrases, each draft picks its own register. One reads like documentation, the next like a vendor landing page.

Unsupported claims are the expensive one. Agents invent customers, metrics, rankings, and credentials when the draft needs authority the research did not supply. A prompt that says "be accurate" does not catch a fabricated proof point. A claim filter plus a human gate does.

Then there is no deterministic enforcement. Style guidance written in prose ("keep titles short") is interpreted differently every run. A script that fails the build on a stuffed title is interpreted the same way every run.

No publishing accountability comes next. If the agent can publish directly, there is no record of what was checked, by what, and when. A PR that requires evidence creates that record for free.

And there is no feedback loop. Publishing without measurement means the same weak pages sit untouched while new drafts repeat their mistakes. Search performance has to flow back into updates and relinking, or output grows while authority does not.

The fix for all six is the same: move knowledge out of the prompt and into the repository.

## An AI Writer Is Not a Content Engine

People hear "AI content system" and picture a prompt that outputs articles. That is the weakest version of the idea. It helps to name the ladder:

**AI writer.** Prompt → article. No memory, no standards, no record. Every draft starts from zero.

**Content workflow.** Research → prompt → article → review → publish. Better, but the standards still live in people's heads and the review catches whatever it happens to catch.

**Content engine.** Strategy → research → decision → content → validation → deployment → measurement → learning. Standards live in files, checks run deterministically, and every publish leaves a record. Output improves because the system remembers.

**Agentic content engine.** All of the above, with coding agents able to operate inside the actual repository, tooling, issue queue, validators, Git workflow, and deployment process. The agent does not describe the system from outside. It works inside it: reading the same files, running the same commands, opening the same PRs.

Each rung removes a category of human babysitting. Most teams arguing about AI content quality are stuck on rung one, debating prompts, when their problem is the absence of rungs two through four.

## The Architecture

The full loop, before any implementation detail:

```text
Strategy
   ↓
Topic Queue
   ↓
Agent Orchestration
   ↓
Draft
   ↓
Editorial + Deterministic Validation
   ↓
Minimum Effective Corrections
   ↓
Git / PR
   ↓
Preview
   ↓
Publish
   ↓
Measure
   ↓
Refresh / Relink
```

Each layer has one job. Strategy decides what is worth writing. The queue decides what comes next. The agent drafts. Validators enforce what can be checked mechanically. Humans judge what cannot. Git records all of it. Measurement restarts the loop.

The rest of this post walks each layer, then maps it to real files.

The important upgrade to that diagram: the engine is a closed loop, not a pipeline. A pipeline ends at publish. A loop feeds every stage back into decisions:

```text
INPUTS        strategy, search demand, existing content, inventory,
              first-party knowledge, analytics, Search Console,
              customer questions, product priorities
   ↓
DECISION      what deserves to exist, what gets refreshed,
              what gets consolidated, what gets retired
   ↓
PRODUCTION    research, brief, drafting, examples, internal linking
   ↓
VALIDATION    editorial, factual, SEO, structural, technical,
              build, link integrity
   ↓
DEPLOYMENT    branch, PR, preview, human review, merge, production
   ↓
FEEDBACK      search performance, AI-search visibility, engagement,
              conversions, decay, internal-link performance
   ↓
NEXT DECISION (back to the top with more evidence than last time)
```

Publishing is the midpoint of the loop, not the end of it. That is what makes this an engine instead of an automated publishing pipeline.

## The Repository Is the Agent's Operating Environment

A coding agent with repository access reads files more reliably than it follows chat instructions. Files are complete, ordered, and re-readable mid-task. Chat instructions are paraphrased from memory halfway through a long run. That asymmetry is the whole argument: put everything the agent must obey where it already looks.

Here is the structure this site uses. Every path below exists in the repository.

```text
docs/CONTENT_ENGINE.md      strategy + orchestration rules, read before any content task
docs/EDITORIAL.md           voice, budgets, title rules, claim policy
config/content/             pillars, inventory, backlog reference
features/blog/content/      *.mdx source of truth, one file per post
config/seo/                 search config, kept out of visible prose
modules/anti-slop/          vendored quality-control engine (L1–L4 scanner, quality gate)
config/anti-slop/           project overrides: protected terms the scanner must not touch
scripts/editorial-audit.py  `npm run editorial`, fails on blacklist and stuffed titles
```

Each layer earns its place. `CONTENT_ENGINE.md` is the entry point: one paragraph states the rule (fewer pages, stronger relationships), then a file map, pillar list, workflow, linking rules, metadata checklist, and quality gate. `EDITORIAL.md` owns voice: direct, technical, calm, with word budgets per surface and a banned-phrase list. `config/content/` owns structure: what exists, what is planned, what links to what. The MDX files own the words. The validators own enforcement.

If you reproduce this, keep the same separation. One file per concern, each short enough that an agent reads all of it in one pass.

Two honest notes on this implementation. First, pillars here are conceptual, not routing: URLs stay flat at `/blog/[slug]` and frontmatter `pillar` is advisory until hub pages ship. Do not rename URLs to match taxonomy. Second, the related-content scorer (`features/blog/lib/content-graph.ts`, rendered by `RelatedArticles.tsx`) exists as a component, but the post template currently renders chronological prev/next navigation, a `ConsultationCTA`, and featured products. Inline contextual links in the body therefore carry internal linking today; the scoring engine is the next wiring step, not a current one.

## Building the Editorial Control Plane

The control plane is everything the agent consults before writing a word. It has five parts.

Strategy first. Name a small set of pillars and refuse to add one without merging or retiring another. This site runs six: `ai-agents`, `ai-engineering`, `mcp`, `ai-automation`, `enterprise-ai`, `lab`. Each pillar has entry queries it must answer. Every post gets exactly one primary intent: informational, technical, comparison, commercial, strategic, experimental, news, or navigational. One intent per post forces the draft to serve one reader and one query.

Voice second. Write the voice down as budgets and bans, not adjectives about tone. This site's rules fit on one page: excerpts run 12–25 words, headlines 1–8, any paragraph holds at most two adjectives. Banned openers and title stuffers live in the voice doc as explicit lists:

```text
banned openers: Ultimate, Complete, Definitive, "Everything You Need to Know"
title stuffers: Free, Modern, Official, Production-Ready
```

Stack details belong in metadata, not prose. Concrete constraints beat paragraphs of tone guidance because a script can check them.

Inventory third. Keep a typed audit of every post: slug, pillar, class, intent, cluster, next action. In this repo that is `config/content/inventory.ts`, covering all 30 posts with verdicts like `core-authority`, `supporting`, `field-note`, `merge-candidate`. The agent looks up a post before touching it and learns whether to keep, refresh, merge, or leave it alone. Without an inventory, agents edit blind and duplicate what exists.

The inventory answers sharper questions than a spreadsheet ever could: what exists, why it exists, which query and intent it serves, which cluster it belongs to, what supports it, whether it is authoritative or redundant, whether it needs refreshing, and what should happen next. Combined with the content graph and search performance, plus business priorities, it becomes an editorial decision system: content inventory + content graph + search performance + business priorities = what to write, refresh, consolidate, or retire.

Priorities fourth. A frozen title backlog (`config/content/backlog.ts`, P0 through P3) works as a seed reference, but live work enters through GitHub Issues labeled `content`, one issue per topic. Issues beat a static list for agent execution: explicit ownership, status (`queued` → `in-progress` → `in-review` → `done`), ordering by priority then age, stale-lock reversion, and a direct branch-to-PR relationship. An orchestrator picks the next queued issue on a schedule; the claim comment carries the target branch name.

Metadata and linking fifth. Each post carries a frontmatter contract: specific title, 12–25 word description, controlled category and tags, `seo` keywords, image plus alt text, canonical `/blog/[slug]`, and BlogPosting JSON-LD with `datePublished`, `dateModified`, and a Person author. For example, the frontmatter `seo` array on a systems post names the search concepts (`content engine coding agents`, `editorial validation pipeline`) so visible prose never has to. Search terms live in metadata and structured data, never stuffed into titles or prose. SEO terminology should still appear naturally where it helps the reader; metadata reinforces relevance, it does not compensate for weak prose.

Internal linking is a graph problem, not a writing task. It should never be "agent, find three related posts." It should be topic relationship → cluster relationship → reader journey → authority relationship → deterministic recommendation → human review. The rule here: same cluster first, then same pillar, then shared tags, capped at four related plus one continue-reading link. No FAQ schema unless a genuinely useful visible FAQ exists. The goal is not maximum internal links. The goal is useful relationships.

## Building the Validation Layer

Separate what code can check from what humans must judge, then automate everything in the first group. This repo runs five commands, all real:

```bash
npm run editorial                 # blacklist phrases, superlatives, stuffed titles, budgets
npm run content:validate          # editorial + anti-slop audit over blog, projects, shop
npm run anti-slop:scan -- <files> # report only, never rewrites, exits 0
npm run anti-slop:strict -- <files> # fail on any hard gate (factual integrity)
npm run check-types               # tsc --noEmit, for content-adjacent code changes
```

Deterministic checks enforce metadata completeness, schema shape, required fields, prohibited patterns, title constraints, protected terminology, structural rules, and link validity. The editorial auditor fails on marketing-default phrases, unsupported superlatives, stuffed titles, and banned openers across frontmatter, bodies, and chrome copy. It skips code fences, tables, blockquotes, and lines containing URLs, so quoted third-party text and code samples are never flagged as site voice.

The anti-slop engine (`modules/anti-slop/`, vendored from `@forwardos/anti-slop`, provenance pinned in `modules/anti-slop/PROVENANCE.md`) is a quality-control layer, not a good-writing detector. It runs an L1–L4 scan, explains each finding, applies the minimum effective edit capped at two passes, and rechecks against a 13-item quality gate. Describe it that way to your agents: it catches concrete defects, it does not confer taste.

That distinction deserves names, because teams constantly collapse them. Deterministic quality is machine-checkable: metadata present, patterns absent, links valid, schema correct. Editorial quality is judgment: clarity, specificity, voice, whether the examples earn their place. Strategic quality is business judgment: does this page deserve to exist, who is it for, what does it change. The scanner owns the first category. Agents assist the second under human review. Humans alone own the third. The moment a team expects the scanner to determine whether writing is good, they have mistaken a defect detector for an editor — and the article quietly becomes an advertisement for an AI detector instead of a system for publishing.

The minimum-effective-edit rule matters more than it sounds. The goal is not to rewrite every article into one artificial voice. The goal is to identify a concrete problem, change only what is necessary, and rerun validation. Good content is left alone. That restraint is what keeps validated output sounding human instead of homogenized.

Respect the exemptions explicitly: code fences, frontmatter, URLs, quoted material, and `config/anti-slop/protected-terms.json` (product names, brand terms, and SEO terms like MCP, n8n, RAG, llms.txt) are never flagged or rewritten. Taste overrides go in the engine's feedback store, never as one-off exceptions in prose.

Treat results in three tiers. Hard failures block the PR: factual-integrity violations, missing required metadata, banned title patterns. Warnings get fixed or justified in the PR body. Everything else is editorial review, which belongs to a human reader doing one plain-English pass: shorter, concrete nouns, facts over claims, delete when in doubt.

## Wiring Coding Agents Into the Workflow

The agent in this system is not the writer. It is the orchestrator. Given a claimed issue, it reads the strategy docs, inspects the inventory for neighbors and duplicates, researches when the topic needs it, drafts the MDX, checks internal links, runs every validator, corrects failures with minimum edits, creates the branch, opens the PR, and attaches the evidence. That is a twelve-step job, and all twelve steps are specified in files, not in a prompt.

The same agent patterns used for production systems apply here. If you orchestrate [AI agents with n8n and LangChain](/blog/building-ai-agent-workflows-n8n-langchain), or design [agent workflows](/blog/ai-agent-workflows) with retries and observability, run content the same way: separate orchestration from reasoning, log every step, retry corrections, alert on hard failures.

MCP access to the validators is an optional acceleration layer. This repo ships a zero-dependency stdio server:

```json
{ "mcpServers": { "anti-slop": {
  "command": "node",
  "args": ["<repo>/modules/anti-slop/dist/mcp/server.js"]
} } }
```

Rebuild first (`npm run anti-slop:build`; `dist/` is gitignored). Available tools include `anti_slop_scan`, `anti_slop_audit`, `anti_slop_rewrite`, and `anti_slop_quality_gate`, alongside design, SEO, code, and comparison tools. MCP is convenience, not control: it gives the agent callable access to deterministic capabilities, but the underlying CLI validators remain the source of truth. A run that passed only through MCP and never through `npm run content:validate` has not passed.

Practical setup notes from this implementation: run agents against a local stack ([private AI stack setup](/blog/private-ai-stack-setup-in-minutes) covers the pattern) so drafts, scans, and rebuilds stay fast and private. Keep tool permissions narrow: the agent needs file read/write in content paths, shell access for the five commands, and Git for branching. It does not need production credentials or direct publish rights. Ever.

## The Git-Based Publishing Loop

All publishing goes through Git. The sequence:

```text
Issue → claim → branch → draft → validation → PR → preview → review → merge → deployment
```

Concretely: work happens on `content/<issue>-<slug>`. The PR references `Closes #<issue>` and carries an evidence section: editorial result, anti-slop result, type-check result when code changed, preview URL, and the internal links used. Missing evidence keeps the issue open under a `needs-evidence` label. On merge, Vercel builds from GitHub and the orchestrator closes the issue with the evidence bundle. Content changes require a redeploy since the sitemap builds at deploy time. Verify on the preview deployment, on desktop and mobile, before merging.

`created` never changes. `lastUpdated` moves only on substantive edits. Old URLs are never silently renamed; if hub pages ship later, migrations go through 301s with canonical updates.

Why this matters for agents: the PR is the one place where machine checks and human judgment meet with a durable record. The agent cannot bypass validation because CI runs it independently. The reviewer cannot wave through slop because the evidence section is empty. Git history shows what changed, the PR shows why, and rollback is one revert. That is what publishing accountability looks like, and it costs nothing beyond discipline.

## What the Agent Should Control vs. What Code Should Control

This distinction is the thesis in table form:

| Agent (interpretation)               | Code / System (enforcement)            |
| ------------------------------------ | -------------------------------------- |
| Interpret the topic and angle        | Validate frontmatter schema            |
| Research when evidence is thin       | Enforce required metadata              |
| Draft prose and examples             | Detect prohibited phrases and patterns |
| Explain trade-offs and failure modes | Validate links and canonical URLs      |
| Choose useful code samples           | Enforce title and budget constraints   |
| Propose minimum effective edits      | Preserve Git history and PR evidence   |
| Summarize validation for review      | Block merges on hard-gate failures     |

**Agent judgment handles interpretation. Code handles enforcement. Git handles accountability.**

Anything appearing in the right column that you currently leave to agent judgment is a defect in your engine, not in your agent. Move it into code. Anything in the left column that you try to encode as a rule will produce stilted output; leave it to the agent and review it as a human.

The operational version of that rule is even shorter: if a requirement can be expressed deterministically, do not ask the agent to remember it. Two examples from this repo's history:

BAD: "Please remember to keep titles short."

GOOD: a validator rejects titles outside the defined constraint, every run, with the same message.

BAD: "Please remember to never invent metrics."

GOOD: claims requiring evidence enter a validation state and cannot pass the factual-integrity gate without evidence or explicit human disposition.

Prompts ask. Code enforces. When the two disagree about what happened, the gate wins and the discrepancy gets investigated — not the other way around.

## Agentic Does Not Mean Everything Is an Agent

The least credible sentence in AI consulting right now is "we'll solve it with multi-agent everything." Agents are one tool in this engine, used where interpretation is the actual work:

CODE handles deterministic validation, parsing, formatting, metadata checks, link checking, schema validation, and deployment. Code is cheap, exact, and the same every run.

AGENTS handle interpretation, research, synthesis, editorial judgment, prioritization, and explaining trade-offs. Agents are for the parts where the answer is not computable in advance.

HUMANS handle strategy, factual accountability, high-impact claims, brand judgment, and final editorial approval. Humans are for the parts where someone has to be responsible.

That division is the reason this system stays trustworthy as it scales. Every task migrates down the ladder the moment it becomes deterministic: today's agent judgment becomes tomorrow's script, and the script never forgets, never improvises, and never has an off day.

## What I Would Not Automate

The same philosophy draws a hard line. I would not fully automate strategic positioning, original opinions, important factual claims, customer promises, legal claims, sensitive topics, major consolidation or deletion decisions, or the final judgment about whether a piece of content deserves to exist.

The principle: automation should remove mechanical work, not responsibility. A system designed to manufacture pages at scale — no QA, no accountability, traffic as the only goal — is exactly what search engines now penalize as scaled content abuse ([Google's spam policies](https://developers.google.com/search/docs/essentials/spam-policies) say this plainly). This engine is built for the opposite direction: fewer pages, stronger relationships, every publish defensible. The machinery exists to increase quality and capacity, not to manufacture volume.

## Content Deserves Software-Grade Quality Controls

If content is a software artifact, it deserves software-grade quality controls. The mapping is nearly one to one:

```text
SOFTWARE                  CONTENT
tests                     metadata validation
CI                        editorial + anti-slop gates on every change
type checking             frontmatter schema checks
linting                   prohibited-pattern detection
deployment gates          PR evidence requirements + preview review
rollback                  Git revert
```

Nobody serious ships software by pasting code into production and hoping. Content should not ship that way either. Tests do not make software correct; they make defects visible early and fixes cheap. Validators do the same for prose: the editorial pass is the linter, the preview deployment is staging, the PR is the code review, and the revert is the rollback. Teams already understand this discipline. They only need to notice it applies to words too.

## Claim Provenance: Every Important Claim Needs a Path

The most load-bearing addition to the system above is claim provenance. Agents invent statistics, customers, rankings, performance numbers, credentials, and results the moment a draft needs authority the research did not supply. The fix is structural: every important claim carries a path back to its evidence.

For externally sourced claims: claim → source → evidence → article. Name the source, keep the evidence inspectable, link it.

For first-party claims: claim → implementation → repository or measurement → article. The proof is a file, a command output, or a recorded observation — something a reviewer can open.

For opinions: claim → clearly identified as opinion or experience. No laundering of guesses into facts through confident tone.

This connects directly to the validation tiers. A claim without a path is not a draft problem; it is a gate problem. It enters a validation state and cannot pass the factual-integrity gate without evidence or explicit human disposition. Missing proof gets an em dash and a label, never an invented number. Drafts that fail originality or expertise get deleted, not patched. That is how the engine stays defensible: not by trusting agents more, but by making fabrication structurally impossible to publish.

## First-Party Experience Is the Moat

AI can summarize existing information. It cannot manufacture genuine first-party experience. That asymmetry is the entire commercial argument for this kind of system — and for the content it produces.

The engine therefore prioritizes actual implementation over summaries of other people's implementations: original observations, failures, trade-offs, architecture, measurements, and lessons learned. Screenshots of real output beat stock diagrams. A named file path beats a paragraph describing an approach. A candid limitation beats a confident generality.

This is also what keeps the content on the right side of search guidance. Google's [helpful-content guidance](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) rewards first-hand expertise and content that satisfies the reader's goal. No prompt engineers that into existence. You either built the thing or you did not — and readers, human or machine, can tell the difference between a system description with file paths and a paraphrase of ten similar articles.

## SEO Still Matters. AI Search Changes the Surface.

Traditional SEO remains the foundation. AI search adds retrieval and presentation surfaces on top of it. The practical objective therefore does not change as much as the tooling vendors claim:

Crawlable content. Clear information architecture. Strong entities. Original information. Useful structure. Authoritative sources. Clear authorship. Descriptive headings. Internal relationships. Structured data where appropriate. Accessible content. Strong user experience.

There is no secret ranking algorithm for AI search to chase, and this article will not pretend otherwise. Google's own guidance keeps returning to the same point: ordinary SEO fundamentals plus valuable, non-commodity content. The tactics that promise guaranteed AI citations, citation manipulation, or special markup that forces visibility are the new keyword stuffing — same game, new vocabulary.

Two honest extensions follow from that position. First, a modern site should be understandable to four audiences, not two: humans, search crawlers, AI retrieval systems, and browser agents. Agent-driven browsing interacts through rendered pages, DOM structure, and accessibility information, which reinforces everything already listed above: semantic HTML, accessible controls, clear navigation, stable URLs, descriptive links, structured information, machine-readable metadata, useful rendered content. No exotic markup required — just the fundamentals, done properly.

Second, this site exposes `/llms.txt`. Treat it for what it is: an optional machine-oriented documentation layer that may help systems understand a site's important resources — one file among many signals, not an SEO requirement. It does not directly improve rankings. The stronger argument was always the underlying information architecture, not the existence of the file. ([Why this site has one anyway](/blog/llms-txt-for-ai-agent-discovery-and-optimization).)

## The Content Lifecycle

The loop in this article ends at refresh and relink. Stated fully, the lifecycle is:

```text
CREATE → PUBLISH → MEASURE → REFRESH → CONSOLIDATE → REDIRECT → RETIRE
```

A mature content engine does not maximize page count. It maximizes the usefulness and authority of the portfolio. Some posts earn refreshes. Some get consolidated into stronger neighbors. Some get 301s. Some get retired outright. Every one of those is a decision the inventory, the content graph, and search performance inform together.

This is the operational meaning of "fewer pages, stronger relationships." Growth in published URL count is not the metric. Growth in per-page authority and per-reader usefulness is. An engine that cannot delete is not an engine; it is an accumulator.

## Measuring the Engine

"Measure" needs teeth. The feedback stage should track traditional search first — impressions, clicks, queries, positions per URL in Search Console — because that data is mature, per-page, and directly tied to refresh decisions.

AI search measurement layers on top, where available: generative-AI performance reporting as Search Console expands into AI surfaces, AI citations and mentions where measurable, branded search movement, referral traffic from AI products, assisted conversions, and conversions. Google is folding AI Mode data into overall Search Console reporting and expanding coverage into multimodal surfaces such as Lens and image discovery — follow the [Search Central blog](https://developers.google.com/search/blog) for the current state rather than building dashboards on frozen assumptions.

The honest target is a measurement surface that eventually covers TEXT + AI + IMAGE + VIDEO + SOCIAL and distribution. Do not overbuild it today. Instrument what decides the next action — refresh, consolidate, redirect, retire — and let the numbers schedule the work. If you cannot name the metric that would trigger a rewrite, you are publishing into the void. Start from [measuring AI agent ROI](/blog/measuring-ai-agent-roi) thinking: define the outcome, instrument it, then let the numbers schedule the refresh.

## What Didn't Work

Nine failures, observed in the wild and in this repo's own history. Presented as failures, because the architecture did not emerge perfectly:

Trusting one giant prompt. It drifts, it cannot be diffed, and no validator can check it. Files first, prompts reference files.

Letting agents invent proof. Every claim must answer: can we prove this? Missing proof gets an em dash and a label, never an invented number. Drafts that fail originality or expertise get deleted, not patched.

Allowing agents to bypass validation. CI must run the same commands the agent runs. If the agent's local pass and the gate disagree, the gate wins and the discrepancy gets investigated.

Hand-picking related content when deterministic linking exists. Use the scorer (`getRelatedPosts`, max four; `getContinueReading` for the next step) instead of whichever posts the agent remembers. This repo's scorer is built; wiring it into the template is the remaining step.

Auto-merging near-duplicates without human review. This inventory carries two merge candidates right now. Merging changes URLs, canonicals, and reader expectations. Humans decide.

Editing content directly in production. There is no CMS shortcut here for a reason. Every change is a branch, a PR, and a preview.

Stuffing technology names into titles. Stack goes in `coreStack` and `techStacks` metadata and JSON-LD, not in the headline. Titles name the subject and the angle.

Treating anti-slop detection as a substitute for editorial judgment. The scanner finds defects. It does not know whether the article is worth existing. The quality gate's commercial-fit and intent-match items need a human.

Publishing without a measurable feedback loop. If you cannot name the metric that would trigger a rewrite, you are not running an engine. You are publishing into the void. Start from [measuring AI agent ROI](/blog/measuring-ai-agent-roi) thinking: define the outcome, instrument it, then let the numbers schedule the refresh.

## The Real Advantage

The advantage is not generating more words. Word count was never the constraint; judgment was. The advantage is a repeatable system where agents move quickly inside constraints that persist, quality is inspectable instead of vibes-based, every change is reversible, publishing is accountable, and the system compounds: each measurement cycle makes the inventory sharper and the rules more precise.

A technical founder can build a simplified version in a weekend: one strategy doc, one voice doc with budgets, a typed inventory of existing posts, an issue queue, two scripts (one for prohibited patterns, one for metadata), and a PR template requiring evidence. Add the layered scanner when the basics hold. Wire MCP when agents complain about shell friction. Each step pays for itself before the next one costs anything.

What that architecture proves matters for who reads it next. If your team already runs AI workflows that work but still depend on long prompts, manual QA, and tribal knowledge, you do not need a better prompt. You need the operating system around your agents: persistent rules, deterministic checks, agent orchestration, and measurable outcomes. That is the service behind this article, and the article itself is the evidence it can be built.

An agent with an operating system beats a better prompt every time. The competitive advantage is not that an agent can write an article — everyone will have that. The advantage is building a system in which agents operate against persistent strategy, real evidence, deterministic constraints, version control, human judgment, and measurable outcomes. Build the system first. The drafts take care of themselves.

## How AI Participates in This Workflow

Because this article is about AI-assisted content operations, the participation deserves a plain statement. Coding agents research topics, draft MDX, run validators, correct failures, and prepare pull requests with evidence sections. Humans set strategy, judge factual claims, approve voice and positioning, and merge. Nothing publishes without human review and a passing validation record.

That division is stated here for the same reason every claim in this system carries provenance: readers — and retrieval systems — should be able to tell what the machinery did and where the responsibility sits. AI assistance in research and structuring is compatible with quality when accuracy, originality, and user value are enforced and a human remains accountable for publication. No apology attached; no credit dodged.

## Build and Execution Support

I build systems like this with my team: engine setup, rule porting, validator wiring, and ongoing content operations on your stack. The [MCP server pattern](/blog/mcp-ai-server-for-highquality-ai) and the agent workflow behind it transfer directly.

This fits when content is strategically important, changes frequently, lives in a repository, and has to connect to business outcomes — and when a technical owner exists to hold the strategy. It does not fit when you publish a handful of articles a year, have no technical owner, or have no reason to automate. Saying so is part of the same honesty the engine enforces: automate the mechanical work, keep the responsibility.

Bring one content workflow to a 30-minute call. We scope the engine and leave your agents with checks that hold.


Last updated on October 7, 2026