writing /

The Factory — Cloud Agents, Cursor, Grok, and the Gates That Keep You Honest

In my post on running a fleet, the whole point was to stop babysitting a single chat and start dispatching many. That fleet still lived on my laptop: tmux panes, worktrees, a loop watching CI. This post is the next leap. A factory — cloud agents that implement while I'm not looking, and gates that keep the product honest.

Cursor's own numbers tell you why the leap matters. Lines added per PR (p75) are up about 2.5× year over year. Mega-PRs of a thousand lines are a growing share of merges. Mean tool calls per session keep climbing. Human review capacity did not scale with any of that. If you only speed up writing, you ship plausible garbage faster.

The AgentCore post was about giving a product a loop and hands. This one is about giving engineering the same thing — and then refusing to merge on the implementer's say-so.

The factory loop

A factory is not "an agent that codes." It is a loop with a job at each station:

  1. Plan — a human-approved approach, written down, before anyone writes files
  2. Cloud implement — an isolated VM clones the repo, edits, tests, and opens a PR
  3. Local review — while the session is still warm, before GitHub becomes the conversation
  4. PR review — the team contract: shared rules, shared history, a shared merge gate
  5. Safety and security — permissions, secrets, auth, data paths; a separate track from "does it compile"
  6. Quality — tests, artifacts, the thing you would be ashamed to ship if it were red
  7. Merge or hold — low-risk copy/config can auto-approve; everything else waits for a person

The rule that makes the loop real: the implementer never grades its own homework. A second model, a second harness, or a required check has to be able to say no.

Nobody ships this as one named product. Cursor publishes Cloud Agents, Bugbot, Security Agents, Approval Agents. xAI publishes Grok 4.6 and Grok Build. You compose them. That's the playbook.

The stack that earns a seat

I don't need twelve tools. I need three seats, filled on purpose.

The implementer is a Cursor Cloud Agent. It runs in an isolated VM with a full environment — cloned repos, dependencies, secrets, network — then plans, edits, runs tests, and can push a branch or open a PR. Kick it off from the desktop Cloud dropdown, cursor.com/agents, iOS, Slack, a @cursor comment on GitHub or Bitbucket, Linear, or the API. It used to be called Background Agents. The rename is the point: this is not a local session you left running. It is a worker with a machine.

The local/CLI twin is Grok Build. Plan mode before writes. Parallel subagents and worktrees when the job fans out. Headless -p when you want the same agent inside a script or CI. AGENTS.md, hooks, skills, MCP — it picks up the repo's conventions instead of asking you to paste them into a chat.

The model is Grok 4.6, because it is the long-running coding worker that actually sits in both of those harnesses. It was trained to stay with a job across many steps — research, codebase, idea-to-working-first-version — and the self-reported High numbers (DeepSWE v1.1 65.9%, CursorBench v3.2 69.9%) are why it earns the implementer seat. They are not why it earns the merge.

The rest of the 2026 landscape is the same shape wearing different logos. Claude Code (local, Anthropic cloud, or a browser remote-controlling your machine). OpenAI Codex in a cloud container that checks out the repo and loops until a PR exists. GitHub Copilot's cloud agent on an ephemeral Actions VM. Devin on a saved Linux snapshot. If you already live in one of those, the factory loop still applies. The running example here is Cursor plus Grok because that's the pair I can dispatch from a ticket and review from a terminal without switching religions.

Cloud Agents in practice

Use a Cloud Agent when a local session is the wrong computer. Overnight PRs. Three tickets in parallel. A change that spans frontend and backend repos. A bug you filed from your phone. The agent does not need your laptop online, and it can build, click, and screenshot inside its own VM.

What you must set up, or the agent is a smart autocomplete with nowhere to stand:

  • An environment — agent-led setup, a saved snapshot, or a Dockerfile pointed at from .cursor/environment.json. Cursor's line is the right one: not giving the agent an environment is like not giving an engineer a computer.
  • Secrets in the Cloud Agents dashboard, injected at start. Do not bake .env into the snapshot and hope.
  • Network allowlists. Outbound should be the package registry, the docs, the APIs the task needs — not the open internet plus prod.
  • Hooks in .cursor/hooks.json in the repo. Formatters, audit scripts, policy checks. User-level ~/.cursor/hooks.json does not follow the VM home directory, because there isn't one.

What you must not let them touch unsupervised: production credentials, unrestricted egress, and merge. Artifacts and remote desktop are how you see the product. They are not a substitute for a gate.

Model routing, not a god model

A factory that uses one model for everything is a student grading their own exam.

  • Implement with Grok 4.6 in the Cloud Agent. Long jobs, messy codebases, "turn this spec into a working slice."
  • Review with a different frontier model — Claude Fable 5 / Opus 5, or GPT-5.3-Codex — so the second pass does not share the first pass's blind spots. Grok Build in read-only plan mode is the CLI form of the same idea.
  • Triage and dedup with a cheap Flash-class model. Cursor already uses that pattern to collapse semantically duplicate multi-agent findings. Do not spend frontier tokens restating the diff.

Leaderboards are harness- and effort-dependent. DeepSWE, CursorBench, SWE-bench, Terminal-Bench — useful as a scoreboard, useless as a religion. Pick the model for the job, then make the job survive a model that is trying to refute it.

The four gates

This is the heart of it. Speed without gates is just a faster way to be wrong. Keep the tracks separate so a green test suite cannot launder a security hole, and a security bot cannot pretend to be product quality.

Review

AI review only helps when the reviewer has repo context, tests, and rules — not a bot staring at a hunk and nagging about names. Cursor's path, which matches how I want the factory to feel, is local then PR then fix:

After the agent finishes, while the thread still knows why it did what it did:

/agent-review
/review-bugbot
/review-security

Then Bugbot on the GitHub, GitLab, or Bitbucket PR. Put team invariants in .cursor/BUGBOT.md. Measure resolution rate — at merge time, was the flagged issue actually fixed? — not comment volume. Cursor published Bugbot moving from ~52% resolution to about 80% by May 2026, across millions of PRs. If resolution falls while comments rise, the bot is generating noise and people will start ignoring it.

Autofix may spawn another Cloud Agent to repair a finding. That repair is a new change. It goes through the same gates. It does not get a friendship bracelet.

A second-model pass belongs here too, from a harness that is not the one that just wrote the diff:

grok -p "Review this branch for bugs, security issues, and behavior that contradicts the plan. Return only confirmed findings." --permission-mode plan --sandbox read-only

Read-only. Plan mode. No writes. If you script it, fail closed: exit 1 unless the review is clean.

Safety

Safety is who the agent is allowed to be, not whether the code is pretty.

Cursor Auto-review (allow / ask / deny) on shell, plugins, computer use, automation writes, and Cloud Agent launches is a convenience layer. The docs are blunt: the classifier is not a security boundary. It can allow a call a human would have blocked. Pair it with deterministic allow/deny, hooks, and a sandbox, or you are role-playing security.

Grok Build's order is the one I want in CI: PreToolUse hooks, then deny / ask / allow. Deny always wins — including under --always-approve. Plan mode is the human gate before writes. --sandbox read-only is the posture for exploration and review. Secrets stay in the dashboard, not in the snapshot.

If a step cannot name the permission it needs, it does not run unattended.

Security

Security is a different track from review. Bugbot will catch some of it. That is not the same as a security program.

On the PR: Cursor's Security Reviewer. On the repo at rest: the Vulnerability Scanner. In CI: the same Grok read-only review, pointed at auth, secrets, and data paths. Approval Agents may auto-approve only when those findings are clear. Cursor's own published pipeline keeps quality and security on different tracks for a reason — a canary deploy at the end is still a production safety gate, not a vibe.

A human still owns anything that touches authentication, authorization, or production data. Agents draft. They do not bless.

Quality

CI tests fail closed. That is the floor, not the product.

Cloud Agents can attach screenshots, videos, and logs, and you can take the remote desktop and click the thing yourself. If the change is UI and nobody looked at an artifact, you did not review the product. You reviewed a diff.

Public evals are the scoreboard, not the shipping checklist. SWE-bench asks whether an agent resolved a real GitHub issue so the failing tests pass and the passing tests stay green. SWE-bench Pro stretches that into longer, multi-file, enterprise-shaped work. Terminal-Bench asks whether the agent can drive a real environment, not just emit a patch. Cybench and CyberSecEval measure cyber skill, which is not the same as "did we ship a nice product." Use them to pick a model. Then write the five checks that would make you ashamed if they went red — checkout, auth, the money path, the empty state, the mobile layout — and put those in CI.

A copyable loop

One feature, spec to merge, with the automations named:

  1. Plan in Grok Build plan mode (or Cursor plan). Approve it. Comment on steps. Do not skip this because the ticket "is small."
  2. Dispatch a Cursor Cloud Agent on a dedicated branch. Pick Grok 4.6. Let it test in its VM and open the PR. @cursor on the issue is enough of a kickoff.
  3. Local review on the resulting diff: /agent-review, then /review-bugbot and /review-security, before you even open GitHub.
  4. PR contract: Bugbot + Security Reviewer as required checks. Invariants in .cursor/BUGBOT.md. Bugbot's GitHub check can sit at neutral and never block merge unless you make unresolved issues fail the build — so make them fail the build.
  5. Second-model review, headless and read-only, with the grok -p command above. Different model than the implementer. Fail closed.
  6. Quality: CI green, plus the agent's screenshots. If it is UI, look at it. A green suite is not a product.
  7. Merge policy: Approval Agents may auto-approve copy and config with zero Bugbot/security findings. Everything else waits for a human. Autofix Cloud Agents may push repair commits. They may not merge.

That is a week of work you can steal. The first time you run it on a real ticket, the local fleet from the last post starts to feel like the workshop, and this starts to feel like the floor.

Failure modes

The factory fails in predictable ways. Name them so you can see them coming.

  • Rubber-stamp review. Same model, same session, "looks good." If the implementer is the reviewer, you do not have a gate.
  • Eval theater. Chasing SWE-bench while your checkout flow is untested. Public benches pick models. Product checks ship products.
  • Agents rewriting the gates. An Autofix agent that "fixes" a failing security check by deleting the check. Hooks and required status checks exist so this is a conflict, not a merge.
  • Mega-PRs nobody can read. The 2.5× line-count problem. Split the work. Cloud Agents are cheap to run in parallel; thousand-line dumps are expensive to trust.
  • Classifier as boundary. Auto-review said allow, so it must be safe. It isn't. Deny rules and sandboxes are the boundary.
  • Neutral checks. Bugbot commented, the check is neutral, GitHub is green, you merged a known bug. Fail closed or admit you are collecting comments for sport.

Where to start

You don't get a factory by buying a seat. You get it by putting one real ticket through the loop:

  1. This week: one Cloud Agent on a ticket you would have done locally, and Bugbot on that repo. Feel what it's like to come back to a PR instead of a chat.
  2. Next: the read-only Grok review in CI, fail closed. Different model than the one that wrote the diff.
  3. Then: Approval Agents, but only for copy and config, and only with zero findings.

The autocomplete question was "what can this model do for my current line of code?" The fleet question was "what are the six things that could be running right now?" The factory question is "what is running without me, and which gate would I be ashamed to skip?"

That last one is the whole job. Learn to ask it and the agents stop being a faster keyboard. They become a floor you can actually ship from.