Skip to content
D-CSIL

AI Guide · 2026-10-11 · 11:00 AM CT

Run your AI agents like an engineering team

TL;DR

SpaceXAI engineer Lauren Tan runs 15+ GrokBot agents the way you'd run an engineering org: a Chief of Staff bot, managers, and workers. Her 55-minute talk lays out the whole system — hire bots like roles, teach them skills once, verify everything, eval on a schedule, and write code agents can actually work in. This guide turns the talk and its companion playbook into steps you can follow, with what's official xAI docs and what's her own method marked as you go.

A woman presenting on stage with a microphone at a tech conference (stock photo)
Photo: Vadim Mazko / Pexels

Bots are hires, not prompts

GrokBot is xAI's take on AI teammates: each bot gets a name, a job, and its own persistent cloud computer with a browser, files, and a terminal. It uses real tools the way you would, keeps working after you close your laptop, and coordinates with other bots. There are apps for macOS, Windows, Linux, iOS, and Android, and access is currently bundled with paid plans — the companion playbook puts entry around $200+/month via SuperGrok Heavy or Cursor's Ultra/Teams tiers, with no standalone plan. Treat pricing as the playbook's figure, not an official price page.

The core shift in Tan's talk: a prompt is a request, but a bot is a role. Name bots like jobs — Inbox Manager, Research Lead, Bookkeeper — and give each one mission and one narrow role. Narrow roles stay reliable, are easy to correct, and build sharp memory. Don't start with specialists, though. Start with a Chief of Staff: one coordination bot that gets your docs, inbox, and Slack, audits the operation, and names the three roles that would move the needle first. Prove each hire — run one task through the Chief, review it by hand, then have it build the specialist bot. Later the Chief routes work to specialists and you only talk to the Chief.

The charter is the hire. Write it as a living job-description file with three parts: what the bot owns (acts alone on routine work), what good output looks like (checks, not adjectives), and where it stops. Keep two standing lists: always-OK-alone — anything reversible in under a minute — versus must-park: send, spend, publish, delete. If it can't be undone in a minute, it waits for you.

Teach it once: skills

Do a task once while the bot watches and it saves the flow as a named, repeatable skill. Pick tasks that are recurring, span two or more tools, and have stable steps, then give the skill a trigger — a schedule like every morning at 7, or an event like a matching email. Tan's own example is pstack, a stack-of-diffs workflow skill she says replaced git log in her daily life and earned 100+ upvotes internally at SpaceXAI.

Her rule for a good skill: it's “management philosophy in a bottle” — the training you'd give a human report. Three ingredients: a CLI with verbs (she puts it as “the agent's hands are your hands”), exactly one task per skill (small 20–30-line CLIs the agent chains together), and built-in verification (the CLI prints evidence to stdout and the skill validates it — pstack runs about five verifications before anything uploads).

This part is confirmed in xAI's official docs: skills are reusable instruction sets capturing steps, decision rules, expected output, and safety boundaries, taught by demonstration. The docs add a warning worth heeding — “the learned skill is a draft” — and tell you to add decision rules, failure handling, and approval boundaries yourself. A recorded skill captures clicks, not judgment. There is also a skill marketplace for sharing, and private skills are shared across all your bots.

Trust is a system: verification and evals

The framing question of Tan's whole talk is “how do you trust it?” Her answer is a trust curve: start agents on small, low-stakes tasks and invest more as they earn it. Verification is the mechanism that moves you along the curve. Her story: giving an agent freedom to create GitHub PRs — trust dipped during incidents (a PR deleted and recreated) and recovered through verification loops, not through hoping harder.

Then comes the part that separates her method from the manual: evals. Write tests for your skills and run them before committing, including pre-commit hooks. Build evals from real examples, using the agent's first attempt as the “gold patch.” Run evals on a schedule and don't merge without passing — because agents regress over time, a drift she calls hill climbing. Agents can write the evals; humans review them. Be clear-eyed here: evals are Tan's own methodology. I checked xAI's official GrokBot docs and found no eval framework, gold patches, or scheduled eval runs — the docs advise testing and re-testing skills, but the eval discipline is hers.

One standing rule from the playbook that belongs in every setup: make the bot show you the tape. Require evidence with every result — the source behind every number, guesses listed separately, skips explained. If it can't show how it got a number, the number stays out.

Scale it: fleets and agent-friendly code

Once a skill is trusted locally, scale it: run agents on cloud machines as a fleet. Tan's practice is 20+ agents over a weekend, one agent per workstream, each holding one workstream's context — and she personally reviews every PR, “like being the manager of a team of people.” Bots can hand work to each other directly: put several in one thread, give the group an objective rather than a checklist, and they split it with one owner per stage, pulling you in only for judgment calls. Group bot chats are in the official docs (2–6 bots per shared outcome, with visible handoffs).

The most opinionated part of the talk is about code itself. Don't rewrite brownfield apps for agents, she says — greenfield, especially vibe-coded projects, is the biggest agent opportunity and the biggest risk. “Architecture is not neutral”: a monolith is one big ball of code that agents fumble. Her Dune architecture — code designed for agents — means one agent per component, no dependency-injection frameworks, boring predictable layout, every component independently checkable, one idea per file, clear component boundaries.

Enforce it with CI, not with pleading: automated checks agents can't argue with. Her examples are a 150-line file limit, import restrictions, no DI frameworks, tests must pass, and agents must produce checkable outputs. Roll new teams out in stages — the playbook's cadence is week one drafts only (you read everything), week two you approve each action, week three routine cases with exception escalation, week four scheduled runs where you read the weekly summary. When something fails, fix the charter, the routine, or the handoff — not just the output. And don't interrupt runs: context compounds only if left alone.

The honest limits

None of this is magic, and the playbook is upfront about the edges. GrokBot is early beta — redesigned pages and popups can still trip bots. Entry is around $200+/month with a card required for trial, and there's no standalone plan. All your bots share one cloud computer: shared files, sessions, and logins, which the playbook calls “a real blast radius” — keep banking and sensitive logins off it. Approvals stop an action but don't reverse one, and there's no full audit log yet.

So start where the downside is capped: reversible, low-stakes, multi-tool recurring work. Pick one such task, create the bot, write a four-sentence charter, teach it by doing the task once, walk away, and come back the next morning. If the tape looks good, you've got your first hire — and a template for the other fourteen.