Writing · Field guide · October 2026
Start with one dot.
Build a team around it.
How founders can give one AI assistant real responsibility, and spend less time being the glue between tools.
Part 1
Stop being the glue
You get a useful answer in a research chat. Then you copy it into a task tool, turn it into an engineering brief, and paste a summary into a status update. By the end, you are the one keeping track of which version is current.
We wanted less of that. Not a larger cast of bots on day one, just one assistant with a clear job: prioritize, turn goals into bounded work, review what comes back, and keep useful results where we can find them again. Specialists came later, after that loop felt ordinary.
Handoffs you made0
ChatGPT dot, Slack and Grok Bot are our example stack. We have one dot and six specialists, all coordinating within Slack. That lets us do everything in the cloud.
Part 2
One chief with a real job
Treat one assistant as chief of staff for a narrow slice of your week. Its job:
- Hold this week’s priorities.
- Turn one goal into a clear brief.
- Review the handback against a written definition of done.
- File the useful result in your existing company system.
- Ask you only when the action is consequential or authority is unclear.
When the assistant owns routing and record-keeping inside approved tools, you spend less time restating the same handoff. You still own goals, taste, and decisions that can create real damage.
Put the strongest model in the chief’s chair
Winky, our chief, is a dot on purpose. We wanted frontier intelligence managing the team. Deciding what to do and whether the work is right are the hardest calls in the loop, so that is where the strongest model sits.
- Chief
- Winky, an OpenAI dot. A frontier model. Prioritizes, writes the briefs, reviews every handback, and decides what reaches Greg.
- Specialists
- Six Grok Bots. A substantially weaker model. Fine for bounded lanes with a fixed handback shape, and cheaper to run around the clock.
The stronger model checks the weaker ones’ work, never the other way round.
Start smaller than you think
You do not need our six-agent setup. You need three things:
- One recurring outcome. For example, every Friday, a one-page week in review: top three open priorities, what moved, what is blocked, and one Monday focus.
- One place of record. The task hub you already trust.
- One stop rule. When to ask you before acting.
Prove a loop like that before you design an org chart of agents.
Decide what it may do alone
The stop rule is the most important line you will write. Ours sorts every action into three groups. Pick one to see where it goes.
Chief can do it
Ask Greg first
Never
Pick any action to see where it goes.
Silence is not approval. A previous approval for a different action is not approval. A prompt is instructions, not a grant of app access or a way around product controls.
Part 3
How we run it today
Greg sets goals and approves anything consequential. Winky prioritizes, delegates, reviews and keeps useful results. A private Slack channel carries explicit tasks out and evidence back, and six Grok specialists each cover one bounded lane. Here is a day of it, sped up.
GGregAwake0 waiting
Winky and six Grok specialists work through the day and night. Consequential decisions wait for Greg.
ClickUp owns priorities, tasks and evidence. GitHub owns code state. The Grok desktop app is one interface to a cloud computer; it is not proof that every specialist runs only on Greg’s Mac, or that it works without quotas and permission gates.
Winky runs three active assignments by default, and a fourth only when the work is independent. That is our operating choice, not a platform limit: we started with two and raised it after reviewing the workload. One writer owns each target, and the approval gates stay. Specialists hand back to Winky; they do not invent follow-on work for each other.
The stack, and why each piece
Every tool has one job. Two of them cost money; the rest of the coordination layer runs on free plans.
- ChatGPT dotPaid
- The chief. A frontier model that is always on, works from its own cloud computer and browser, learns from feedback, and can be reached in ChatGPT and in Slack. We pay for the best judgment where judgment matters most. Dots come with ChatGPT Pro or Business Premium, rolling out to eligible accounts; at the time of writing, Pro is $100 or $200 a month and Business Premium $125 a seat.
- Grok BotPaid
- The six specialists. Each Bot runs on a persistent cloud computer, so work continues with the laptop closed. Each gets only the plugins its lane needs, and routines let it wake on a schedule or a narrow Slack keyword. Its weaker model is cheaper to run around the clock. It needs a paid Cursor plan or a linked SuperGrok subscription; each plan has a weekly usage allowance, with optional on-demand billing you can cap.
- ClickUpFree plan
- The record. Priorities, tasks, decisions and evidence links live here, where Greg and every agent read the same thing. A handoff only counts once it is saved here and read back. The Free Forever plan covers a small team: unlimited tasks and members, with caps on storage (100 MB), automations (100 runs a month) and Spaces (five).
- SlackFree plan
- The handoff channel. One private channel,
#agent-ops, where every assignment and handback is visible and searchable. The dot joins as a contact and the Grok Bots post through Slack’s plugin. The free plan shows 90 days of history and allows 10 app integrations, which is one more reason the record lives in ClickUp, not in chat. - GitHubFree plan
- Code state. The source of truth for code. Sync reads it to match engineering evidence; nothing merges without matching approval.
The running cost is two subscriptions, plus whatever usage you allow beyond them. Agents use far more tokens than chat, so set usage caps from day one and treat a hit cap as a signal to narrow the work, not to raise the limit.
The six lanes
- Sync
- GitHub ↔ ClickUp engineering evidence. Stops at merging code or changing labels without matching approval.
- Ledger
- ClickUp HQ and business-development prep. Stops at spending, outreach, and financial claims.
- Scout
- Market and claims research, with facts, inferences and unknowns kept apart. Stops at publishing claims.
- Forge
- Open-source R&D that separates existing requirements from missing proof. Stops at installing or changing code.
- Herald
- DevOps content and visual drafts. Stops at publishing anything.
- Desk
- Feedback and support drafts. Stops at messaging anyone outside the approved path.
What it has genuinely achieved so far: corrected an outdated specification status, updated engineering evidence with acceptance caveats, produced R&D that separated existing requirements from missing proof, and one explicitly approved original DevOps post.
Part 4
Why it works
Dots and Grok Bot launched in September 2026, so there are no studies of them yet. There is solid research on the pattern underneath: several AI agents, one coordinator, and people checking the work. Each reason below is one of our rules, and the finding behind it.
- +90.2%
Put the strongest model in charge.
Anthropic’s research system paired a stronger lead model with lighter subagents, and beat the stronger model working alone by 90.2% on its internal evaluation. That is the case for a frontier-model chief directing weaker specialists.
Anthropic Engineering, 2025 (internal evaluation)
- 4.4×
One coordinator contains errors.
In 180 agent setups, a mistake grew 17.2 times as it passed between independent agents, and 4.4 times with a central coordinator. One chief, not seven bosses, and no bot-to-bot handoffs.
- +81%
Split work in lanes; keep judgment central.
The same study found coordinated agents improved work that splits into parallel parts by 80.9%, and made step-by-step reasoning 39 to 70% worse. Specialists get separable lanes; sequential judgment stays with the chief and Greg.
Same study
- 14
Most failures are design failures.
Across seven multi-agent frameworks and more than 200 tasks, failures fell into three groups: unclear specifications, agents misaligned with each other, and missing verification. Written boundaries, one coordinator, and evidence read back from the hub answer each one.
- Outside
The check has to come from outside.
Language models struggle to correct their own reasoning without outside feedback, and people over-trust automation even with training. So every handback carries evidence someone else reviews, and the approval line is written down.
- −19%
Measure; don’t trust the feeling.
In a randomized trial on 246 real tasks, experienced developers with AI tools took 19% longer, and still believed they had been 20% faster. Run one outcome for two weeks and count the minutes before adding anything.
- 15×
Review time is the real limit.
Multi-agent systems use about 15 times the tokens of a chat. More parallel work than you can verify is just unverified work, which is why active assignments are capped.
Why now: METR’s measurements of how long a task AI agents can finish on their own have been doubling every few months (METR research). Agents that work through the night are becoming practical, which makes clear boundaries more important, not less.
Part 5
Set it up, step by step
Follow the phases in order. Each has the exact actions, how to tell it worked, and what to do if it didn’t. Steps marked You need a person: an assistant cannot buy plans, create accounts or connect apps for you.
Phase 0Decide before you build15 minutes
- Write down one recurring outcome, like the Friday week in review.
- Write down one place of record: the task hub you already trust (ClickUp, Linear, Notion, Asana…).
- Write down one stop rule. Start from the “Ask Greg first” list above.
- Write down the sources it may use, usually the hub plus links you paste in.
✓ Done when all four answers fit on one page and none of them is vague.
If not an answer reads like “help with stuff.” Make it concrete before you continue.
Why Most multi-agent failures start as unclear specifications. These four lines are the specification.
Phase 1Create the dot and run one loop by handDay 1
- You Check that dots are available on your account in OpenAI’s getting-started guide. No dot yet? Use any capable assistant; the pattern is the same.
- You Create a dot by following that guide, and name it. Ours is Winky.
- Paste in the chief brief (below), with the brackets filled from Phase 0.
- Ask it to produce the outcome once, right now.
Show the chief brief
You are my Chief-of-Staff assistant for [company] this week. Single outcome: [the recurring outcome from Phase 0]. Sources you may use: [task hub], plus any links I paste in this chat. Write the result into [task hub / doc] and paste a short summary here. Separate facts from your inferences. List unknowns instead of guessing. You may: research in-scope sources I named, draft the summary, organize notes, and prepare the next task brief for my review when the work stays inside this outcome. You must ask before: publishing, emailing, DMing, installing anything, spending money, changing permissions, or publishing claims about customers, legal commitments, or financial outcomes. Prompts do not grant app access or override product controls. Use only tools and permissions already available to you.
Show what a good handback looks like
[Outcome] — ready for you Facts: 3 open priorities listed from [task hub]; 2 items moved to done; 1 blocked on [named waiting input]. Inference: Monday focus should be unblocking [item] before new work. Unknown: no update yet on [item] — not in the hub. Saved to: [link or task title in your hub] Approval needed from you: none to read; yes before sharing externally.
✓ Done when you get back facts, inferences, unknowns, where it was saved and what needs your approval, filled with real information.
If not it keeps asking questions: tighten the outcome and sources. If it guesses, add “List unknowns instead of guessing.”
Why A single loop is cheap to fix. A team of agents multiplies whatever this loop does, good or bad.
Phase 2Make the hub the recordDay 1
- You Connect only the task hub this outcome needs, using the official connector. Messaging, apps and computers are separate connections (computers and apps).
- Ask the dot to save the result into the hub.
- Open the hub yourself and find the item.
- Ask the dot to read the item back and quote it.
✓ Done when the item is in the hub and the read-back matches it. “Saved” in a chat message is not enough.
If not check the connector’s permissions and retry once. If it still fails, file it by hand this week and note the failure.
Why Agents cannot reliably check their own work. A record you can open and read back is outside evidence.
Phase 3Repeat for two weeks and measureDays 2–14
- Run the same outcome every cycle, for example every Friday.
- Each time, note the minutes you spent routing or restating work, and anything wrong in the handback.
✓ Done when your notes show time saved, not just a feeling.
If not fix the brief and boundaries before adding anything.
Why Developers in a randomized trial felt 20% faster and measured 19% slower. Measure.
Phase 4Optional: put the dot in SlackDay 15+
- You Create a private channel for agent work, for example
#agent-ops. - You Add the dot as a Slack contact, then add it to that channel. The separate @ChatGPT Slack app is not your dot.
- In the channel, tell the dot exactly what to watch and when to act. Adding it does not start monitoring on its own.
✓ Done when you mention the dot in the channel and it replies there.
Why One shared channel puts every assignment and handback in one readable place.
Phase 5Add one specialist and prove one round-tripDay 15+
- You Check your plan on Grok Bot plans and billing. It needs an eligible paid plan.
- You On a Mac: download Grok Bot for Apple silicon or Intel, open the disk image, drag it into Applications and open it. Choose Sign in and finish in the browser.
- You Choose New in the sidebar (or press ⌘N), then Create new Bot. In Edit Profile, set a name, a label (its job in a few words) and a description (the specialist description below).
- You Add only the plugins this lane needs from Settings → Plugins. The Slack plugin posts as the Slack user you connected, so that person owns the posts.
- Have the dot send it one bounded task.
- Optional: set a routine on one narrow Slack keyword in one channel.
Show the specialist description
You are [Name], a specialist for [one lane, e.g. market and claims research]. You take tasks only from [Chief name]. You do not start work for, or hand work to, other agents. If another lane is needed, tell [Chief name] with a complete evidence packet. Return every task in this exact shape: Task ID: [id] | Status: ready for review Result: [one sentence] Evidence: [links or file names] Facts / inferences / unknowns: [three short lists] Checks: pass | fail | not run — [which] Actions taken: [only what actually happened] Not applied: [drafts or proposals] Blocker: [or "none"] Recommended next action + who must approve: Ask before: publishing, outreach, new accounts, installing anything, spending money, changing permissions, or code changes.
Show the bounded task the chief sends
Task ID: [your-id] Owner: [specialist name] Deliverable: [exact artifact] In scope: [sources / systems] Out of scope: publishing, outreach, new accounts, code changes Acceptance: [tests or checklist] Return: one-sentence result, evidence links, facts vs inferences vs unknowns, blockers, recommended next step for human review.
✓ Done when one full round-trip is recorded in the hub: assignment out, bounded work, evidence back, chief review, decision saved. Do it twice before adding another lane.
If not two agents talk past each other: stop and route through the chief. If quotas or approval cards stall the work, narrow the task or wait; never bypass controls.
Why The first handoff is where setups break quietly. Prove the shape once before you copy it. Your chief should be the stronger model; the specialist can be weaker because its lane is narrow and its work is reviewed.
Phase 6Expand only on verified handbacksOngoing
- Add one lane at a time, each with its own narrow tools and fixed handback shape.
- Cap active assignments across the team. We started at 2 and now run 3 by default, after reviewing the workload.
- Give each target one writer: a document, a task, a repository area.
✓ Done when you still read every handback that matters, and the hub is current.
If not review is slipping: lower the cap before adding lanes.
Why Your review time is the real limit. Expanding only on evidence keeps the team as large as your attention, not larger.
When it goes wrong
- It keeps pausing.
- Tighten responsibilities and approval boundaries before adding tools.
- A great reply changed nothing in the tracker.
- Require read-back before calling the loop successful.
- Text looks fine; the page or visual does not.
- Check the surface humans see.
- Quotas or approval cards stall work.
- Narrow the task or wait. Do not bypass controls.
- Two bots talk past each other.
- Stop and route through the chief.
Part 6
Honest limits
We have not built a fully autonomous company, and we are not claiming proven profitability from this setup. Assistants still hit quotas, permission prompts and awkward file transfers. A polished reply can still fail to file. Human approval stays on public, paid, permission-changing or customer-facing actions. Product access and interfaces change, so prefer live official pages over anything frozen here.
Dependable help comes from clear ownership and verified handoffs. Start with one dot, one real outcome, and a task hub as the record. Give that assistant enough authority to remove you from routine routing, and keep yourself on the decisions that matter.