2026-08-29 · Guide
Learn Grok Bot: A Curriculum From First Run to Your Own Boundary Line
Learning Grok Bot starts with the environment it can reach, not with a clever prompt. A useful first run is one you can explain before it begins: what state the bot inherits, which action needs approval, where your written boundary sits, and what evidence will prove the run stayed inside it.
This pillar is the hub for a 15-lesson curriculum. The order matters. You first learn what follows a bot into the shared computer, then separate work surfaces from security boundaries, then place approvals correctly, write a boundary that survives pressure, choose an access path without guessing, and finally manage routines as objects with their own lifecycle. Every product statement here is limited to the safe claims in VERIFIED-FACTS-2026-08-25. Use the lesson links for the full exercises and evidence.
Use this curriculum if you will operate, review, or buy a bot
This curriculum is for the person who will be accountable after the output leaves the screen. You might be setting up your first Grok Bot, reviewing a colleague's setup, approving a team purchase, or writing the operating rules for a recurring workflow. You do not need security credentials or programming experience. You do need access to a practice environment, the patience to inspect current account state, and authority to stop a test when its assumptions fail.
The sequence also fits experienced automation users who are new to agent behavior. A workflow graph often makes connections and branches visible. A bot can receive a goal and work through unfamiliar input, so the operator has to make the stop conditions equally visible. Prior automation experience helps, but it does not replace the isolation and lifecycle lessons.
| Reader | Start here | Practice environment | Completion evidence |
|---|---|---|---|
| First-time operator | Lesson 1 | Synthetic files and disposable logins | Written state inventory |
| Team reviewer | Lesson 1, then Lesson 5 | Named test account | Approval and boundary test |
| Buyer or administrator | Lesson 10 | Current billing and seat records | Eligibility worksheet |
| Routine owner | Lesson 14 | One noncritical practice routine | Recovery inventory |
Do not use a production inbox, customer record, financial account, or publishing surface as your classroom. The curriculum is designed so harmless markers and synthetic records reveal the mechanism without making a real person absorb the cost of a mistake.
Run the lessons as tests, not as reading assignments
Complete the lessons in order and keep one learning log. For each lesson, write a prediction, perform the smallest safe observation, record what happened, and state which operating rule changed. A lesson is complete only when your evidence could prove you wrong. “I understand approvals” is not evidence. “The run stopped before sending and produced the exact proposed recipient and body” is.
Use a clean practice case, but do not mistake clean for imaginary isolation. The product documentation says all bots on an account share one persistent cloud computer, with a separate screen for each bot. Your exercise should therefore inventory the shared state even when every file and login is disposable.
| Step | Your action | Evidence to keep | Failure that teaches you something |
|---|---|---|---|
| Predict | Write expected reach and stop | One paragraph before the run | You cannot state what should happen |
| Observe | Run one narrow synthetic case | Screenshot or state card | Result differs from prediction |
| Explain | Name the mechanism | One sentence tied to verified facts | Explanation depends on an unsupported feature |
| Repair | Change one rule or environment condition | Versioned charter line | Same failure repeats |
| Retest | Reuse the original bad case | Pass or fail against a rubric | New wording cannot be scored |
Spend one session on each lesson rather than racing through all 15. The purpose is not memorization. It is to build a chain of decisions you can reuse when the interface, plan list, or workflow changes.
Finish with five abilities you can demonstrate
At the end, you should be able to perform five concrete tasks without relying on product folklore. First, inventory the cookies, signed-in sessions, files, and command-line credentials that a run may encounter on the shared computer. Second, distinguish a bot screen, an instruction boundary, a product approval, and an actual capability restriction. Third, write a boundary line whose forbidden transition and human handoff can be tested.
Fourth, classify an account as eligible, ineligible, or unresolved by its exact documented plan, while refusing to invent the model behind Grok Bot or a usage allowance figure. Fifth, inventory a routine and preserve enough information to rebuild it before deleting its owning bot.
| Ability | Observable test | Passing result |
|---|---|---|
| Inspect inherited state | Compare two bot screens on one account | Shared state is recorded, not assumed away |
| Place an approval | Draw the action timeline | Gate appears before the consequential transition |
| Write a boundary | Run a tempting edge case | Bot produces a review artifact and stops |
| Verify access and limits | Complete plan and budget worksheets | Unknowns remain labeled unresolved |
| Preserve a routine | Rehearse deletion recovery | Workflow, owner, cadence, inputs, and output are recorded |
These abilities are intentionally operational. They do not certify that every future run is safe. They give you a repeatable way to find unsupported assumptions before those assumptions become external actions.
Correct one misconception at every stage
Every stage has one tempting shortcut. During isolation, the shortcut is believing that a new bot or a different screen creates a clean environment. During approvals, it is treating a yes as retroactive protection or as permanent policy. During boundary writing, it is confusing polite language with a testable stop. During access and limits, it is inferring eligibility, model identity, or a cap from a price or interface label. During routines, it is treating a recurring workflow as a team-owned object that survives its bot.
| Stage | Common misconception | Correct operating view | First corrective action |
|---|---|---|---|
| Isolation | Each bot has its own computer | Bots share one account computer | Inventory shared state |
| Approvals | Approval makes prior work safe | Approval governs the proposed action | Place it before consequence |
| Boundaries | “Be careful” creates a hard stop | A boundary is an instruction with an observable trigger | Name verb, object, substitute, and stop |
| Access and limits | Paid means eligible and capped | Exact tiers differ and no bot-specific cap exists yet | Record plan and usage evidence |
| Routines | The team owns the schedule | A routine belongs to one bot | Record the owning bot |
Treat this table as a diagnostic index. When a test surprises you, return to the stage's misconception before rewriting the prompt. Many failures come from the environment or lifecycle model, not from the wording of the immediate request.
What a Pasted Prompt Inherits the Moment It Runs
A pasted prompt does not arrive in an empty room. On Grok Bot, all bots on an account share one persistent cloud computer. The separate screen helps organize work, but the browser cookies, signed-in sessions, files, and command-line credentials on that computer can be shared across bots. The first lesson therefore changes the first question from “What does this prompt say?” to “What can this prompt reach when it runs here?”
Start with a state card. Record browser identities, relevant local folders, authenticated command-line tools, and any active sessions. Do not copy secrets into the card. Use a harmless marker file to test reachability, and ask for only that exact path rather than a broad search. The goal is to observe one mechanism without exposing unrelated data.
The most common misconception at this stage is that unfamiliar text begins with no context. In practice, the environment supplies context and authority that the text itself may never mention. A cautious prompt review paired with a careless environment review is incomplete. If the state card contains anything inappropriate for the exercise, stop and clean or relocate the work before execution.
Continue with What a Pasted Prompt Inherits the Moment It Runs for the full marker-file exercise and before-and-after state card.
Screens Are Work Surfaces, Not Security Boundaries
The second lesson corrects the visual metaphor. Grok Bot gives each bot its own screen on the account's shared computer. That screen is a work surface, useful for keeping jobs understandable and returning to the right context. It is not evidence of a separate browser profile, filesystem, credential store, or virtual machine. The verified documentation explicitly warns against using separate bots as a security boundary.
Test the distinction with two columns. In the first, list what a screen visibly separates: current work, navigation context, and the operator's organization. In the second, list what the shared computer can retain: sessions, files, cookies, and command-line credentials. Then ask what control would actually be needed for isolation. A different bot name does not satisfy that requirement.
The misconception is easy to hold because visual separation feels like account separation. Many products use tabs, profiles, workspaces, and containers differently. Do not transfer a familiar product's isolation model into this one. Use the documented unit, which is the computer assigned to the user account, and design the environment around the most sensitive bot that can reach it.
Continue with Screens Are Work Surfaces, Not Security Boundaries to practice separating organizational convenience from security evidence.
Where a Bot Cookie Actually Lives and How Long It Stays
A browser cookie belongs to the browser state on the shared account computer, not to the label of the bot that happened to create it. If one bot signs into a service, another bot on the same computer may encounter that signed-in session. The safe facts establish that browser cookies and signed-in sessions are shared across bots. They do not establish a universal lifetime for every cookie, because each service can set and revoke its own session behavior.
Map identity as a chain: service account, browser session, shared computer, bot screen. Record the service name and account identity without recording cookie values. Then move to another bot screen and check whether the same identity is present using a harmless practice account. Your observation should answer whether the session persists in this environment, not promise how long the service will keep it alive.
The misconception here is that closing a screen, ending a task, or switching bots ends authentication. None of those actions is the same as signing out or revoking a session. Cleanup needs an explicit service-level action and a verification pass. If session lifetime matters, verify it with that service's current primary documentation rather than assigning a number from memory.
Continue with Where a Bot Cookie Actually Lives and How Long It Stays for the identity map and session cleanup exercise.
Why Deleting a Bot Leaves the Files and the Sessions
Deleting a bot removes the bot object, but the verified product facts say it does not remove files or browser sessions from the shared computer. This is a lifecycle problem. The bot, its routines, the account computer, local files, and external sessions are related objects with different deletion behavior. One delete action cannot safely be treated as cleanup for all of them.
Before deletion, build an inventory with one row per object. Record the bot's routines, relevant files, browser accounts, and command-line credentials. Decide which items must be preserved, signed out, revoked, or removed. Then perform the deletion and verify each remaining object independently. The lesson is not that deletion is broken. It is that its scope is narrower than a beginner often assumes.
The common misconception is that deleting the visible thing cleans the environment that served it. That belief is especially risky during handoff, incident response, or retirement. Files can remain useful, and sessions can remain powerful, after the bot disappears. Treat preservation and revocation as explicit work with named owners.
Continue with Why Deleting a Bot Leaves the Files and the Sessions for the full lifecycle map and deletion checklist.
What an Approval Actually Governs, and What It Cannot Undo
An approval controls a proposed action. It does not reverse work already completed. Place that fact on a timeline: the bot reads input, interprets it, drafts or prepares a change, proposes a consequential action, and then either stops for approval or crosses the transition. Approval belongs immediately before the transition it governs.
This placement reveals both its value and its limit. Approval can keep a draft from being sent or a proposed change from being applied. It cannot make earlier browsing, file reading, credential use, or drafting unhappen. Those earlier steps need their own scope and environment controls. A late approval is still useful, but it is not a general privacy shield.
| Timeline point | State of the work | What approval can govern | What it cannot change |
|---|---|---|---|
| Before reading | Input is untouched | A later proposed action | State already present on the computer |
| After reading | Information has been accessed | Whether the next consequence occurs | The completed read |
| At proposal | Exact action should be visible | Accept or deny that action | Earlier drafts and observations |
| After action | Consequence has occurred | A future action only | The action already completed |
The misconception is that a visible approval prompt certifies the whole run. Instead, read the exact proposed action, target, account, quantity, and consequence. If the proposal is vague, refuse it and repair the workflow so the next request is reviewable. Record what remains true after denial, including any files already created or information already read.
Continue with What an Approval Actually Governs, and What It Cannot Undo for the action-timeline exercise and denial-state review.
Approval Fatigue and the Blanket Yes That Undoes Your Boundary
Approval quality declines when every prompt looks routine and the reviewer has no compact basis for comparison. The dangerous response is a blanket yes that turns a meaningful decision into ceremony. A boundary can say “never send without approval,” yet repeated, low-information approvals can make that line functionally weak because the human no longer evaluates the send.
Improve the queue rather than asking the reviewer to concentrate harder. Every request should show the exact action, target, reason, evidence, and consequence. Group only proposals that share the same risk and review criteria. Separate unusual destinations, large quantities, missing evidence, and changed account identities. A reviewer should be able to reject one item without accepting or rebuilding the rest.
The misconception at this stage is that more prompts mean more safety. Prompt count measures interruption, not judgment. A smaller number of information-rich gates can preserve the decision better than dozens of repetitive confirmations. Track denials, corrections, and proposals returned for missing context. If every request receives instant approval, test whether the gate is still carrying a real decision.
Continue with Approval Fatigue and the Blanket Yes That Undoes Your Boundary for queue design and blanket-approval failure tests.
How to Write a Boundary Line a Bot Cannot Argue With
A useful boundary line names four things: the trigger, the forbidden action, the allowed substitute, and the stop behavior. “Be careful with customer email” names none of them. “When a reply could leave the draft folder, never send it. Save the recipient, subject, and body for review, ask the named operator, and stop the run” can be observed and tested.
Write the line before writing the method. That order prevents the desired outcome from quietly swallowing the stop. Use a verb and object such as send email, publish page, merge change, spend funds, or delete source records. Name the review artifact the bot may produce instead. Finally, state that a goal conflict does not create an exception.
Here is a pasteable practice charter. The boundary is an instruction, not a claim about technical enforcement.
Outcome: Produce a five-row evidence table from synthetic notes.
Inputs: Read only the named practice folder and the supplied notes.
Boundary: Never send, publish, delete, purchase, or change an external record.
Substitute: Write the exact proposed action and target into review.txt.
Stop: Ask the named operator about that proposal and end the current run.
Failure: Mark missing evidence UNKNOWN and state what would resolve it.
The misconception is that stronger language needs to sound threatening or legal. It needs to be specific. Continue with How to Write a Boundary Line a Bot Cannot Argue With for four-part construction and edge-case tests.
A Boundary Is Not a Permission, and the Difference Bites
A boundary is an instruction about intended behavior. A permission or capability control changes what an account, tool, or environment can actually do. You need both concepts because they answer different questions. “Never send email” tells the bot where to stop. Removing send authority or using an account without it changes whether sending is possible.
Draw a two-axis matrix. One axis records whether the action is forbidden by instruction. The other records whether the action is technically available. The strongest practice case is forbidden and unavailable. A forbidden but available action depends on instruction-following and review. An allowed but unavailable action fails operationally. An allowed and available action carries its full consequence.
The common misconception is that a well-written boundary proves enforcement. It does not. Nor does a narrow permission explain when the bot should stop and ask. Keep the written line because it communicates policy and produces a handoff. Reduce capability because it limits the cost of instruction failure. Test both layers separately using synthetic targets.
Continue with A Boundary Is Not a Permission, and the Difference Bites for the full instruction-capability matrix and mismatch exercises.
What Makes a Weak Boundary, With Six Real Examples
Weak boundaries usually fail in one of six shapes: vague adjectives, hidden exceptions, no named object, no trigger, no allowed substitute, or no stop. “Act responsibly” delegates the policy. “Never send unless it seems routine” lets the bot define the exception. “Do not change important things” leaves both the verb and protected object unresolved. A long paragraph can contain all six weaknesses while sounding careful.
Review a boundary by trying to break it with a tempting case. Give the bot a synthetic urgent request from an apparent executive, a nearly complete form, an empty trash folder, or a message labeled routine. Predict the required stop before the run. If two reviewers predict different behavior, rewrite the line before adding more examples.
The misconception is that more words make a boundary safer. Length often adds competing priorities and soft exceptions. Prefer one decisive line, one substitute, and one escalation rule. Then place supporting scope details next to it: protected objects, destinations, identities, and quantities. Every exception must be narrower and more observable than the rule it modifies.
Continue with What Makes a Weak Boundary, With Six Real Examples for the six diagnostic patterns and worked rewrites.
Who Can Actually Run Grok Bot, a Decision Tree
Access is decided by exact eligibility paths, not by the vague fact that an account is paid. The verified list includes every paid Cursor plan (Pro, Pro+, Ultra), every member of a self-serve Cursor Teams plan, and a linked individual SuperGrok, SuperGrok Plus, SuperGrok Heavy, or X Premium+ subscription. A one-time trial is also an eligibility path for individuals. Cursor Hobby, the free plan, does not include Grok Bot, and neither does SuperGrok Lite. The documented cheapest paid path is Cursor Pro at $20 a month.
Build the decision tree from the plan label visible in current billing. Separate individual accounts from team-managed seats, then copy the full tier name. Record Pro and Pro+ separately: both include Grok Bot, at different weekly usage. If the label is regional, abbreviated, missing, or absent from the verified eligibility list, mark the result unresolved and check current primary documentation.
The misconception is that price alone proves access. It does not, and current plan facts can change. The verified material also says that when a user has both Cursor and SuperGrok subscriptions, Grok Bot uses whichever has more usage. That rule does not publish the size of either allowance.
Continue with Who Can Actually Run Grok Bot, a Decision Tree for the complete eligibility worksheet and unresolved branch.
Why the Model Behind Grok Bot Is Not Published
Grok Bot has no model picker for members or administrators, and the supplied facts say there is no plan to allow user or admin choice. The product uses a fixed model set per surface with automatic failover, and billing follows the actual serving model. The identity of the model behind a particular Grok Bot run is not published in the verified facts.
This means you should not infer the model from the Grok name, a separate model catalog, Grok Build documentation, output style, speed, or a screenshot from another surface. Grok 4.6 is real and powers Grok Build, but that does not establish that Grok Bot runs Grok 4.6. Surface names are part of the evidence.
The misconception is that every product carrying one brand shares one model. Instead of writing a model name into your test record, capture the surface, date, account plan, task, observable output, and any published billing evidence. This lets you compare behavior without turning an inference into a product fact. Also note the documented inconsistency: settings language mentions a default model when selection is available, while team documentation says there is no picker. Do not resolve that tension by inventing a feature.
Continue with Why the Model Behind Grok Bot Is Not Published for the evidence hierarchy and claim-audit exercise.
What You Cannot Cap, and How to Budget Anyway
There is no Grok Bot-specific spend cap yet. Subscriptions include a weekly usage allowance, and overflow is on-demand, billed from model and token cost. The verified facts do not publish a universal allowance amount, so a responsible budget cannot begin with an invented credit figure or a control the product does not provide.
Budget operationally. Record the owning subscription, routine cadence, number of active workflows, input size, and observed on-demand charges. Start with one narrow bot and a low-frequency practice routine. Review usage on a fixed schedule, pause work when the evidence crosses your internal threshold, and require a human decision before adding another recurring job. These are management controls, not claims of an enforced product ceiling.
The misconception is that a budget document or written boundary blocks billing. It does not. Separate observation, policy, and enforcement in your worksheet. If the product later ships a specific cap, verify its scope before relying on it. Until then, design for detection and a named stop owner.
Continue with What You Cannot Cap, and How to Budget Anyway for the monitoring worksheet and overflow response plan.
Which Surface Reads SKILL.md, and Why It Is Not This One
Grok Build and Grok Bot are different product surfaces. The verified facts attribute Claude Code compatibility, automatic reading of skills, plugins, MCPs, agents, hooks, and related instruction files to Grok Build. The Grok Bot documentation does not mention SKILL.md, CLAUDE.md, or Claude Code compatibility. Therefore, do not claim that placing a SKILL.md file on the shared computer configures Grok Bot.
Even on Grok Build, accepted metadata is not automatically enforced. The verified facts say Grok accepts but does not apply SKILL.md fields for model, effort, license, and compatibility, and that allowed-tools neither grants nor restricts tools. A parsed field and an enforced control are different claims.
The misconception is that file compatibility travels across every surface bearing the Grok brand. Always attach the surface to the claim: “Grok Build reads” rather than “Grok reads.” Test configuration behavior in the product that documents it, and keep bot boundaries in the bot's actual operating instructions rather than relying on an unverified file discovery path.
Continue with Which Surface Reads SKILL.md, and Why It Is Not This One for the surface matrix and metadata enforcement test.
What a Routine Is, and Where It Dies With the Bot
A routine assigns a workflow to one bot. The verified facts set a maximum of 50 routines per bot and say the app keeps the 20 most recent run records per routine. Routines are not team-level objects, and deleting a bot deletes its routines. This lifecycle differs from shared-computer files and browser sessions, which can remain after bot deletion.
Inventory every routine with its owning bot, purpose, cadence, inputs, output destination, reviewer, boundary, and last known failure. Preserve the human-readable workflow and a synthetic test case outside the routine before deleting the owner. This does not assume an export feature. It creates enough evidence for a deliberate manual reconstruction.
The common misconception is that recurring means durable or shared. Recurrence describes when work starts, not who owns the saved assignment or what survives deletion. The recent run record window is also not a permanent archive. If history matters, preserve the needed evidence elsewhere under your own retention policy.
On the phone app, the verified facts allow pausing, resuming, approving, reading run history, and deleting a routine, not editing. Editing and testing require desktop. Continue with What a Routine Is, and Where It Dies With the Bot for the lifecycle inventory and recovery rehearsal.
The Five Questions to Answer Before Your First Bot
The final lesson compresses the curriculum into five questions. What exact artifact should exist at the end? Which evidence may support it? Which consequential action must remain human? What structured result should appear when the bot cannot finish? Who reviews the output, and against which visible criteria? Together, the answers form a bot charter.
Use synthetic data and choose one narrow artifact, such as a five-row evidence table. Name allowed files and sources. Write the boundary with a forbidden action, substitute, handoff, and stop. Define useful failure labels such as missing source or conflicting source, making clear that your labels and row counts are exercise choices rather than product limits. Give the reviewer a binary rubric that checks both desired output and forbidden consequences.
The misconception is that a broad goal lets the bot demonstrate more intelligence. It usually makes the lesson impossible to score. A narrow first charter creates a stable baseline. Keep the failed case as a regression test, change one line at a time, and version the reason for each revision.
Continue with The Five Questions to Answer Before Your First Bot to build the complete first-run charter and prediction worksheet.
Continue with one narrow bot and a review cadence
After the last lesson, choose one job that produces a draft or review queue without external consequence. Do not begin with a fleet. A single narrow bot lets you see whether your environment inventory, evidence rules, boundary, approval placement, and usage review work together. Suitable catalog examples include Bookmark Skill Grader, which gives you a bounded review artifact, and Personal CFO, whose catalog boundary keeps financial movement human.
Run the job first with synthetic inputs whose expected answer you already know. Review every run until the output shape and stop behavior become boring. Keep a weekly state card for shared sessions and files, a usage note for the owning subscription, and a routine inventory if you schedule it. Add a second job only when the first has a named owner, a passing regression set, and an incident response that someone other than its author can follow.
Your next learning path should follow the first real weakness you observe. If the boundary breaks, study instruction and capability controls. If the output drifts, strengthen the rubric and evidence schema. If the routine surprises you, rehearse lifecycle recovery. The curriculum has succeeded when a failure sends you to a specific test, not to a larger prompt.
Keep reading: build your first Grok Bot in an hour, then use the bot failure modes field guide to turn the first bad run into a regression test.
Frequently Asked Questions
What should I learn first about Grok Bot?
Learn the isolation model first. All bots on an account share one persistent cloud computer, while each bot gets a separate screen. Cookies, signed-in sessions, files, and command-line credentials can be shared across bots. Before pasting a prompt, inventory that inherited state with harmless practice data. This foundation prevents later lessons about approvals and boundaries from resting on the false assumption that a new bot starts inside a clean, isolated computer.
How long does this Grok Bot curriculum take?
Plan one focused practice session per lesson rather than setting a fixed product rule. A session is complete when you can predict a result, run a narrow synthetic test, save evidence, and explain any mismatch. Fifteen rushed readings will teach less than five completed experiments. The useful measure is not elapsed time. It is whether you can demonstrate inherited-state inspection, approval placement, boundary testing, access verification, budget monitoring, and routine recovery without guessing.
Does a written boundary technically prevent Grok Bot from acting?
No. A boundary is an instruction that states intended behavior, a forbidden transition, and a human handoff. It is not itself a permission control. Pair the written boundary with the narrowest available account and tool capabilities, then test both layers separately. The boundary tells the bot and reviewer where work must stop. Capability restriction limits what can happen if that instruction fails. Neither layer should be described as the other.
Can I use SKILL.md to configure Grok Bot?
The verified Grok Bot documentation does not say that Grok Bot reads SKILL.md. The documented skills and Claude Code compatibility claims apply to Grok Build, a different surface. Do not transfer that behavior to Grok Bot. Put the operating boundary in the bot's actual instructions and verify behavior there. On Grok Build, also remember that some accepted SKILL.md metadata is not applied as enforcement, including allowed-tools, model, effort, license, and compatibility fields.