2026-08-28 · Guide
Score Whether This Grok Bot Paid for Itself
You cannot score whether this grok bot paid for itself from a green morning and a subscription receipt. The receipt buys access for every bot on the account. The green morning proves a routine fired. Neither number is hours you avoided. Neither number is overflow you copied off an invoice.
Log the human hours you actually did not spend. Log overflow you actually saw. Cap the evidence at twenty run records, because that is all the product keeps. Do not invent a weekly allowance in dollars. Do not print a savings percentage as a product fact. There is no audit view of Bot actions outside Enterprise, so the sheet is the payback ledger.
Eligible plans include a weekly usage allowance. Past that pool, work is billed on demand from model and token cost (Grok Bot FAQ). No published page prints the allowance as dollars, credits, or runs. There is no Grok Bot-specific spend cap yet (teams and enterprises).
| Question you actually have | Page that answers it | What this score will not do |
|---|---|---|
| Did this bot return hours against overflow I copied | This page | It will not replace a timesheet you refused to keep |
| Why did usage explode | Grok Bot cost | It will not trace frequency, retries, or long documents |
| What do plans cost today | What AI bots cost | It will not reprint the price grid |
| How do I prove what the bot proposed | No audit view outside Enterprise | It will not stand in as an action log |
Stay until one bot has a window of hours versus overflow. Then keep, pause, or rewrite.
Score this grok bot from hours you avoided, never from a savings percentage nobody published
Payback is two quantities. Quantity one is hours you would have spent, minus hours you spent reviewing or rewriting. Quantity two is overflow you copied from your own invoice while this bot ran. If you cannot name both, you have a story, not a score.
Vendors have not published a Grok Bot savings percentage. Roundups that print "bots save forty percent" are not citing docs.x.ai. Do not paste that figure into a budget, a charter, or a prompt. A ratio you compute from your sheet is your arithmetic for this window, labeled as yours. It dies when the next window uses different minutes.
Score one job with one boundary. Inbox Triage never sends. You already know how long a morning of drafts takes. A bot that "handles operations" has no clock. Split it until the hours have one.
All bots share one persistent cloud computer assigned to the user, not to a bot (computer and apps). Overflow follows that grain. Score one bot at a time. Averaging the roster hides the bot that burned the pool while another returned the hours.
Write a verdict with a date, a window length, and the two quantities. Keep and pause are verdicts. "It feels worth it" is not.
Log each morning as minutes you spent versus minutes you would have spent without the bot
Open the sheet the night before morning one. Column A is minutes this job took you last week, timed with a phone clock. Column B is minutes you spent this morning with the bot in the path: reading, correcting, rewriting.
Freeze column A for the window. If last week's mornings were 28, 35, 31, 33, and 30 minutes, the baseline is the median, 31, or the mean if you write which one you used. Maren's 32 minutes later are an arbitrary example she chose. Use yours.
Column B includes waiting, hunting a draft, and explaining a correction the bot will ignore tomorrow. Logging only the final send fakes a win.
Hours avoided equals baseline minutes minus minutes you spent, summed across the window. Zero or negative means the bot did not pay for itself in hours, even if overflow is zero. Review time and rewrite time live in the same column so rewrite cannot hide inside "I used the bot today."
A third column flags whether the output would have been wrong if a human had not caught it. Wrong means you would not have sent, posted, or used it. Tone you would have edited is not automatically wrong. A wrong vendor name is. A date you had not promised is.
Ten mornings fit an inbox drafter and still fit inside twenty run records if you score this week. Declare the window before morning one. Do not stretch it after a bad day.
Count a wrong inbox draft as time you still worked, not as time the bot returned
A draft that would have been wrong is not a partial save. Log nearly the full baseline, plus discovery time, or the sheet lies.
Words are not hours avoided. If you discarded the draft and wrote from scratch, column B sits next to column A. The bot used quota. You used the morning. That row is near zero, or negative if discovery ran long.
Two wrong mornings in ten is not a rounding error. Do not promote it into "Grok Bot is eighty percent accurate." That is a savings-percentage cousin. It describes this window only.
Mail Cleanup Assistant misses cheaper: rejecting a label costs seconds, not a rewrite. Still flag a proposed permanent delete, even when the boundary blocked it. The proposal took review time.
If you cannot tell whether a draft would have been wrong, the row is unscored. Unscored rows do not count toward payback. Guessing "probably fine" is how a sent mistake arrives after you already declared a win.
Copy overflow from your own invoice and refuse every invented allowance figure
Overflow is the on-demand line after the weekly pool. Copy it from the invoice you can open today. Do not take a dollar figure from a thread. No published page prints the weekly allowance as dollars, credits, or runs. Anyone quoting a pool size is guessing.
There is no Grok Bot-specific spend cap, but the account-level On-demand monthly limit applies. Scoring as if a ceiling will fire is scoring a product that does not exist. Zero overflow plus real hours avoided is a keep for this window, not a promise about next week.
If you cannot tell which bot wrote the overflow, log the account line and write allocation unknown. Do not invent "this inbox bot caused sixty percent." There is no audit view of Bot actions outside Enterprise to support a split (teams and enterprises). A keep is weaker when allocation is unknown. A pause is easier to defend.
Overflow as a phase lives on the weekly allowance page. What burns it lives on on-demand usage. If the overflow line is large and the hours are small, pause first, then open those pages.
Paste the amount with a date and a source: "copied from invoice, Friday 22 August 2026." From the phone app (iPhone or Android) you can approve steps, pause or resume a routine, read its run history, and delete it, but not edit or test it (mobile). Do not declare payback from a train seat.
Treat the twenty run records as a cache that cannot hold a payback ledger
A routine assigns a workflow to one Bot. The app keeps the 20 most recent run records per routine. Max 50 routines per Bot. Deleting a Bot deletes its routines. Nothing is team-level (skills, routines and automations).
Those rows show that a run fired. They do not show minutes you spent, whether a draft would have been wrong, or overflow dollars. A clock that fires every hour will push the first row off the list before tomorrow morning. A once-a-weekday job keeps roughly a month of rows, which is why a ten-morning score still fits if you fill the sheet this week and fails if you wait until week five.
Cite the twenty to debug a silent morning. Never cite them as proof this grok bot paid for itself. If you wait three weeks to score ten mornings, the early rows are gone. Fill the sheet the morning of the run.
On the phone you can open that history and pause. Write the score at a desk, on the sheet. There are Linux desktop and Android apps as of September 2026, and the iOS app also runs on iPad (iPadOS 18 or later) (Grok Bot FAQ).
An audit packet proves what the bot proposed and who signed. A payback sheet proves hours versus overflow. Do not merge them into one document that does neither.
Walk Maren through ten weekday inbox mornings where two drafts would have been wrong
Maren ran one inbox drafter for ten weekday mornings from Monday 11 August 2026 through Friday 22 August 2026. Ten was an arbitrary window that still fit inside twenty run records if she scored the same week. She froze column A at 32 minutes, the median of five timed mornings. That baseline is hers, not a product benchmark.
The charter matched Inbox Triage: classify, draft, never send. An approval controls the proposed action. It does not reverse work already completed (approvals, security, and privacy).
Wednesday 20 August 2026 is the dated miss. The bot drafted a thank-you to a vendor it named Northline, for a renewal Northline had not asked for, and offered Friday delivery, a date Maren had not promised. She spent 4 minutes catching the invented renewal and 28 minutes writing the real reply. She logged 32 minutes in column B.
Friday 22 August was the second miss: Reply All on a thread where she was only copied, quoting an internal pricing aside. She spent 29 minutes writing a one-to-one note. That Friday she copied 4.10 USD of overflow from her invoice. That figure is her invoice, an arbitrary example of a copied amount, not the size of the weekly allowance.
| Morning | Minutes without the bot | Minutes she spent | Wrong if sent | Overflow copied |
|---|---|---|---|---|
| Monday 11 August 2026 | 32 | 7 | No | None yet |
| Tuesday 12 August 2026 | 32 | 6 | No | None yet |
| Wednesday 13 August 2026 | 32 | 8 | No | None yet |
| Thursday 14 August 2026 | 32 | 7 | No | None yet |
| Friday 15 August 2026 | 32 | 9 | No | None yet |
| Monday 18 August 2026 | 32 | 7 | No | None yet |
| Tuesday 19 August 2026 | 32 | 8 | No | None yet |
| Wednesday 20 August 2026 | 32 | 32 | Yes | None yet |
| Thursday 21 August 2026 | 32 | 7 | No | None yet |
| Friday 22 August 2026 | 32 | 29 | Yes | 4.10 USD from her invoice |
Baseline across ten mornings: 320 minutes. Minutes she spent: 120. Hours actually avoided: 200 minutes, which is 3 hours 20 minutes. Overflow she actually saw: 4.10 USD, allocation unknown. She did not write a savings percentage. She wrote a keep with a warning: two misses, unknown allocation, send never left. She scheduled the next ten-morning window instead of declaring the bot paid for itself forever.
Chief of Staff Briefing needs a different column A: minutes you spent assembling a pack by hand. Time three packs before you score it. Inventing "this would have taken two hours" after you liked the output is how every briefing looks like payback.
Keep every draft unsent so a failed morning cannot become an outbound cost
A sent wrong draft is an incident this method does not price. Send stays off the bot while you score. Auto-send turns a rewrite row into a customer-facing cost. Stop the score. Handle the incident.
Ask is the gate. Keep the draft rejectable. Standup Scribe posts only to your own DM, never to a shared channel, so a bad note stays off the team while you still measure typing time.
A wrong draft that never sent still spent quota. Maren's Wednesday row returned nothing and still drank the pool. Two of those in a week can fail the score with nobody outside the company seeing a word.
If the job cannot sit on ask (a calendar change that already landed, a purchase already completed), this method is the wrong tool. Approvals do not undo completed work. Pause. Read approval reversibility. Come back only on jobs that still stop in front of you.
Fill the score sheet on the morning of the run, not when the Friday invoice arrives
Friday memory will recall that the bot helped and forget twenty minutes hunting a draft. Fill column B and the wrong-if-sent flag before you leave the desk. Copy overflow when the invoice shows it, with that date.
Skip three mornings and those rows are unscored. Three unscored rows in a ten-morning window fail the window. Do not impute 7 minutes because the other days were 7.
Put the sheet where the company owns it. Maren used one document: window start, window end, baseline rule, overflow source, verdict date. The bot may append that a run finished. It may not fill minutes or grade its own drafts.
If you travel, pause rather than reconstructing after you return. A paused week is an honest gap. A reconstructed week is invented payback.
Answer the claim that the subscription already paid for itself on day one
The strongest objection: you already pay Cursor Pro, Pro+, Ultra, Teams, or a linked SuperGrok, so every bot is incremental-cost-zero, so this grok bot paid for itself the first morning it drafted anything.
That wins on SKU math and loses on overflow math. The subscription is a floor. Overflow is on-demand after the weekly pool. There is no Bot-specific cap. A bot that returns 3 hours and writes a large overflow line can still fail. A bot that takes your full baseline and writes zero overflow returned nothing. The objection skips both measurements.
It also treats the roster as one product. Lead Scout can empty the week in a long browser pass while the inbox drafter looks innocent. "The subscription paid for itself" scores the account, not this bot. Score the scout on a separate sheet.
If you are not on an eligible plan, the objection does not apply. Every paid Cursor plan includes Grok Bot, from Cursor Pro at $20; Cursor Hobby, the free plan, does not, and an individual SuperGrok, SuperGrok Plus, SuperGrok Heavy or X Premium+ subscription can be linked instead. The cheapest individual paid door is Cursor Pro at 20 USD, checked 23 September 2026 on cursor.com/pricing. Upgrade math lives on Is Grok Bot worth it. This page starts after the door is open, or on a timed trial job.
Grant this much: zero overflow, real hours, SKU already paid for other reasons, and the keep is easy. Still write the hours. Still write zero. Still date the window. Next week overflow can appear without a plan change.
Pause this routine after two consecutive failing score windows, not after a feeling
A window fails when hours avoided are at or below zero, when overflow is large with unknown allocation and thin hours, or when unscored rows break the rule you wrote on night zero. Two consecutive failures is the pause trigger. One can be a login wall. Two is a pattern.
Pause from the phone if needed. At a desk, read the last twenty records, coarsen or delete the clock, then rewrite the charter. Operating without a per-bot spend cap is the Friday roster ritual. This pause is one bot, two failed payback windows.
Do not add bots during a failing window to amortize the SKU. Extra bots share the pool and the computer. Cookies, sessions, files, and CLI credentials are shared. Separate bots are not a security boundary. Deleting a bot does not remove shared-computer files or sessions. A new bot will not create a spend cap.
After pause, morning one starts with a fresh baseline. Payback does not bank.
| Symptom | Likely cause | Fix |
|---|---|---|
| Hours avoided near zero, overflow zero | You still rewrite most drafts | Pause. Narrow the charter. Time a new baseline |
| Hours avoided real, overflow large, allocation unknown | Another bot may have burned the pool | Pause clocks you do not score. Do not invent a split |
| Three or more unscored rows | You filled the sheet from memory, or not at all | Fail the window. Do not impute minutes |
| Twenty records gone before you score | You waited too long, or the cadence is too tight | Score each morning. Coarsen the clock |
| Verdict written from iPhone | The phone shows run history, not your minutes | Pause from there. Finish the sheet at a desk |
| Savings percent on a slide | Someone wanted a product fact | Delete the percent. Keep hours and copied overflow |
Paste a payback charter that forbids the bot from declaring it earned its keep
The bot may append that a run finished. It may not compute payback, invent an allowance, or write a percentage.
Name: Inbox morning score clerk
Job: Draft inbox replies for one mailbox, then append a run heartbeat.
Each weekday morning, for this mailbox only:
- Classify unread messages into the four queues I already defined.
- Draft replies only where the draft can quote a sentence from the thread.
- Never send. Never reply. Never forward. Never Reply All.
- Never invent a vendor name, a delivery date, or a price.
After drafts are ready, append one row to SCORE.md in the company folder:
date, run finished at (timestamp), draft count, and the words
"human fills minutes and wrong-if-sent".
Boundary: You never send. You never write hours avoided. You never write
overflow dollars. You never write a savings percentage. You never claim
this bot paid for itself. You never guess the weekly allowance.
Stop when drafts are in the folder and the heartbeat row is appended.
If a thread is ambiguous, skip it and say skipped.
If the bot writes "saved you 40 percent this week," flag the morning even if the drafts were usable. Teach-by-demonstration records up to ten minutes of visible computer work, no microphone audio, draft skill, browser workflows only, unavailable on iPhone. It will not fill column B. You still type the minutes.
Hand bill shape, published prices, and missing audit packets to the twin pages
This page stops when the question is no longer hours versus copied overflow.
Bill shape (frequency, retries, long documents) is Grok Bot cost. Published plan prices and unpublished figures are What AI bots actually cost. Proof of what the bot proposed is Grok Bot has no audit view outside Enterprise. Hours versus overflow will not tell a controller which vendor name was invented on 20 August.
How to stop overspending is the hour you pause everything. Operating without a per-bot spend cap is the Friday roster test. Use them when overflow is the emergency. Come back when the clocks are quiet and you want to know whether the surviving bot still returns hours.
SpaceX acquired xAI (announced 2 February 2026). SpaceX acquired Anysphere/Cursor (closed 14 August 2026). Neither added a Bot-specific spend cap, an allowance dollar figure, or an audit view outside Enterprise.
Stop this score when the job is not hours you can name in advance
A compliance packet is evidence, not hours. A send you cannot undo is an incident. A brief you would never have written by hand has no honest column A. If column A is a guess, time the job by hand three times or drop the score.
Privacy Mode (Legacy) blocks Grok Bot entirely. If it is on, there is no bot to score. Hosted MCP sign-in tokens stay with Cursor's backend. Browser cookies still sit on the one VM every bot shares. Those facts change blast radius. They do not fill column B.
Shipped since the August docs previewed them: a team-level ceiling on local execution for Teams and Enterprise admins, and Terminate, which lets Enterprise organization admins delete a member's computer while the durable disk is kept. Terminate is not a ledger. Do not log admin controls as hours.
When the job is hours you can name: run the window, copy overflow, refuse invented allowances and product-level savings percentages, keep send on ask, pause after two failing windows.
| Verdict | Hours avoided | Overflow copied | Allocation | What you do |
|---|---|---|---|---|
| Keep | Clearly above zero | Zero, or small next to hours | Known or honestly unknown | Keep ask. Start the next window |
| Keep with warning | Above zero | Visible | Unknown | Keep ask. Pause other clocks. Re-score |
| Pause | At or below zero | Any | Any | Pause this routine. Rewrite later |
| Pause | Thin hours | Large | Unknown | Pause. Do not invent a split |
| Fail the window | Unscored rows | Any | Any | Do not verdict. Log next window cleanly |
| Wrong page | Job is not clockable hours | n/a | n/a | Leave this article. Use the twin |
Keep reading: Grok Bot Cost: What You Pay and How Usage Adds Up, What AI Bots Actually Cost, Grok Bot Audit View: Keep Your Own Receipts Outside Enterprise.
Frequently Asked Questions
Can the twenty run records prove this grok bot paid for itself?
No. A routine keeps the twenty most recent run records, then older rows vanish. Those rows show that a run fired. They do not show minutes you spent, minutes you would have spent, whether a draft would have been wrong, or overflow you copied from an invoice. Individual accounts and self-serve Teams still have no audit view of Bot actions; Enterprise has audit logs and Action Recording. Fill the score sheet the same morning. If you wait until the cache has slid, you cannot reconstruct payback from the product, and you should not invent the missing minutes.
How do I score one bot when overflow is billed to the whole account?
Log hours avoided for this bot, then copy the overflow line for the account and write allocation unknown unless you paused every other clock for the window. All bots share one persistent cloud computer and one weekly pool. There is no action log to split the overflow. Do not invent a percentage split. A keep verdict is weaker when allocation is unknown. If overflow is large and hours are thin, pause this routine and the unmeasured clocks before you guess which bot spent the pool.
May I put a savings percentage on a budget slide as a Grok Bot fact?
No. Grok Bot documentation does not publish a savings percentage, and it does not publish a dollar size for the weekly allowance. A ratio you compute from your own hours and your own copied overflow is your arithmetic for that window, labeled as yours, dated, and retired when the next window uses different minutes. Printing it as "Grok Bot saves X percent" turns a personal sheet into a product claim. Keep the slide to hours avoided, overflow copied, misses flagged, and the verdict date.
When should I stop this score and open a different page instead?
Stop when the question is bill shape, published prices, or an audit packet. Frequency, retries, and long documents live on the cost page. Plan prices and unpublished figures live on the price hub. Proof of what the bot proposed lives on the no-audit-view page, because twenty records are not a ledger there either. Also stop when the job has no honest baseline of hours, or when a send already left. Those cases need a different method, not a stretched payback window.