2026-08-25 · Guide

What AI Bots Actually Cost

Two questions hide inside "what does an AI bot cost", and mixing them is why almost every answer you find is useless. The price is what a vendor charges for access: published, verifiable in one click, and liable to change without warning. The cost is what your setup consumes once it is running, which nobody publishes because it depends on choices you have not made yet.

The price is a lookup. The cost is a design decision, and most of it is decided by a dropdown you click in four seconds.

This page is the hub for both. It prints the prices that are genuinely published, names the figures that are not, and then spends the rest of its length on the half you control.

On this page

Separate the price you pay from the cost you cause

The price is what a vendor charges to let you in. One number on a page, and you should read it there today rather than in any article, this one included.

The cost is what your bots consume once they run. It is published by nobody, for nobody, because it is a function of how often your bots wake, how much they read when they do, how many tool calls they make, and how often they retry a step that will never work. Those are your settings.

Nearly every page answering "how much does Grok Bot cost" takes the first question, usually with a stale number, and stops. That is the wrong half. The price costs you one lookup. The cost shape is the difference between a roster you forget about and a roster you supervise, which was the whole reason you built one.

Here is the access picture, read from the vendors' own pricing pages on 25 August 2026. Treat the date as part of the fact: eligibility widened on 21 August 2026, which made every article written before that week wrong, and it will happen again.

PlanPrice as publishedIncludes Grok BotWorth knowing
Cursor HobbyFreeNoNot an access path
Cursor Pro20 USD a monthNoThe tier people assume works
Cursor Pro+60 USD a monthYesCheapest paid path for an individual
Cursor Ultra200 USD a monthYesIncludes it, at a very different price
Cursor Teams Standard40 USD per user a monthYesCheapest per-seat route
Cursor Teams Premium120 USD per user a monthYesIncludes it
SuperGrok30 USD a monthNoThe x.ai tier that does not
SuperGrok Plus100 USD a monthYesListed as including Grok Bot access
SuperGrok HeavyNot publishedYesEligible, and we will not guess

Sources: cursor.com/pricing, x.ai/pricing, and the Grok Bot FAQ for the eligibility list.

Three readings matter more than the numbers. The cheapest paid route is Cursor Pro+ at 60 USD a month, or Teams Standard at 40 USD per user, so anyone quoting an entry price of 120, 200 or 300 is describing the world before 21 August 2026. A one-time trial for individuals is cheaper still and left out of most round-ups. And holding both a Cursor and a SuperGrok subscription does not stack into one larger pool: Grok Bot uses whichever has more usage available.

Cursor also lists an India-only Start plan at Rs 649 a month, which is not on the documented list of plans that include Grok Bot, so do not buy it expecting access. The access question in full is in why Grok Bot needs a Cursor account.

Three numbers you will see quoted that nobody has published

A page that invents one figure has invented others, so knowing which numbers do not exist is a fast way to grade a source.

The figureWhere you meet itActual statusWhat you can say instead
Size of the weekly allowanceRound-ups, "real cost" threadsNot published in dollars, credits, or runsIt is weekly, and its size is undisclosed
A SuperGrok Heavy priceComparison tablesNot published on any primary pageEligible, price unlisted
A per-run or per-token rateCost calculatorsNot published; overflow bills from model and token costMeasure your own over three days
The model Grok Bot runsModel comparison postsNot published, and there is no pickerA fixed set per surface, with failover

The first row does the most damage, because an allowance figure makes a whole article feel authoritative. It is published nowhere: not in dollars, not in credits, not in runs. Any number you have read was invented, and the rest of that page deserves the same suspicion. We will not guess at it either, which is why nothing here is stated in money except the prices above.

The fourth row has a consequence people miss. The documentation says plainly that Grok Bot has no model picker, for members or admins, that a choice is not planned, and that billing follows whichever model served the request. So the standard lever from every other agent stack, running a cheaper model for the boring jobs, is not available. Three controls remain: how often it runs, how much it reads, and how much it writes. The documentation trail is in the no spend cap guide.

Treat the subscription as a floor and the charter as the ceiling

Your bill has three layers and only one of them is a price.

LayerWhat sets itWhat happens as usage growsWhat stops it
The subscriptionYour plan, at a published priceNothing, it is flatYou, when you change plans
The included weekly allowanceThe plan, at an undisclosed sizeIt is consumed earlier each weekThe clock, weekly
On-demand overflowModel and token cost of what actually ranIt scales with whatever your bots didNothing in the product

The third row is documented and blunt: there is no Grok Bot specific spend cap yet. No setting stops a runaway on your behalf. The only ceiling in the system is the one you write into each charter, which is why half the sections below end in a clause rather than a tip.

That structure explains a confusing experience people report. Under the allowance, halving a bot's schedule changes your invoice by exactly nothing. Over it, the same edit is the entire saving. Your bill is a step rather than a slope, and most people meet cost tuning in the week they cross, which is the worst possible moment to start learning it.

Read the bill as a product, because two levers multiply and the rest add

Usage is not a list of costs, it is a product: runs, multiplied by the work per run, inflated by the retry rate. Treating the inputs as a flat list is why people tune the wrong one, prune a memory file, and change nothing.

LeverEffect on the average billEffect on the worst caseBounded without a clause?
Run frequencySets it, linearlySets it, linearlyYes, the schedule is the bound
Retry rateSmall, most runs succeedUnbounded, one wall can run all weekendNo, the only true no
Material read per runLarge, and it drifts upwardLarge, one document can dominate a monthNo, what it fetched decides
Tool calls per runModerate, grows as the bot improvesModerateNo, "a few sources" has no number
Browser stepsLarge against an API routeLarge and highly variableWeakly
Output lengthSmall, then carried as contextSmallYes, if you state a length

Read the middle column, because nobody does. A lever with a modest average and an unbounded worst case is a risk rather than an expense, and risks get bounded rather than tuned. That reorders the work: set the clock and the retry ceiling first, because those are the multipliers, then prune what each run reads, because that edit now lands on a run count you chose deliberately.

The five drivers one at a time are in the Grok Bot cost breakdown, and the steady-state version for a roster already running is in keeping bot costs predictable.

Every five minutes is the most expensive phrase in bot setup

A five minute schedule is 288 runs a day against one. Nothing about that arithmetic is surprising, and the schedule gets picked anyway, because at the moment you click it the cost is invisible and responsiveness feels free.

The jobWhat five minutes buys over dailyWho acts at 03:00The honest cadence
Competitor pricingA price you will not act on until MondayNobodyOnce a day
Inbox triageA shorter queue, at 288x the runsNobodyTwice on weekdays
Deploy or incident watchReal minutes, if someone is on callWhoever is pagedAn event trigger, never a poll
Social mentionsA faster reply you were not writingNobodyTwice a day, or on event
Long document summariesRepeat payment for a fixed answerNobodyOnce per item
Weekly KPI reportingA number moving less than its noiseNobodyWeekly, or monthly

One row has a defensible case for a tight clock, and even there the answer is an event trigger rather than a poll, because a poll costs the same on a dead Sunday as during a launch.

The question that settles every scheduling argument is not how fresh you would like the data. It is how long the answer stays true. You are not repricing at 3am, which is why Competitor Pricing Watch reads public pages daily, and you are not answering mail while asleep, which is why Inbox Triage runs twice a day. Event triggers are cheapest on average and least predictable by nature, so pair each with a per-day cap.

Choosing between the two is covered in schedules versus event triggers, and the settings in the Grok Bot scheduling guide.

The retry loop is the only cost with no natural end

Every other lever has a per-run maximum you can work out on paper. A retry loop does not, and three properties make it the classic runaway rather than a mild overspend.

It is unbounded, because nothing in the loop is aware of cost.

It is silent. A bot meeting a login wall does not throw an exception. It observes a screen, forms a theory, acts on it, and reports that it is still working. That is the signature: progress without output, and by the time you recognise it the weekend has gone.

It is polite, which is why people leave it running. A bot saying it will try another approach sounds like diligence.

Approvals do not save you, and the documentation says why: an approval controls the proposed action and does not reverse work already completed. Everything done to reach the prompt is done and charged. Approvals protect you from the action. Only a ceiling protects you from the approach.

Two clauses stop it and you need both. A hard attempt limit, two tries at any single step and then stop. And a ban on alternate routes to the same result, because without it an attempt limit means two attempts per route, and routes are unlimited: a mirror, a cached copy, a search result, a help article, a different sign-in URL. Flight Check-In is built this way and stops for a human at every 2FA prompt or captcha rather than trying to get past one.

The loop written out attempt by attempt, including the turn where a retry ceiling stops helping, is in the no spend cap guide, and the same failure as one of seven recurring modes is in the seven ways bot setups fail. A subscription pruner meeting a new device check is the canonical version.

Browser work is what separates an agent bill from an API bill

Grok Bot operates a persistent cloud computer, with each bot getting its own screen on that shared machine. Reading a number out of a web dashboard means loading a page, waiting, observing what rendered, scrolling, clicking, and sometimes observing again because a panel arrived late. Reading the same number from an API means one request and one small response. Identical answer, very different work, and far more variance.

So the design move that saves the most usage is not tuning a schedule. It is one afternoon spent finding a non-browser route to each recurring number.

What you needThe browser routeLook for this firstSetup cost, once
A SaaS dashboard metricSign in, navigate, read the tileA scheduled CSV or emailed exportAn afternoon
Analytics figuresOpen the report, set the rangeA connector or an API pullAn hour
Mailbox contentsDrive the web clientA connected mailboxMinutes
A finance system balanceSign in behind 2FA each timeAn export you refresh weeklyAn hour, then five minutes a week
Social mentionsScroll a feedA feed, list, or search exportAn hour
Competitor pricesLoad the public pageThe public page. No shortcutNothing to do

The last row is the honest one. Some jobs have no other route, which is why Competitor Website Watch reads public pages and never interacts, keeping its per-run work to a visible ceiling.

One second-order effect matters more than the arithmetic. Browser work is where retries come from, since layouts change and sessions expire, so the expensive route is also the flaky one and flakiness feeds the lever with no ceiling. What the shared machine means beyond cost is in one computer, many screens.

Fill the model in with your own numbers, because nobody else's transfer

You are estimating against an allowance whose size is not published, so an estimate built from published figures is impossible in principle. Measurement is the only honest input, and it takes a week.

// MEASUREMENT PROTOCOL, once per bot
1. Run the bot manually, not on a routine, five times over three days,
   against your real data rather than a tidy sample.
2. Note your account usage before run one and after run five. Divide.
3. Take the MEDIAN of the self-report lines, not the mean. One long
   document will drag a mean somewhere useless.
4. Multiply by the cadence you were about to pick, before you pick it.
5. Repeat the whole thing once with your input ceiling removed, on a day
   the source material is unusually long. The gap between the two numbers
   is what the ceiling is buying. No gap means it is too loose to bind.

// THE SELF-REPORT LINE THAT MAKES STEP 3 POSSIBLE
End every report with one line, even on a clean run:
  runs=1 | calls=<n> | pages=<n> | items=<n> | retries=<n> |
  ceiling=<none|which>
TermWhere the number comes fromExampleYour figure
Runs per monthThe cadence you are weighing44
Tool calls per runMedian across five manual runs9
Page loads per runMedian across the same five0
Items read per runMedian across the same five40
Retry rateFailed steps over total steps0.08
Ceiling hits per weekRead off the self-report line0

Then run the check that can fail. Multiply your per-run figure by the schedule you had in mind. If hourly gives 720 runs a month and the arithmetic makes you wince, the schedule was wrong rather than the bot.

Measure before you bound, in that order. A ceiling guessed in week one truncates output that then gets blamed on the model, which is how people conclude a workable bot does not work. The exception is the retry ceiling, which goes in on day one because it guards against a runaway rather than against drift. The full per-run formula is in the Grok Bot cost breakdown, and the habit of proving a setup rather than trusting it is in testing your bot.

Price the roster, because six bots is not six times one bot

One bot is easy. You watch it for a week and nothing surprises you. That stops working around the fourth bot, because the total is no longer a sum of things you understand. It is a sum of things that each vary, one of which occasionally varies a lot, and nothing tells you which moved.

Roster sizeWhat dominates the billWhat breaks firstThe control that matters most
One botWhatever that bot readsNothing, you can feel itA cadence chosen on purpose
Two or threeStill the heaviest single readerYour memory of which is whichOne job per bot, so usage is attributable
Four to sixDuplicated reading, plus one heavy botYour review timeExactly one bot owns each source
Seven to twelveCorrelated weeks, plus review loadYour attention, before the invoiceA weekly review and a stand-down order
More than twelveThe roster itselfKnowing what any of them are forDeletion

Two effects make growth superlinear and both are avoidable. Three bots reading the same newsletter pay for it three times, and the fix is an ownership rule rather than a cost optimisation: one bot reads each source and writes a digest the others read. And every new bot adds review load, which is the constraint that actually binds.

Splitting earns its place for a second reason. With no audit view of bot actions yet, a bot doing four jobs blends four cost profiles into one signal you cannot tune. One job per bot is the only way to get attributable usage from an environment that attributes nothing. The structural version is in running a team of bots without chaos, and the forecast arithmetic is in keeping bot costs predictable.

Correlated weeks are what actually cross the allowance

Rosters do not cross an allowance because one bot went wrong. They cross because several bots reacted normally to the same unusual week. Each stayed inside its own range. Together they moved.

The weekBots that rise togetherWhy they move togetherStand down first
A launchInbox, mentions, support, researchOne event creates volume everywhereThe research bot
An incidentMonitoring, support, standup, commsOne incident feeds all of themThe digest bots
A press momentInbox, mentions, lead research, CRMInbound arrives in every channelLead research
Quarter endFinance, reporting, reconciliation, KPIThe calendar, which you knew aboutNothing, plan for it
A vendor adds a login stepEvery bot touching that vendorOne wall, several retry loopsThe bot that hit the wall
A source got slowerEvery browser-driven botEach run waits and observes moreThe tightest clock

The bottom two rows are different in kind. They are failures wearing a cost costume, and tuning a schedule in response fixes nothing. Telling the two apart is the subject of the complete reference on when bots go wrong.

The cheap defence against the top four is a written stand-down order: which bots pause, in which sequence, when the week goes sideways. Decide it while calm, because the week you need it is the week you have no time. The charter block below carries one.

Your attention is the second bill and you cannot top it up

Usage is refillable. Your reading is not, and it decides how many bots you can actually run. Here is one illustrative week, with numbers to replace with yours.

BotMinutes you spend on it weeklyDecisions it changedMinutes per decision changed
Inbox triage2593
Lead scout2037
Newsletter digest150Undefined
Competitor watch818
KPI report616
Standup scribe50Undefined

The two rows with no denominator are the finding, and neither is a usage problem. A bot producing correct output on schedule that changes nothing you do costs you twenty minutes a week forever, and no cadence tuning addresses that.

The rule that follows is unpopular. A seventh bot that pushes you past the reading you can genuinely do makes the other six less valuable, because now you skim all of them. Before adding one, delete one or move an existing bot to the sampling regime in watching what your bot did, where you read three runs a week at random instead of every output.

Measure cost per decision changed, because cost per run flatters everything

There are two ways a bot wastes money and only the cheap one worries people. A broken bot is loud: it fails, you notice, you fix or delete it. The expensive failure runs correctly, produces accurate output on schedule, and changes nothing. Nothing looks wrong, so it survives for months.

Decide the counting rule before you count. A changed decision is a message you sent, a call you made, a task you dropped, a meeting you moved, a price you went back to check. Reading the output and thinking "good, nothing to do" does not count. That may be worth something as reassurance, and reassurance should be priced as reassurance rather than smuggled in as impact.

BotRuns a monthDecisions changedRank by cost per runRank by cost per decision
Standup scribe220CheapestLast, and undefined
Inbox triage4436Most expensiveFirst
Competitor watch44CheapSecond
KPI report11Cheapest per monthThird
Newsletter digest221MiddleSecond to last

The ranking inverts, which is the whole point. On a cost-per-run basis you would cut the inbox bot and keep the standup scribe. On a cost-per-decision basis the inbox bot is the one earning its usage and the scribe is the one to delete. Count over four weeks and let the second ranking decide. What to do about each row is in the Grok Bot cost breakdown.

Spot the bot that stopped earning before it stops being obvious

A bot rarely announces that it has stopped being worth its usage. It degrades into background noise, and you stop reading it before you decide anything about it. These are the signals, each with a test that can fail.

The signalWhat it usually meansThe test that can failThe decision
You skim it and act on nothingThe value is event-shaped, not scheduledTurn it off for a weekConvert it to an event trigger
You open the source yourself anywayIt does not answer your real questionWrite down the question you open the source forRewrite the charter, keep the cadence
It agrees with what you already decidedIt is confirming, not informingCount the times it changed your mindKeep it as reassurance, priced as such
Counts falling, nobody decided thatIt succeeds on a shrinking sliceCompare this month's examined count with month oneFix the scope before judging value
No handoff in two monthsIts stop conditions are unmeasurableTest the rules against two ambiguous itemsFix the rules, then re-evaluate
You would not build it again todayThe job changed and the bot did notRewrite its charter in five minutesDelete it, or replace it

Row four is not a value problem at all, and reading it as one deletes a broken bot while leaving the breakage in place. A quietly narrowing bot looks worthless and is actually failing, which is why an examined count belongs in every report. Row five is the same family: a bot that never stops is usually missing instrumentation rather than performing well, argued in designing the handoff. Row one carries the only test here that cannot be gamed.

Bot Advisor is a reasonable place to put the listing work, and its boundary is the right one: it never deletes or rewrites another bot without your explicit say-so. Automate the review, keep the killing manual.

Write the budget block into every charter, then write the stand-down order

Since the runtime has no cap, the charter is where it goes. Two things in the block below turn a bad week into one instruction rather than a roster-wide audit: a cost class on line one, and a stand-down order.

// COST CLASS, the first line of every charter
CLASS: B          // A = must run, B = should run, C = nice to have
STAND-DOWN: if I say "roster over budget", class C stops until I say
resume, class B drops to its slowest listed cadence, and class A does
not change. Confirm in one line which rule you applied.

// CADENCE
Run twice on weekdays at 08:00 and 15:00 Europe/London.
Slowest listed cadence: once on weekdays at 08:00.
Never run twice within one hour. If I ask you to check again, refuse
and tell me when the next run is.

// PER-RUN CEILINGS
At most 12 tool calls and at most 6 page loads per run.
Read at most the 20 newest items. Never re-read an item you covered.
For any document over 20 pages, read the summary and the tables only,
then list what you skipped.
Prefer an export, a feed, or an API over opening a page, every time.
When you hit a ceiling, stop, report what you covered, and name what
you did not reach. Never continue past a ceiling to finish the job.

// RETRY CEILING
Two attempts at any single step, then stop and record it as failed.
Never a third attempt, and never an alternate route to the same result.
Stop immediately at a captcha, a 2FA prompt, or a login that fails once.

// SELF-REPORT
End every report with one line, even on a clean run:
  class=<A|B|C> | runs=1 | calls=<n> | pages=<n> | items=<n> |
  retries=<n> | ceiling=<none|which>

// WHERE YOU STOP
Never start a purchase, an upgrade, a paid trial, a credit top-up, or a
subscription. Never create an account. Never accept terms.
If a task needs spend, describe it in one line and wait for me.

The self-report line caps nothing and is the clause people cut first. It exists because no audit view does. A counter the bot writes itself is the only per-bot number you will ever have, and when a page count doubles between two Tuesdays you have found the change before the invoice does.

The last block is the boundary, doing double duty. The line that keeps a bot from buying something is also the only hard stop between an enterprising run and an invoice you did not authorise, and there is no toggle for it. The case for writing that line before the workflow is in the bot boundaries guide, and a charter you can fill in from scratch is in the charter template.

A hosted bill and a self-hosted bill are the same arithmetic on different invoices

The arithmetic underneath does not change when you self-host. What changes is which invoice each line lands on, and which levers you are allowed to pull.

Line itemHosted, Grok Bot shapeSelf-hosted, Rakazo shapeWho you call when it spikes
AccessOne subscription, published priceNothing, the runtime is yoursYour card statement
Model usageWeekly allowance, then overflowWhoever serves the modelYour model provider
Model choice as a leverNot available, no model pickerYours to changeYourself
ComputeIncluded, a managed Linux VMYour machine, container, or cloudYour hosting bill
StorageIncludedYoursYour hosting bill
Upgrades and breakageThe vendor's problemYoursYourself, on a Sunday
Your timeSetup and reviewSetup, review, upgrades, breakageNobody

The third row reorders the decision. On Grok Bot the standard lever of running a cheaper model for the boring jobs does not exist. On a self-hosted runtime it is your main lever, which makes picking a model a cost decision as much as a quality one, as choosing a model for Rakazo works through. The last row is the one self-hosting comparisons omit, and it is usually the largest.

The runtime comparison is in Rakazo versus Grok Bot, what self-hosting involves is in the self-hosting walkthrough, and the wider field is in open source bot runtimes compared.

The cheapest bot is the one you decided not to build

Every bot has a build cost before it has a usage cost: writing the charter, testing it, reading every output for a week, and two rounds of tuning. Call it a few hours. That is the number to compare against, and it kills more candidate bots than any usage figure ever will.

The jobVerdictWhy
Recurs on a schedule, and you always actBuild the botRecurrence is the whole case
Recurs often, decided differently each timeA saved prompt you runThe judgment is the work
Happened three times this yearNeitherYou will forget the bot exists
Takes you four minutes a weekA saved prompt at mostSetup costs more than the year
Somebody already reports it to youNothingPaying twice for one answer
You want it because a bot couldNothingThe reason is the tell

Row two is the one people get wrong most, and the giveaway is a charter you cannot finish writing. If you cannot state what the bot owns, what good output looks like, and where it stops, you do not have a role yet. You have a task, and tasks belong inside an existing bot or in your own hands.

Which three to build first is argued in the starter roster, and the first week of running any of them is in your first week with Grok Bot.

Answer the argument that none of this ever reaches an invoice

The strongest objection is straightforward. For a solo operator running three or four bots on daily schedules, usage sits inside the included allowance every week, none of this arithmetic becomes money, and the afternoon spent on it returns nothing. A fifth bot would have been worth more.

That is correct under four conditions together: a small roster, nothing tighter than hourly, no bot reading long documents, and an allowance you have never crossed. If all four hold, skip to the retry ceiling and go build something.

There is a sharper version aimed at this page. You told me not to trust any article's prices, then wrote thousands of words about cost. The answer is the distinction the page opens with: the price is a ten second lookup, and the cost is a design decision you make in a dropdown and live with for a year. One of those deserves an article, and it is not the one everybody writes.

The four conditions break quietly, which is the real argument. Rosters grow past eight bots because each one seemed free. A schedule gets tightened during a busy week and never loosened. A research bot starts being handed PDFs. A login expires and the retry loop finds it at 2am. So split the work in two. Tuning is optional and often not worth the hour. Insurance is two clauses and four minutes, and the retry ceiling plus the no-spend boundary are worth writing even if the objection holds completely, because they stop the runaway rather than the drift.

Where cost is the wrong lens on the problem

Three kinds of work sit outside everything above.

Investigative work loses most of its value under a ceiling. A research bot chasing an unclear question does variable work because the question is variable, and a hard cap turns a real answer into a partial one delivered on time. Run those attended, and let your attention be the ceiling.

Work with a deadline attached to money is second. Where being late costs more than being expensive, use report-and-continue rather than a hard stop: the bot says it passed its expected volume and keeps going, so you learn about it instead of finding a half-done job.

The third matters most, and a usage screen cannot see it. The most expensive run a bot ever performs is not the one that consumed the most usage. It is the run that sent the wrong thing to the wrong person, posted from your account, or filed a payment against bank details that arrived inside an invoice. Those cost credibility and rework, neither of which appears on a billing page.

Which is why the counterpart to this page is the complete reference on when bots go wrong, and why the cheapest insurance here is a capability you never granted. A bot that cannot send cannot produce that run at any price, which is the case for building a bot that drafts but never sends first and connecting the minimum, not the maximum after.

Keep reading: The Starter Roster, Grok Bot vs Zapier, Running a Team of Bots Without Chaos.

Frequently Asked Questions

How much does an AI agent cost per month?

Access and usage are separate numbers. For Grok Bot, the cheapest published paid route as of 25 August 2026 is Cursor Pro+ at 60 USD a month for an individual or Cursor Teams Standard at 40 USD per user for a team, with SuperGrok Plus at 100 USD a month as the x.ai route, plus a one-time trial. Usage on top of that is not predictable from any published figure, because it depends on how often your bots run and how much each run reads. Measure your own per-run consumption over three days and multiply by the cadence you want.

Does Grok Bot have a spend cap or a budget limit?

No. The documentation states directly that there is no Grok Bot specific spend cap yet. Eligible subscriptions include a weekly usage allowance, its size is not published anywhere, and usage beyond it is billed on demand from model and token cost. Because no setting stops a runaway for you, the only ceiling is the one you write into each charter: a stated cadence, numeric limits on tool calls and page loads, an input ceiling for long documents, a hard retry limit, and a boundary forbidding the bot from initiating spend.

What makes one AI agent run cost more than another?

Run frequency multiplies everything else, so the schedule decides more than any other setting: a five minute clock is 288 runs a day against one. After that it is the material pulled in per run, since one long transcript can outweigh a week of ordinary work. Then tool calls, because every source checked is separate work. Browser workflows cost more than API calls for the same answer, because each step must be loaded and observed. Retries are the only driver with no natural ceiling, which is why an explicit attempt limit matters most.

How do I know whether a bot is worth what it costs?

Count decisions changed rather than runs completed, over four weeks. A changed decision is a message you sent, a call you made, a task you dropped, or a price you went back to check. Reading the output and thinking nothing needs doing does not count. A bot that changed a decision most days it ran is earning its usage. One that changed three to eight decisions has real value at the wrong cadence, so halve the frequency. One that changed nothing while producing correct output is working and worthless, and tuning cannot fix a job you do not need.

What AI Bots Actually Cost | botskills.sh