grok botgrok 4.6xai+17

Grok Bot and Grok 4.6: Cloud Agents, API, and Pricing

Grok Bot and Grok 4.6 launched a day apart in August 2026, and the name now points at three different products. This guide separates them, walks through installing Grok Bot and creating your first agent, covers skills, routines, approvals, and the shared-computer security model, provides working API code with real cost math, and examines the benchmark rows, pricing cliffs, and hallucination figures that most of the launch coverage quietly left out.

Parash P

Sep 30, 2026
51 min read
34 views

Grok Bot and Grok 4.6: Cloud Agents, API, and Pricing

Two Different Products, One Confusing Name

If you searched for "Grok bot" in August 2026, you landed on a name that now points at three separate things. Getting them straight before you spend money is the single most useful thing this guide can do for you.

1. The @grok account on X. The reply bot. You tag @grok under a post and it answers in the thread with the post as context. Free to try, rate-limited, and by far the most-used Grok surface. It is a conversational assistant, not an agent — it cannot do anything except reply.

2. Grok Bot, the desktop product. Launched in early beta on August 11, 2026. This is a completely different thing: persistent, named AI agents that run on a cloud computer with a browser, a filesystem, and a terminal. They sign into your tools, work across apps, run while your laptop is closed, and hand work off to each other. It is paywalled behind three high-tier subscriptions and has no free tier.

3. A bot you build yourself on the Grok API. The grok-4.6 model behind an API key, wired into your own application, Slack workspace, Discord server, or agent loop. Pay-per-token, no subscription required.

Layered underneath all three is Grok 4.6, released one day after Grok Bot on August 12, 2026, as SpaceXAI's frontier model. It powers the API, it is the default model in Grok Build, and it shipped in Cursor on all plans.

This guide covers all of it: what actually changed in 4.6, how to set up and run Grok Bot end to end, how to build your own bot on the API with working code, what everything costs, and — the part most launch coverage skipped — where the numbers do not hold up.

A note on freshness before you read further: nearly every fact here dates to a two-day launch window in mid-August 2026. Prices, plan names, and platform support are moving fast. Verify anything you are about to pay for.

Grok 4.6: What Actually Shipped on August 12, 2026

SpaceXAI — the entity formerly known as xAI, which SpaceX acquired in February 2026 — released Grok 4.6 as a direct successor to Grok 4.5, which had shipped roughly a month earlier on July 8. That is an aggressive cadence by frontier-lab standards.

The framing from xAI is specific: 4.6 is built for long-running agents and for more ambitious interactive and visual work. The company's stated goal is a model that stays coherent across many steps — researching a topic, working across a codebase, turning an idea into a finished application or work artifact — rather than one that answers a question well and stops.

The Specification Sheet

Property

Value

API model name

grok-4.6

Context window

500,000 tokens

Knowledge cutoff

February 1, 2026

Modalities

Text and image input; text output only

Output limit

No stated text output limit

Input price

$2.00 per 1M tokens

Output price

$6.00 per 1M tokens

Cached input

$0.50 per 1M tokens

Reasoning effort

low, medium, high (default), xhigh

APIs

Responses API, Chat Completions

Parameter count

Not published

Note the last row. SpaceXAI did not publish a parameter count for Grok 4.6, and any article giving you one is speculating.

How It Was Trained

xAI was unusually candid in its launch post about the training recipe, and the details matter if you are trying to predict where the model is strong.

Grok 4.6 got a longer supplemental training run than 4.5 did, using curated model-generated data aimed at reasoning and advanced technical concepts, plus what the company calls high-quality engineering data, plus an improved optimizer and training recipe.

Then the interesting part: xAI used Grok 4.5 itself to regenerate the supervised fine-tuning trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work — then filtered out problematic traces with model-based checks. The previous generation taught this one.

The reinforcement learning stage was pointed at a wide range of agentic tasks: knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and computer-aided design.

The behavior xAI highlights as emerging from this process is self-verification — on longer trajectories, the model checks its own work before moving on. That claim is from xAI and has not been independently verified. Treat it as a hypothesis you should test on your own workload.

The Benchmark Table, Read Honestly

xAI published a ten-row eval table. Marketing quoted the winning rows. Here is the full table as published:

Evaluation

Grok 4.6 High

Grok 4.5 High

GPT-5.6 Sol Max

Fable 5 Max

AA Intelligence Index

61

56

61

62

GDPVal-AA v2

1753

1526

1728

1741

CursorBench v3.2

69.9%

66.7%

67.2%

70.5%

DeepSWE v1.1

65.9%

54%

73%

70%

FrontierCode v1.1 (Extended)

61.3%

56.6%

60.6%

63.6%

APEX-Agents

57.5%

47.1%

56.7%

59.2%

Terminal-Bench v3.0

26%

15.7%

34.6%

34.1%

APEX-SWE

56.4%

53.6%

—

58.8%

AA-Briefcase

1577

1313

1502

1574

Harvey LAB (Vals)

15.8%

12.9%

2.5%

11.3%

The Artificial Analysis Intelligence Index is a composite of nine benchmarks. Grok 4.6 scores 61, up five points from Grok 4.5's 56, tying GPT-5.6 Sol Max and landing one point behind Claude Fable 5 Max at 62. Several third-party trackers place Claude Opus 5 at 63, above all of them, though Opus 5 does not appear in xAI's own comparison table.

Where Grok 4.6 genuinely leads: GDPVal-AA v2 (real-world professional tasks) at 1753 Elo, a 227-point generational jump. AA-Briefcase (long-horizon analyst work) at 1577, up from 1313. Harvey LAB, a legal eval, at 15.8% against GPT-5.6 Sol's 2.5%.

Where it clearly loses: DeepSWE v1.1 at 65.9% against GPT-5.6 Sol Max's 73% — a 7.1-point gap. Terminal-Bench v3.0 at 26% against 34.6% — an 8.6-point gap. Both are the benchmarks that matter most for autonomous software engineering.

The honest read, which several independent reviewers reached the same week: this is a knowledge-work model that was announced as a coding model. Its wins cluster in research, analysis, and document-heavy agentic tasks. Its losses cluster in exactly the terminal-driven, repo-wide engineering work that "agentic coding" implies. Grok 4.6 did improve on every coding benchmark over 4.5 — the generational gains are real — but improvement and leadership are different claims.

The Row Nobody Printed

Artificial Analysis measured something xAI's launch post did not include: an AA-Omniscience accuracy of 48.2% with a non-hallucination rate of 65.7%.

Unpack that second number, because it is the one that should shape your deployment decision. A 65.7% non-hallucination rate means that when the model produces a non-correct response, roughly a third of the time it states a confident fabrication instead of acknowledging uncertainty.

This is a continuation of a pattern, not a one-off. Grok 4.5 showed the same shape at launch: AA-Omniscience accuracy climbed sharply from the prior generation while the hallucination rate more than doubled. The model learned more and got worse at knowing what it did not know.

Two important caveats in both directions. First, the 65.7% figure comes from a single independent evaluator and had not been reproduced by a second measurement pass at the time of writing — treat it as a strong signal, not settled consensus. Second, hallucination benchmarks are gameable in a specific way: a model that refuses more questions scores better on non-hallucination while being less useful. The metric rewards caution, and cautious models can be worse products.

That said, the practical guidance is straightforward. For high-volume agentic work where output gets reviewed anyway, Grok 4.6's price-to-intelligence ratio is genuinely excellent. For customer-facing output, medical or legal summaries, or anything where a confidently wrong answer is expensive, you want a verification layer or a different model.

Reasoning Effort and the New xhigh Level

Grok 4.6 introduces a fourth reasoning tier. This is one of the more useful controls in the API, and most people leave it at the default without thinking about it.

Setting

Behavior

Best for

low

Some reasoning tokens, still fast

Latency-sensitive agentic loops, simple tool calling

medium

More thinking, moderate latency

Complex data analysis, long-context reasoning

high (default)

Deeper thinking, more reasoning tokens

Hard problems, complex math, multi-step logic

xhigh

Maximum depth, correspondingly higher latency

The hardest problems, where quality beats speed

Three things worth knowing:

  • Reasoning cannot be disabled. If not specified, reasoning_effort defaults to "high".

  • xhigh is new to 4.6. On models that do not support it, such as grok-4.5, an xhigh request is silently treated as high.

  • Reasoning tokens are billed. They count toward your total consumption at output rates. Setting xhigh on a high-volume loop is a real cost decision, not a free quality upgrade.

Also note: presencePenalty, frequencyPenalty, and stop cannot be used with reasoning models. Requests including them return an error. And logprobs/top_logprobs are silently ignored on grok-4.20 and newer.

Grok Bot: The Agent Product

Grok Bot launched on August 11, 2026 — one day before Grok 4.6 — in what xAI labels an early beta. It began as an internal prototype at SpaceXAI, spread across the company, and was opened up externally.

The premise: instead of prompting an agent, you hire one. Bots are persistent, named teammates with a job, a conversation, and working context that accumulates.

What Makes It Structurally Different

Each Bot runs on a persistent cloud VM with a browser, filesystem, and terminal. It can use connectors and MCP where those exist, and fall back to computer use — literally clicking through a UI — for apps and websites with no clean API. Work finishes in the actual tool rather than as a chat draft you have to paste somewhere.

There is no workflow builder. You create a Bot, message it, grant access as needed. Nothing to configure up front. The same Bot is reachable from desktop and iOS with synced conversations.

Bots coordinate with each other. Multiple Bots share one user-scoped computer and run in parallel. They can message each other directly, share context in threads or group chats, and pass ownership of work. You are not the router between tools.

Bots learn from demonstration. Ask a Bot to follow along while you do a multi-step workflow once. It persists that path as a routine and can re-run it on a schedule.

State is durable. Named Bots keep memory, files, browser sessions, and preferences across turns. Context compounds instead of resetting to a fresh environment every task.

The Cursor Dependency

Here is the structural detail that most launch coverage buried, and it changes how you should evaluate the product.

Grok Bot authenticates with a Cursor account. The download link points at downloads.cursor.com. Sales routes to Cursor's contact form. Privacy and data-sharing choices are managed through Cursor account settings. Training opt-out follows applicable Cursor account and privacy settings. Backend retention follows Cursor terms. Enterprise access goes through your Cursor account team.

xAI's own security documentation explicitly tells you to review Cursor's published security documentation rather than treating the user-level computer assignment as a broader guarantee.

Two of the three eligible plans are Cursor plans, not xAI plans. If your organization has a vendor review process, Grok Bot is a Cursor procurement question as much as an xAI one. Budget time for that.

Access, Platforms, and What It Costs

There is no free tier and no published trial. Access requires one of exactly three subscriptions:

Plan

Approximate price

Notes

SuperGrok Heavy

~$300/month

xAI's top consumer tier; not listed on xAI's public pricing page

Cursor Ultra

$200/month

Cheapest individual route

Cursor Teams Premium

$120/seat/month

Cheapest per-seat route

Free Grok, standard SuperGrok (~$30/month), SuperGrok Plus (~$100/month), and ordinary Cursor Pro are not eligible. This is the single most common misunderstanding about Grok Bot, and it is the reason "SuperGrok Heavy" started trending again on launch day.

Supported platforms at launch:

  • macOS on Apple silicon and Intel

  • Windows on x64 and Arm64

  • iPhone on iOS 18 or later

Linux desktop, Android, and iPad are not supported at initial launch. Note that some launch coverage claimed a Linux build; xAI's own FAQ says otherwise.

The Billing Detail That Should Shape Your Pilot

Eligible subscriptions include a weekly Grok Bot usage allowance. Beyond that allowance, eligible accounts can add on-demand usage billed from model and token cost.

xAI does not publish the size of the included allowance. It does not publish the difference in Bot limits between the three eligible plans. And, critically, there is no Grok Bot-specific spend cap documented at launch. There is also no model picker inside Grok Bot, so you cannot route work to a cheaper model when the meter climbs.

The subscription is an access floor, not a budget. If you are piloting this on a team, treat the monthly seat price as the smallest line item and instrument your actual usage from week one.

Tutorial: Setting Up Grok Bot End to End

This walkthrough follows xAI's documented flow. Budget about fifteen minutes for setup and another fifteen for a first useful result.

Step 1: Install the Desktop App

Open the Grok Bot access page at cursor.com/bot/onboarding and choose the download for your machine.

macOS:

  1. Choose the Apple silicon or Intel download.

  2. Open the downloaded disk image.

  3. Drag Grok Bot to Applications.

  4. Open Grok Bot. If macOS asks for confirmation, choose Open.

To check which build you need: Apple menu → About This Mac. A Chip field means Apple silicon; a Processor field means Intel.

Windows:

  1. Choose the x64 or Arm64 download.

  2. Run the installer.

  3. Open Grok Bot from the Start menu.

To check architecture: Settings → System → About → System type.

The app checks for updates automatically. You can also force one from Settings → Beta → Check for Updates.

Step 2: Sign In

  1. Choose Get started on the welcome screen. If you are already in the app, use Sign In with Cursor from Settings.

  2. Complete authentication in the browser window that opens.

  3. Return to Grok Bot after the browser confirms sign-in.

If your organization requires SSO, complete the normal organization sign-in flow.

⚠️ Blocker to know about in advance: Grok Bot requires cloud data storage and does not support Legacy Privacy Mode. If your Cursor account is on Legacy Privacy Mode, you must move to a supported data setting before Grok Bot will start. Change this at cursor.com/dashboard/settings?openPrivacy=true before you begin, or you will hit a wall at first launch.

On first use, the app introduces Bots, the shared computer, and routines, then asks which tools you use. Those answers only shape the initial teammate suggestions — they do not connect or modify anything. Computer setup runs in the background while you finish onboarding.

Step 3: Create Your First Bot

Pick a suggested teammate from Meet a future teammate, or choose Create your own.

Or create one manually:

  1. Choose New in the sidebar, or press Cmd/Ctrl+N.

  2. In New chat, select Create new agent.

  3. Grok Bot creates and opens a Bot named New Agent.

  4. Open Bot actions → Edit Profile to set name, title, description, and avatar.

Give it three things: a short name, one primary job, and a description of how it should work.

Here is xAI's own example, which is a good template because it shows the shape of a well-scoped Bot:

Name: Piper Job: Product performance Description: Investigate product-performance questions using our observability tools. Preserve links and screenshots, separate evidence from hypotheses, and return a short summary with the highest-impact issue first. Never change production settings.

Notice what that description contains: a domain, a source system, an output format, an epistemic rule (separate evidence from hypotheses), and a hard boundary (never change production settings).

Scope your Bots narrowly. Good jobs look like Talent Scout, Expense Manager, Bug Reproduction. A job called General Helper gives the Bot less guidance and makes its accumulated context harder to reuse. Focused Bots build more useful memory.

Limits: an account can have up to 50 Bots and group chats combined.

Step 4: Give It a First Task

A strong request has five components. This is the single highest-leverage habit to build:

  1. Outcome — what should be finished?

  2. Sources — which apps, websites, files, or conversations matter?

  3. Constraints — what must the Bot avoid or ask about first?

  4. Deliverable — what should it return?

  5. Review point — when should it stop for you?

For a five-minute first result that needs no connector or login, attach a document and try:

Summarize this document in five bullets. List every date, decision, and open
question in a separate section. Cite the page or section for each item. Do not
change the source file.

Then step up to something in one of your tools:

Open our analytics dashboard and compare new-user activation for this week
with the previous four weeks. Identify the largest step-level change and
draft a short investigation plan with links to the relevant charts. Do not
change any dashboards. Ask me to sign in if needed.

Read that second prompt again and note the structure: outcome, source, comparison window, deliverable, two explicit constraints, and an escape hatch for authentication. That is the pattern.

⚠️ Review the approvals documentation before you allow any external changes. Do not start with a task that sends, publishes, or purchases anything.

Step 5: Sign In to the Tools It Needs

When the Bot hits an app that requires authentication, it will ask you to take over the computer.

  1. Open Agent Computer from the conversation.

  2. Choose the takeover control.

  3. Enter the password, passkey, two-factor code, or complete the CAPTCHA yourself.

  4. Return control to the Bot.

The browser session persists on the shared computer afterward, so other Bots on your account can use the same signed-in session. That is a convenience feature and a security consideration at the same time — more on that below.

For supported services, you can install a connector instead from Settings → Plugins and authenticate it in your browser. Prefer connectors where they exist: they are structurally more reliable than clicking through a UI that can redesign itself out from under the agent.

Step 6: Review and Correct

Ask the Bot to revise anything incomplete or badly formatted. When you find a preference you want to persist, name it explicitly rather than just correcting the output:

Use this format for future weekly reports: five bullets, source links inline,
and a final section called "Decisions needed."

There is an important distinction between the two places instructions live:

  • The Bot description holds rules that should remain true. "Never send external messages without approval."

  • The conversation holds task-specific instructions. "Draft follow-ups for these twelve accounts."

Putting a durable rule in a message means it decays. Putting a one-off task in the description means it pollutes every future run.

When the process is stable, save it as a skill.

How the Shared Cloud Computer Works

This is the part that most determines whether Grok Bot is safe for your use case, so it is worth understanding precisely.

All of your Bots use the same computer. Not one per Bot — one per user account. That means:

  • Browser cookies and signed-in sessions are shared

  • Files are visible to every Bot

  • Command-line credentials are shared

  • One Bot can continue from work another Bot saved

Each Bot gets its own screen on that shared computer, so several Bots can drive browser and desktop tools in parallel. One Bot can run one computer-use task on its screen at a time. But screens are separate work surfaces, not separate security boundaries.

xAI states this plainly, and it deserves to be stated plainly here too: do not use separate Bots as a security boundary. If you sign a Bot into your production admin console, every Bot on your account can reach that session.

Watching and Interrupting Work

Open Agent Computer from a conversation to see the shared desktop. The preview shows clicks, typing, navigation, and current status. You can close the preview, the app, or your laptop and cloud work continues.

Files

The computer has a shared workspace at /workspace. Ask Bots to keep durable project files there in clear project folders.

Files, browser state, and supported sign-ins are designed to survive normal computer updates and recovery. Treat temporary directories, manually installed packages, and uncommitted application state as replaceable — copy important results into /workspace or attach them to the conversation.

Recovery Options

When the computer becomes unreachable, use Recover computer from the error state. For planned maintenance, Settings → Beta offers three escalating options:

  • Update Agent Computer — rebuilds with the latest image while preserving durable state

  • Recover Agent Computer — replaces an unreachable computer while preserving durable state where offered

  • Reset Agent Computer — returns to the most recent durable snapshot and can discard recent unsaved work

Use them in that order. Wait for active work to finish before recovery where possible.

Your Local Machine Is Separate

The Grok Bot cloud computer is not your Mac or PC. Access to your local machine is a distinct capability, controlled at Settings → General → Agent → Execution on Local Computer, with three settings: always require approval, always allowed, or never allowed.

The default is Ask every time. Unless a Bot has a specific reason to touch your local files, set this to Never allowed. It does not affect the cloud computer at all.

Skills and Routines: Making Work Repeatable

Grok Bot has two building blocks for automation, and conflating them is a common early mistake.

  • A skill is a reusable set of instructions for how to do a task.

  • A routine tells one Bot when to run a workflow — on a schedule, or where supported, after an event.

The recommended sequence is deliberate: run a one-time task, make it reliable, save the method as a skill, and only then automate it. Skipping to automation is how you end up with a scheduled job that confidently produces garbage every Monday at 8 AM.

Saving a Skill

Ask the Bot directly:

Save the process we just used as a skill called "Weekly account health."
Include the source systems, risk definitions, output format, and the rule that
customer contact always requires approval.

A useful skill states six things:

  1. When to use it

  2. Required inputs and access

  3. The sequence of work

  4. How to validate the result

  5. What to return

  6. What requires approval

Skills are available across your Bots, though a Bot may need the relevant connector or login to actually use one. Discover and install supported connectors and packaged skills from Settings → Plugins.

Composer shortcuts: type / in the desktop composer to reference a saved skill; type @ for Bots, groups, routines, and connectors. If an installed private skill does not appear in the / menu, open Settings → Plugins → Yours and enable it for that Bot.

Teaching a Workflow by Demonstration

Where Teach a task is available, you can demonstrate a browser workflow instead of describing every step:

  1. Open a one-to-one Bot conversation and its computer view.

  2. Choose Teach a task.

  3. Describe the result you are about to demonstrate.

  4. Perform the workflow once.

  5. Stop the recording and review the skill the Bot creates.

  6. Test it on a safe example before scheduling it.

Constraints worth knowing: teaching records visible computer interaction for up to ten minutes and does not record microphone audio. Do not expose secrets during the demonstration — use the secure handoff flow for credentials instead.

The learned skill is a draft. One demonstration cannot teach decision rules, failure handling, or approval boundaries. Add those yourself before you trust it. And the feature may be rolling out gradually — if the control is not visible, ask the Bot to create a skill from written instructions plus the completed task instead.

Creating a Routine

Ask the Bot that should own the recurring job:

Every weekday at 8:00 AM, run the Daily customer-risk skill against the
current account list. Post a linked watch list in this conversation. Do not
contact customers. If the source data is unavailable, report the failure
instead of using old data.

That last sentence is not decoration. A stale-data policy is the difference between a routine that fails loudly and one that quietly reports last week's numbers as if they were current.

Confirm six things when the routine is created: the owning Bot, the schedule and time zone, the input source, the expected result, the approval boundary, and what should happen when a source is missing.

Background routines run while your laptop is closed.

Event Triggers

Cursor account integrations can start a routine from an event — a Slack message, a GitHub notification. These are separate from Slack or GitHub plugins and may need their own connection flow.

Define a narrow matching rule and a clear response:

When a message in #customer-escalations contains a support ticket link and
the phrase "needs repro," open the ticket, reproduce the issue in staging,
and post a repro pack in this conversation. Never post back to Slack without
approval.

Avoid broad listeners like "every new message." They generate noise, consume usage against an uncapped meter, and increase the chance of acting on irrelevant input.

Test Before Enabling

Use Test run after creating or editing a routine.

⚠️ A test run performs real work. It can navigate websites, change files, and call connected tools. Use safe inputs and keep write actions behind approval.

Review five things: whether it selected current inputs, whether output meets the required format, whether every action has a source or audit trail, whether it stopped at the intended approval point, and whether failure states are explicit.

Managing Routines

Open the Bot, choose View conversation details, then Routines. From there you can enable or pause, run a test, edit the schedule or instructions, inspect recent success and failure history, and delete.

Limits: a Bot can own up to 50 routines, and the app keeps the 20 most recent run records per routine. Deletion is immediate with no undo. Deleting a Bot also removes routines it owns.

Grok Bot may ask whether to keep routines running after a long period away and will pause them if there is no response — a sensible guard against an uncapped meter running unattended. Review paused routines when you return.

Design Principles for Routines You Can Trust

  • Automate preparation before execution

  • Have the Bot draft, reconcile, or recommend first

  • Require approval for sending, purchasing, deleting, publishing, or changing production

  • Include a no-data and stale-data policy

  • Make retries idempotent where possible

  • Tell the Bot where to report partial completion

  • Re-test after a website, connector, or source format changes

Approvals, Security, and Least Privilege

Grok Bot's security model puts a lot of weight on you writing good boundaries. Here is how to use the controls it provides.

Set the Boundary in the Request

Tell the Bot what it can do and where it must stop:

Reconcile the campaign data and draft a recommended budget change. Do not
change the campaign or message the agency. Ask for approval after showing the
current value, proposed value, and expected impact.

Prefer explicit boundaries for: sending messages or invitations, publishing content, purchases and financial transfers, deleting or overwriting data, changing permissions, production changes, and accepting legal terms.

An approval controls the proposed action. It does not reverse work already completed. That asymmetry is the whole reason to front-load boundaries rather than relying on catching things at the approval prompt.

Reviewing an Action

When an action needs approval, the conversation shows the proposed operation and its inputs. Review target, scope, and values before approving.

  • Desktop: Allow once continues, Deny blocks, Always allow can save a matching rule.

  • iPhone: Approve once and Deny.

Do not approve an action whose target or effect you cannot identify. Ask the Bot to explain it in plain language or produce a draft first.

Auto Review Rules

Where Auto Review enforcement is available, Grok Bot evaluates tool calls and computer actions before they run. Add rules at Settings → General → Auto-review.

  • Require Approval rules always stop matching actions

  • Always Allow rules let matching actions proceed only when automated review does not identify another reason to stop

  • When both match, Require Approval wins

Write narrow rules tied to a known action and scope:

✅ Require approval before sending any external email ✅ Require approval before changing a production dashboard ✅ Always allow running git status in /workspace/reports

❌ Allow everything in the browser

Websites and tool behavior change over time, and Auto Review is model-based. It should complement least privilege and explicit approval boundaries, not replace them.

One gotcha: personal Auto-review rules are stored on the current desktop and synced to its Grok Bot computer. Verify them separately on another desktop installation.

Credentials

For passwords, passkeys, two-factor codes, CAPTCHAs, and payment confirmations, the Bot should hand you control of the computer. Open Agent Computer, take control, complete only the blocked step, return control.

Never send a password or one-time code in ordinary chat.

If the Bot presents a secure secret request for a supported connection, enter the value there — it is masked, excluded from the transcript, and not shown to the model. It is not a general-purpose password manager, though.

Removing Access Properly

When a project or login should no longer be available, the order matters:

  1. Pause or delete related routines

  2. Sign out of websites on the shared computer

  3. Uninstall connectors and revoke their authorization in the source service

  4. Remove sensitive project files from /workspace

  5. Hide or delete Bots that should no longer appear

  6. Use the account settings flow if you need to delete the Cursor account

⚠️ Deleting a Bot does not remove shared-computer files or browser sessions. This is the most likely place to get a false sense of security. The Bot is gone; the signed-in Salesforce tab is not.

A Least-Privilege Checklist

  • Connect only the tools a workflow needs

  • Use scoped service accounts where the source system supports them

  • Start with read-only tasks and draft outputs

  • Keep sending, publishing, purchasing, deletion, and production changes behind approval

  • Review installed connectors and active routines regularly

  • Pause a routine when its source system or workflow changes

  • Preserve source links and an action log for important decisions

Running a Team of Bots

The pattern that emerged inside SpaceXAI is a chief-of-staff Bot sitting on top of specialists — one each for inbox management, expenses, recruiting, bug fixes, operations. Bots message each other, share context in threads, and can be placed in a group chat where they coordinate on their own, passing work and assigning ownership, pulling you in only for judgment calls.

Build up to that rather than starting there:

  1. Give one Bot ownership of an end-to-end outcome

  2. Add another Bot only when the work has a stable specialist role

  3. Put Bots in a group chat when the handoff itself needs to be visible

  4. Keep external actions behind a clear approval boundary

Duplicating a Bot is useful for scaling a role across scopes — one Account Health Bot per region. The copy carries profile, settings, enabled skills, routines, and avatar, but not conversation history, learned memory, or chat attachments. Rename it and give it the new scope before assigning work.

Pinning and hiding: pin active Bots to the top of the sidebar; Hide from sidebar removes a Bot from the main list without deleting its work. Restore from Show hidden chats → Unhide. Note that hiding does not pause the Bot or its routines — a hidden Bot with an active routine is still running and still billing.

What a Bot Remembers

A Bot retains stable working preferences, important facts, and summaries from its work, which lets it hold a role over time without replaying every message. Its conversation and learned role are separate from other Bots, while shared files, browser sessions, group messages, and direct handoffs move context between them.

Memory is not an authoritative source. Keep changing facts in the source system, ask the Bot to cite or reopen current data for consequential decisions, correct stale assumptions directly, and put explicit safety boundaries in the description rather than hoping memory holds them.

Eight Roles Worth Copying

xAI's documented use cases are a good starting roster, and they share a design principle: each owns a repeatable outcome, not a loose category of questions.

  • Sales Outbound — account research, contact prioritization, review-ready outreach. Returns a review list; does not send or enroll anyone.

  • Talent Scout — sourcing, candidate research, outreach drafts, scheduling prep. Does not contact anyone without approval.

  • Paid Media — campaign monitoring and budget recommendations with supporting numbers. Does not change budgets.

  • Expense Manager — weekly reconciliation, receipt matching, policy exception flagging with citations. Does not send or change reimbursements.

  • Product Performance — targeted investigations returning screenshots and direct links, separating facts from hypotheses. Does not change production settings.

  • Bug Reproduction — turns reports into reproduction packs with exact steps, environment details, and a minimal test case. Does not use production customer data.

  • Account Health — ranked watch list combining usage, support escalations, renewal timing, and stakeholder activity. Does not contact customers or edit the CRM.

  • Chief of Staff — a source-linked digest of what changed and what needs a decision. Does not send messages or change meetings.

Every single one of those ends with a constraint. That is not a coincidence — it is the design.

To turn any of them into something durable: put the job, sources, output format, and standing boundaries in the description; run one real task with safe scope; correct until reviewable; save as a skill; test on a second input; create a routine only when retries and failure cases are defined; keep consequential external actions behind approval.

The @grok Bot on X

The oldest and most-used Grok surface is also the least like the others. Tag @grok in a reply or post with your question and it answers in-thread with the tagged post as context. Anyone on X can use it, with usage limits on the free tier.

The dominant use case is fact-checking. Academic research covering March through September 2025 found 447,083 tweets tagging the bot specifically to request fact checks of other posts, and estimated nearly 1.4 million verification requests across Grok and Perplexity combined during that window — around 7.6% of all interactions with the two tools. "@grok is this true" became one of the most common messages sent to the bot.

That volume comes with a documented reliability record you should factor in before treating it as a verification tool. The bot has produced antisemitic output that xAI publicly apologized for, attributing the root cause to an update in a code path upstream of the bot rather than the underlying model. It has injected unrelated political claims into unrelated answers. It was briefly suspended from X in 2025. And through late 2025 into 2026, Grok's image capabilities generated sexualized imagery at scale, triggering formal investigations by Ofcom and the European Commission, action from dozens of US state attorneys general, and regulatory scrutiny in multiple additional jurisdictions.

None of that is a claim about Grok 4.6's behavior specifically — there is no evidence in the material reviewed here that 4.6 reproduces those particular incidents, and xAI says 4.6 shipped with its widest-ever pre-deployment safeguard testing. But it is directly relevant context if you are evaluating xAI as a vendor for regulated or brand-sensitive work, and it is the kind of thing procurement will find whether or not you mention it.

Practical guidance: the @grok bot is a fast way to get context on a post. It is not a citation, and it is not a substitute for checking a primary source.

Tutorial: Building Your Own Bot on the Grok 4.6 API

This is the route most developers actually want, and it has no subscription requirement. Get an API key from console.x.ai and set the model name to grok-4.6.

The Minimal Request

cURL:

bash

curl https://api.x.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-4.6",
    "input": "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"
  }'

Python, xAI SDK:

python

import os
from xai_sdk import Client
from xai_sdk.chat import user

client = Client(api_key=os.getenv("XAI_API_KEY"))

chat = client.chat.create(model="grok-4.6")
chat.append(user("Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"))

response = chat.sample()
print(response.content)

Python, OpenAI-compatible:

python

from openai import OpenAI

client = OpenAI(
    api_key=os.getenv("XAI_API_KEY"),
    base_url="https://api.x.ai/v1",
)

response = client.responses.create(
    model="grok-4.6",
    input="Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}",
)

print(response.output_text)

JavaScript, Vercel AI SDK:

javascript

import { xai } from '@ai-sdk/xai';
import { generateText } from 'ai';

const { text } = await generateText({
  model: xai.responses('grok-4.6'),
  prompt:
    'Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}',
});

console.log(text);

The OpenAI-compatible base URL means most existing OpenAI client code migrates by changing two lines. Both the Responses API and Chat Completions are supported.

Setting Reasoning Effort

python

import os
from xai_sdk import Client
from xai_sdk.chat import system, user

client = Client(
    api_key=os.getenv("XAI_API_KEY"),
    timeout=3600,  # reasoning models need a long timeout
)

chat = client.chat.create(
    model="grok-4.6",
    reasoning_effort="xhigh",
    messages=[system("You are a highly intelligent AI assistant.")],
)
chat.append(user("Find all prime numbers p such that p^2 + 2 is also prime. Prove your answer."))

print(chat.sample().content)

On the OpenAI-compatible client, the shape is reasoning={"effort": "xhigh"}. On the Vercel AI SDK, it is providerOptions: { xai: { reasoningEffort: 'xhigh' } }.

Set the timeout. The default client timeout will bite you on high and xhigh requests. xAI's own examples use 3600 seconds.

Streaming the Reasoning Summary

Grok 4.6 exposes summarizations of its internal reasoning, which is genuinely useful for showing progress in a UI during a long task:

python

import os
from xai_sdk import Client
from xai_sdk.chat import system, user

client = Client(api_key=os.getenv("XAI_API_KEY"), timeout=3600)

chat = client.chat.create(
    model="grok-4.6",
    messages=[system("You are a highly intelligent AI assistant.")],
)
chat.append(user("A projectile is launched at 30 m/s at 37 degrees above horizontal from a 45 m cliff. Find its speed on impact. (g=10 m/s^2)"))

for response, chunk in chat.stream():
    if chunk.reasoning_content:
        print(chunk.reasoning_content, end="", flush=True)

On the Responses API, watch for response.reasoning_text.delta and response.reasoning_summary_text.delta events.

You can also retrieve encrypted reasoning content by passing include: ["reasoning.encrypted_content"] to the Responses API, and send it back to give a later turn more context. With the Vercel AI SDK this happens automatically unless you set store: false.

The Caching Setting You Should Not Skip

This is the highest-value optimization in the entire API and it is one line.

xAI strongly recommends setting a prompt_cache_key on the Responses API — or the x-grok-conv-id header on Chat Completions. It routes a conversation's requests to the same server, which makes cache hits reliable. Without it, you frequently land on a cache-cold server and pay full input price on tokens you already sent.

Cached input runs at $0.50 per million against $2.00 uncached. On a multi-turn agent loop where the system prompt and tool definitions repeat every turn, that is a 75% discount on the bulk of your input tokens. Skipping it is the most common and most expensive Grok API mistake.

For long agent loops, also look at context compaction, and for tool-heavy workloads, the function calling documentation.

Server-Side Tools

Grok 4.6 supports xAI-hosted tools you do not have to implement yourself. Two components get billed: token usage, and per-invocation tool costs.

Tool

Tool name

Cost per 1k calls

Web Search

web_search

$5

X Search

x_search

$5

Code Execution

code_execution / code_interpreter

$5

File Attachments

attachment_search

$10

Collections Search (RAG)

collections_search / file_search

$2.50

Image Generation

image_generation

Imagine API rates

Image Understanding

view_image

Token-based

X Video Understanding

view_x_video

Token-based

Remote MCP Tools

set by each MCP server

Token-based

⚠️ Critical default: Grok has no knowledge of current events beyond its training data — February 1, 2026 — unless you explicitly enable search tools. If you are building anything that touches current information, Web Search or X Search is not optional. A lot of "Grok gave me outdated information" reports trace directly to this.

X Search is the genuine differentiator here. No competing frontier model has structured access to X posts, profiles, and threads. If your product depends on real-time social signal, that is a reason to pick Grok that has nothing to do with benchmark scores.

Two API-surface caveats: all tool names work in the Responses API, but in the gRPC API (Python xAI SDK), code_interpreter and file_search are not supported. And for view_image and view_x_video, you are not charged the invocation fee but are charged for the image or video tokens processed.

Long Context: The Pricing Cliff

This is the gotcha that produces surprise invoices.

Prompt size

Input

Cached input

Output

Under 200K tokens

$2.00/M

$0.50/M

$6.00/M

200K tokens and above

$4.00/M

$1.00/M

$12.00/M

The mechanic that catches people: once a request's prompt reaches the long-context threshold, long-context rates apply to all tokens in that request — not just the ones above 200K.

A 199,000-token prompt costs $0.398. A 201,000-token prompt costs $0.804. Crossing the line roughly doubles the bill for the whole request.

Grok 4.6's headline feature is a 500K window. Its headline price only applies below 40% of that window. Both things are true, and only one of them appears in the marketing.

Mitigation: monitor prompt size in your agent loop and compact or summarize before you cross 200K rather than after.

Other Cost Levers

  • Batch API — asynchronous processing at a discount, most requests completing within 24 hours, and batch requests do not count toward rate limits. Note: grok-4.6 is not in the discounted list at time of writing. Models without a listed discount have none.

  • Priority Processing — higher scheduling priority at 2x standard rates across all token types. You are only billed at the priority rate when the response confirms "service_tier": "priority".

  • Usage guidelines violation fee — $0.05 per request for violations caught before generation in the Responses API. Violations caught after generation are billed for the generation.

Worked Cost Examples

A 50-turn agent loop, 30K input and 2K output per turn, all under 200K:

  • Input: 1.5M tokens × $2.00 = $3.00

  • Output: 100K tokens × $6.00 = $0.60

  • Total: $3.60

Same loop with an effective prompt_cache_key and roughly 80% of input served from cache:

  • Cached input: 1.2M × $0.50 = $0.60

  • Fresh input: 300K × $2.00 = $0.60

  • Output: 100K × $6.00 = $0.60

  • Total: $1.80 — half the cost, one parameter

A document-analysis call with a 250K-token prompt and 5K output:

  • Input: 250K × $4.00 = $1.00

  • Output: 5K × $12.00 = $0.06

  • Total: $1.06 per call — versus roughly $0.53 if you had stayed under the threshold

Add ten web searches to that call: 10 × ($5 / 1000) = $0.05, plus tokens for the retrieved content.

Reasoning tokens bill at output rates and are not free. An xhigh request can generate substantially more reasoning than low for the same visible answer. Measure this on your workload before setting it globally.

Grok Build: The CLI Coding Agent

Separate from Grok Bot and separate from the API, Grok Build is xAI's coding agent. It runs as an interactive TUI, headlessly in scripts, or through the Agent Client Protocol in other apps. Grok 4.6 is its default model.

Install:

bash

# macOS / Linux / WSL
curl -fsSL https://x.ai/cli/install.sh | bash

# Windows (PowerShell)
irm https://x.ai/cli/install.ps1 | iex

Start a session:

bash

cd your-project
grok

On first launch it opens a browser for authentication. In non-browser environments, use an API key:

bash

export XAI_API_KEY="xai-..."
grok

Useful first prompts: Explain this repo. or @src/main.rs Walk me through this file.

Headless mode, which is what makes it scriptable:

bash

grok -p "Explain this codebase"
grok -p "Explain the architecture" --output-format streaming-json

Custom models. Grok Build supports any custom model via ~/.grok/config.toml (%USERPROFILE%\.grok\config.toml on Windows):

toml

[model.my-model]
model = "model-id"
base_url = "https://api.example.com/v1"
name = "Display Name"
env_key = "API_KEY"

[models]
default = "my-model"

Then grok inspect shows what Grok discovered in the current directory — config sources, instructions, skills, plugins, hooks, and MCP servers. Switch models with grok -p "Hello" -m my-model or /model <name> inside the TUI.

⚠️ Name collision warning. There are unaffiliated community projects on GitHub called grok-cli. Grok Build is the official tool and the grok command comes from the installer above. If you installed something else, none of this applies.

Grok Build access rides on an xAI or X plan — SuperGrok, X Premium+, or SuperGrok Heavy — or a pay-per-token API key for headless use. There is no free tier and no standalone Grok Build subscription. xAI does not publish a request quota for Grok Build at any tier, so any article stating a specific number invented it.

The API model behind Grok Build, grok-build-0.1, is separately priced at $1.00/$2.00 per million tokens with a 256K context — cheaper than grok-4.6 if you are building your own coding tooling and can accept the smaller window.

The Grok Consumer Plan Ladder

If you just want to use Grok rather than build on it, the plan structure is genuinely confusing, and it is confusing for a structural reason: Grok is sold on two storefronts with names that do not map to each other or to model versions.

Approximate US monthly list prices as of August 2026:

Plan

Approximate price

Storefront

Free

$0

grok.com, app, X

X Premium

$8

X

SuperGrok Lite

$10

grok.com

SuperGrok

$30

grok.com

X Premium+

$40

X

SuperGrok Plus

$100

grok.com

SuperGrok Heavy

~$300

grok.com

⚠️ Verify every one of these before you buy. These prices moved more than once during 2026, xAI does not list Heavy on its public individual pricing page, and community reports describe periodic promotional rates for Heavy well below list. Check the renewal rate, not the promo rate.

Two structural notes. First, tier names do not map cleanly to model versions, and new models roll out to lower tiers in stages — so "which Grok model does my plan get" is a separate question from "what does my plan cost," and the answer changes during rollout windows. Second, a SuperGrok subscription does not include API usage, and API usage does not require a subscription. They are entirely separate billing systems.

Since mid-2026, paid plans share a single weekly usage pool spendable across Chat, Imagine, Voice, and Build, with extra usage credits available on top.

What Grok 4.6 Is Actually Good At

Pulling the evidence together, here is where the model earns its place and where it does not.

Strong fit:

  • Knowledge work and research. The GDPVal-AA and AA-Briefcase results are the most convincing part of the eval table, and they measure long-horizon professional and analyst tasks.

  • Long-document analysis — with the 200K pricing cliff watched carefully.

  • In-editor coding assistance. CursorBench 3.2 at 69.9% is competitive with the frontier.

  • Cost-sensitive high-volume agentic work. At $2/$6 against frontier competitors charging several times more, the price-to-intelligence ratio is the strongest argument for the model.

  • Anything needing live X data. X Search has no real competitor.

  • First-pass generation of visual and interactive projects. xAI's claim that the model establishes structure and visual language in one pass is company-reported, but it is the kind of claim you can test in an afternoon.

Weak fit:

  • Autonomous software engineering. DeepSWE and Terminal-Bench are the benchmarks that predict this, and Grok 4.6 loses both by wide margins.

  • Anything where a confidently wrong answer is expensive. The 65.7% non-hallucination rate is the governing number here.

  • Workloads that genuinely need the full 500K window, unless you have modeled the doubled pricing.

  • Regulated or brand-sensitive deployments where vendor history is part of the review.

Common Mistakes to Avoid

❌ Assuming any SuperGrok tier gets you Grok Bot. Only SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium. SuperGrok at $30 and SuperGrok Plus at $100 do not include it.

❌ Skipping prompt_cache_key. You pay full input price on a cache-cold server. On a repetitive agent loop this is the difference between a $1.80 run and a $3.60 run.

❌ Ignoring the 200K long-context threshold. Long-context rates apply to every token in a request once the prompt crosses the line, not just the excess.

❌ Forgetting to enable search tools. Grok has no knowledge past February 1, 2026 without Web Search or X Search explicitly enabled.

❌ Using separate Bots as a security boundary. All your Bots share one computer, one set of browser sessions, and one set of command-line credentials.

❌ Automating before the one-time task is reliable. Build the task, save the skill, test it on a second input, and only then create a routine.

❌ Assuming deleting a Bot cleans up its access. It removes the profile, conversation, and routines. Files and logins on the shared computer remain.

❌ Writing broad Auto Review or event-trigger rules. "Allow everything in the browser" and "every new message" both fail in the same way — they widen the blast radius against an uncapped meter.

❌ Treating vendor benchmark tables as independent verification. xAI's table is xAI's table. Competitor figures are drawn from those vendors' own published cards and leaderboards.

❌ Piloting without instrumenting usage. There is no Grok Bot spend cap, no published allowance size, and no model picker. Your subscription is a floor, not a ceiling.

Troubleshooting

Grok Bot will not start after install. Most likely a Legacy Privacy Mode account. Grok Bot requires cloud data storage. Move to a supported Cursor data setting first.

The agent computer is unreachable. Use Recover computer from the error state first. Escalate through Settings → Beta in order: Update, then Recover, then Reset. Reset can discard recent unsaved work — it is the last resort, not the first.

A saved skill does not appear in the / menu. Open Settings → Plugins → Yours and enable it for the current Bot. Installed private skills are enabled per Bot.

Teach a task is not visible. The rollout may be gradual. Ask the Bot to create a skill from written instructions plus the completed task instead.

A routine silently produces stale output. You did not write a stale-data policy. Add an explicit instruction to report a failure rather than fall back to old data, then re-test.

The Bot keeps getting blocked at a login. Some sites expire sessions, enforce short timeouts, or re-request verification. Ask the Bot to pause and notify you rather than attempting to work around the check. Where a connector exists for that service, install it from Settings → Plugins instead of using the browser.

A workflow broke after a website redesign. Computer-use skills are brittle to UI changes by nature. Pause the routine, re-run the underlying task manually, update the skill, and re-test before re-enabling.

API requests time out on high or xhigh. Raise your client timeout. xAI's own examples use 3600 seconds.

API request returns an error mentioning penalties. presencePenalty, frequencyPenalty, and stop cannot be used with reasoning models. Remove them.

Auto-review rules are not applying on a second machine. Personal rules are stored on the current desktop and synced to its Grok Bot computer. Configure them separately on each installation.

How Grok Bot Compares to the Alternatives

The always-on agent category got crowded fast in 2026, and the honest comparison is about pricing model and maturity rather than raw capability, because nobody has enough independent reliability data yet.

Grok Bot — bundled subscription plus uncapped usage, three eligible high-tier plans, no free tier, no spend cap, no model picker, launched August 11, 2026 as an early beta with a launch-day documentation set and no revision history.

Claude Cowork — available on lower-priced Claude plans, a materially cheaper entry point for evaluating the "hand off multi-step work" pattern.

Self-hosted options like OpenClaw — infrastructure plus API costs, full control, more setup, no vendor plan gate.

Workflow automation platforms — n8n, Make, Zapier, and similar occupy an adjacent space: more deterministic, far cheaper at volume, considerably more setup, and they cannot handle a service with no API by clicking through its UI.

The genuine Grok Bot differentiator is the combination of a persistent shared cloud computer, computer use for API-less services, and Bot-to-Bot coordination in one product. That is a real capability set. It is also an early beta with an unpublished usage allowance, no spend cap, and a security model that depends heavily on you writing good boundaries.

If you want to evaluate it seriously, the honest recommendation is: pilot one Bot, on one read-only workflow, with instrumented usage, for two weeks, before you buy seats.

Verify These Before You Commit

Everything in this guide reflects a narrow window in August 2026, and several figures are explicitly marked by their sources as provisional. Confirm the following against primary sources before making a purchasing or architecture decision:

✅ Grok Bot plan eligibility and prices — check the live access page and Cursor pricing; these moved during 2026

✅ The weekly usage allowance size — unpublished at launch; ask before you buy seats

✅ Whether a Grok Bot spend cap now exists — none documented at launch

✅ Platform support — Linux desktop, Android, and iPad were unsupported at initial launch

✅ The 65.7% non-hallucination figure — one independent evaluator, awaiting a second measurement pass

✅ API pricing and the long-context threshold — check docs.x.ai/developers/pricing for the live rate card

✅ Batch discount eligibility for grok-4.6 — not in the discounted list at time of writing

✅ Consumer plan tiers and which model each gets — tier-to-model assignment changes during staged rollouts

✅ Your own eval results — run your actual workload before trusting any composite index score

Getting Started: A Practical Sequence

If you have read this far and want to actually deploy something, here is the order that minimizes wasted money.

If you are a developer: start with the API. Get a key at console.x.ai, run the minimal request, add prompt_cache_key immediately, enable Web Search or X Search if you need current information, and benchmark low versus high reasoning effort on your real workload before picking a default. Total cost to find out whether Grok 4.6 fits your problem: a few dollars.

If you want a coding agent: install Grok Build, point it at a repo you know well, and ask it to explain the architecture. You will learn more about the model's actual reasoning quality in twenty minutes of that than from any benchmark table.

If you want always-on agents: do not start with Grok Bot. Start by writing down one workflow you do weekly, in the five-part structure — outcome, sources, constraints, deliverable, review point. If you cannot write it down clearly, no agent will execute it clearly. Then evaluate whether $200 to $300 a month plus uncapped usage is the right vehicle for that one workflow, or whether a cheaper always-on agent or a deterministic automation platform gets you there.

If you just want to use Grok: the free tier on grok.com or X is genuinely useful, and the @grok reply bot costs nothing. Climb the plan ladder only when you hit a limit you can name.

The most useful frame for Grok 4.6 and Grok Bot together is this: xAI shipped a strong knowledge-work model at an aggressive price and an ambitious agent product at a premium one, one day apart, and marketed both as coding tools. The knowledge-work story is the one the evidence supports. Build accordingly.

Parash P

Content Creator

Creating insightful content about web development, hosting, and digital innovation at Dplooy.