agenticonsult logoagent i /consultDocs
Features

Voice — the hands-free assistant

Talk to your platform — realtime conversation, system-wide dictation, and a curated, pausable, audited tool reach across the whole desk.

Voice is the hands-free command surface for the whole platform: a multimodal assistant you talk to that can see your project, remember across sessions, generate media, and act on Command Center — open views, drive terminals, walk the knowledge graph — through a tool reach you curate and can pause at any moment.

Three surfaces share one engine:

SurfaceWhat it is
The Voice viewThe full studio — type or talk, switch providers, generate images and clips, manage conversations
The voice widgetA compact always-on-top orb summoned with Ctrl+Alt+V from anywhere — talk while any other app is focused
Global dictationHold a key in any application, speak, release — your words land at the cursor, fully offline

Voice — the view and the widget — is an Ultra surface; see Tiers: Pro & Ultra. On Pro the view stays visible with a lock and an upgrade note. Global dictation is included in Pro: it shares the same local speech engine but is not part of the Ultra gate.

The provider reality — read this first

Voice's conversational AI does not run on your Claude subscription. The agent fleet is BYO Claude; Voice is its own provider, which you bring and connect:

  • Grok (xAI) — the full-capability provider: text chat, realtime spoken conversation, read-aloud, image and short-video generation, and web plus X search grounding. Connect with your own Grok subscription (an in-app sign-in) or an xAI API key (pay-as-you-go on xAI's rates).
  • ChatGPT (OpenAI) — text chat via your own OpenAI API key. Media generation and read-aloud run on the Grok side.

What costs nothing extra: the local half. Offline dictation, driving windows and terminals, reading your files, memory, and knowledge graph, and the grounding layer all run on your machine with no provider account at all. An unconfigured Voice is still a working dictation and platform-control surface; only talking to the AI needs the connection.

Set providers up in the view's Connections menu — sign in to Grok, or enter a key, and test the connection in place. Realtime spoken sessions are subject to your provider's own usage limits, like any voice product built on a third-party model.

Talk to it

In the Voice view, type or hold the mic. Conversations are named, searchable, and browsable by date, each with its own settings and system instructions.

  • Realtime conversation — natural, low-latency spoken back-and-forth.
  • Read-aloud — any answer spoken in one of 5 built-in voices across 20+ languages.
  • Dictation into the chat bar — three speech-to-text engines: an offline local engine (the default — no cloud involved), plus two cloud modes on your provider.
  • Media in-conversation — ask for an image or a short clip and it generates inline (on the Grok provider).
  • Grounded answers — the assistant can search the live web and X during a conversation, and cites what it used.

Summon the widget with Ctrl+Alt+V when Command Center is not the focused app — same engine, floating on top, always listening while you work elsewhere. Global dictation goes one step further: hold the dictation key in any application — a browser, an email client, a document — speak, and release. It runs on a local speech model, entirely offline: no account, no cloud, no per-word cost.

What Voice can reach — and how you bound it

Voice's real differentiator is that it can act on the platform, through a deliberately curated bridge rather than open-ended access:

  • Project files — read, search, and (sandboxed) write within your workspace. Secret files are denied outright, even inside the project; outputs are size-capped.
  • Memory and the knowledge graph — search shared memory, walk the graph, add findings. Ask "what do we know about X" and it answers from your own fabric, citing which entry.
  • The intelligence desk — query harvested items and reports (read-only).
  • Terminals — drive the visible agent fleet by voice: "run the tests in my focused terminal" resolves which terminal from live focus, without you naming it.
  • Windows and the graph canvas — open views, arrange windows across monitors, steer the knowledge-graph canvas by voice.

The reach is governed, not assumed:

  • Read is broad; mutation is narrow and named. Destructive operations — deleting memories, knowledge-base content, or workspaces — are denied and never reach the assistant.
  • One pause switch halts every mutating tool action instantly, no restart; reads keep flowing.
  • Per-area switches disable a whole capability area — turn terminal reach off for Voice entirely, for example.
  • Everything is logged. Every write and every powerful action appends to an audit log you can read.
  • Outbound stays human-approved. Voice can draft an email or a newsletter; it cannot send one. Anything that reaches other people waits for your click in its review queue — the same operator-model gate as everywhere else.

A per-conversation switch ("allow custom functions") decides whether the assistant has any tool reach at all, and the project brief that grounds it is a plain editable file — you can tune how Voice thinks about your system.

What Voice does not do

  • It does not run on your Claude subscription — conversational AI needs your Grok or OpenAI connection.
  • It does not send anything outward on its own — drafts only, always.
  • It does not touch other applications' windows — it maneuvers Command Center's own windows only, and it cannot close the main window.
  • The ChatGPT provider is text-first: spoken conversation, read-aloud, and media generation run on the Grok provider.

Where to go next

On this page