agenticonsult logoagent i /consultDocs
Features

Voice — the hands-free assistant

Talk to your platform — realtime conversation, system-wide dictation, and a curated, pausable, audited tool reach across the whole desk.

Voice is the hands-free command surface for the whole platform: a multimodal assistant you talk to that can see your project, remember across sessions, generate media, and act on Command Center — open views, drive terminals, walk the knowledge graph — through a tool reach you curate and can pause at any moment.

Three surfaces share one engine:

SurfaceWhat it is
The Voice viewThe full studio — type or talk, switch providers, generate images and clips, manage conversations
The voice widgetA compact always-on-top orb summoned with Ctrl+Alt+V (Cmd+Alt+V on macOS) from anywhere — talk while any other app is focused
Global dictationHold Control+Shift+Space in any application, speak, release — your words land at the cursor, fully offline

Voice — the view and the widget — is an Ultra surface; see Tiers: Pro & Ultra. On Pro the view stays visible with a lock and an upgrade note. Global dictation is included in Pro: it shares the same local speech engine but is not part of the Ultra gate.

The provider reality — read this first

Voice's conversational AI does not run on your agent subscription. The agent fleet is bring-your-own — Claude is the documented path; Voice is its own provider, which you bring and connect:

  • Grok (xAI) — the full-capability provider: text chat, realtime spoken conversation, read-aloud, image and short-video generation, and web plus X search grounding. Connect with your own Grok subscription (an in-app sign-in) or an xAI API key (pay-as-you-go on xAI's rates).
  • ChatGPT (OpenAI) — text chat via your own OpenAI API key. Media generation and read-aloud run on the Grok side.

What costs nothing extra: the local half. Offline dictation, driving windows and terminals, reading your files, memory, and knowledge graph, and the grounding layer all run on your machine with no provider account at all. An unconfigured Voice is still a working dictation and platform-control surface; only talking to the AI needs the connection.

Set providers up in the view's Connections menu — sign in to Grok, or enter a key, and test the connection in place. Realtime spoken sessions are subject to your provider's own usage limits, like any voice product built on a third-party model.

Talk to it

In the Voice view, type or hold the mic. Conversations are named, searchable, and browsable by date, each with its own settings and system instructions.

  • Realtime conversation — natural, low-latency spoken back-and-forth.
  • Read-aloud — any answer spoken in a choice of voices; the live voice catalog comes from your provider, so it is what your account offers rather than a fixed list.
  • Dictation into the chat bar — three speech-to-text modes: an offline local engine (the default — no cloud involved), plus two cloud modes on the Grok provider.
  • Media in-conversation — ask for an image or a short clip and it generates inline (on the Grok provider).
  • Grounded answers — the assistant can search the live web and X during a conversation, and cites what it used.

Summon the widget with Ctrl+Alt+V (Cmd+Alt+V on macOS) when Command Center is not the focused app — same engine, floating on top, always listening while you work elsewhere. Global dictation goes one step further: hold Control+Shift+Space in any application — a browser, an email client, a document — speak, and release. It runs on a local speech model, entirely offline: no account, no cloud, no per-word cost. The chord, the model size and the paste behaviour are all set in Settings ▸ Speech-to-Text.

On macOS, the first dictation of a session must happen with the Command Center window on screen: macOS gates opening the microphone on window visibility, and with the window hidden the request is deferred rather than refused. Once the microphone has been opened, every later dictation works from anywhere, including with Command Center off screen. macOS also prompts for Microphone on first use, and for Accessibility the first time text is typed into another application — answer both while the window is visible.

What Voice can reach — and how you bound it

Voice's real differentiator is that it can act on the platform, through a deliberately curated bridge rather than open-ended access:

  • Project files — read, search, and (sandboxed) write within your workspace. Files matching the assistant's secret-file patterns — your environment file and the generated server configuration among them — are refused even inside the project, and outputs are size-capped. Treat that as a guard against the obvious mistake rather than an inventory: if you keep credentials in a file of your own naming, assume it is readable and keep it outside the workspace.
  • Memory and the knowledge graph — search shared memory, walk the graph, add findings. Ask "what do we know about X" and it answers from your own fabric, citing which entry.
  • The intelligence desk — query harvested items and reports (read-only).
  • Image generation — ask for an image and it dispatches the same image service the rest of the platform uses.
  • Your mailboxes — read, search, triage, move, stage drafts, and send. Sending is real, and the section below says what bounds it.
  • Terminals — drive the visible agent fleet by voice: "run the tests in my focused terminal" resolves which terminal from live focus, without you naming it. It can also create and arrange whole terminal workspaces the way you would.
  • Windows and the graph canvas — open views, arrange windows across monitors, steer the knowledge-graph canvas by voice.

The reach is governed, not assumed:

  • Read is broad; mutation is narrow and named. Destructive operations — deleting memories, knowledge-base content, or workspaces — are denied and never reach the assistant.
  • One pause switch halts every mutating tool action instantly, no restart; reads keep flowing.
  • Per-area switches disable a whole capability area — turn terminal reach off for Voice entirely, for example.
  • Everything is logged. Every write and every powerful action appends to an audit log you can read.

A per-conversation switch ("allow custom functions") decides whether the assistant has any tool reach at all, and the project brief that grounds it is a plain editable file — you can tune how Voice thinks about your system.

Voice can send mail — read this before you connect a mailbox

Voice is the one surface where a send is not staged behind a review queue. The conversational assistant can call send, reply and forward directly on a connected account. That was a deliberate decision, not an oversight: you are already in the conversation, and making you leave it to approve a message you just dictated would make the surface pointless.

What bounds it instead:

  • It is instructed to confirm first. Voice's brief requires it to state the recipient, the subject and the gist, and get an explicit yes, before any send. Drafting needs no confirmation. This is an instruction, so treat it as a strong convention rather than a wall.
  • The pause switch stops it. One click halts every mutating action across every area, including sends, with reads still flowing.
  • The per-area switch removes it. Turn the mail area off for Voice and the tools disappear from the assistant's catalog entirely.
  • The master stop is above both. Settings ▸ Terminals ▸ Agent control planes ▸ Email sending decides whether mail may leave this machine at all. See If Send stops working — it stops your own composer too.
  • Every send is in the audit log with everything else Voice does that changes something.

Voice cannot delete mail; it moves messages to Trash instead. And a conversation with tool reach switched off has no mail tools at all. See Autonomy & controls for how these switches sit together.

This relaxation is scoped to the conversation. Routines, the inbox watcher and the background email agent do not use Voice's bridge, so their own approval rules — the ones on the Email page — are unaffected by anything here.

What Voice does not do

  • It does not run on your agent subscription — conversational AI needs your Grok or OpenAI connection.
  • It does not delete anything irreversibly — deleting memories, knowledge-base content and workspaces is denied outright and never reaches the assistant, and mail can only be moved to Trash.
  • It does not touch other applications' windows — it maneuvers Command Center's own windows only, and it cannot close the main window.
  • The ChatGPT provider is text-first: spoken conversation, read-aloud, and media generation run on the Grok provider.

Where to go next

On this page