Three jobs, separate choices
Quick setup’s speech menu changes the shared speech model; there is no separate meeting-only speech override. A ready local speech runtime is reused. While a meeting holds it, competing jobs and model changes are guarded. Pause retains it; ending releases the meeting reservation without unloading dictation’s selected speech model.
Local meeting models
On-device insights use Qwen 3.5 4B (qwen3_5_4b_4bit, catalog estimate ~3.1 GB). This is independent of the model selected for dictation Polish and uses a separate inference instance.
For local speech and providers without speaker labels, Stenox loads Nemotron-3-Diarization-8bit for meeting-local speaker separation. It is not a dictation cleanup model and does not appear as an interchangeable speech model. Its state is scoped to one meeting and is released when that meeting finishes.
Initial preparation may download these files. No fixed RAM minimum or processing speed is guaranteed. Local speech, diarization, and insights together need more resources than speech alone. Turning automatic insights off reduces the work while retaining transcription and personal notes.
Cloud speech behavior
Deepgram, AssemblyAI, and Gemini use timed meeting transcription with speaker requests. Other adapters use their normal file transcription path, with local speaker handling where needed.- Deepgram follows the selected Nova model and language. Streaming-only Flux uses Nova 3 for meeting audio windows, as it does for file transcription. Requests opt out of model improvement.
- AssemblyAI uses Universal-3.5 Pro with its Universal-2 fallback and selected region/language. Stenox requests deletion of completed jobs.
- Gemini Transcribe uses Gemini 3.5 Transcribe with Interactions storage disabled.
Subscription insights
In Settings → Notetaker, choose Codex, Claude Code, or Google Antigravity. Use Sign in if needed, then choose an available model and reasoning level. Models come from the connected runtime; Stenox does not promise a fixed list or silently replace a missing saved model. Codex and Claude Code use bundled runtimes; they do not require a separate Node, Bun, or CLI installation. Existing eligible coding-runtime sign-ins can be reused. These integrations require subscription authentication, not an API key. Provider limits and usage apply. Google Antigravity requires its officialagy CLI installed separately. This adapter accepts 1.2 releases starting at 1.2.12. Its UI identifies the active CLI login source, not a verified email address.
Selecting another insights provider or turning insights off does not sign you out of the external provider. Stenox does not expose a separate Sign out or Disconnect control for these subscription accounts.
The session’s Insights model menu also lets you change the provider, model, reasoning choice, and Automatic insights setting during capture or while editing saved history. Switching cancels pending insight work and rejects stale results while capture continues. Existing insights and your edits remain. Returning to On-device affects future work; it cannot recall earlier transfers.
Context and failures
Your selected provider handles Google-linked and manual sessions. With a cloud provider selected, the inputs below leave your Mac; On-device processes them locally.
Personal notes guide focus in insights and refinement. Long notes use an excerpt based on a 6,000-byte UTF-8 limit plus an omission marker; the full saved note is unchanged. Notes do not establish facts, commitments, or owners. Editing or removing them invalidates stale results and refreshes the relevant provider context.
Titles and saved content can contain information derived from Google Calendar. Codex and Claude Code can continue provider context across Insights, refinement, and Ask for the same session and selection; that context can contain Personal notes from earlier turns. Provider caching is not guaranteed. Antigravity sends the selected context afresh each time. A request’s listed inputs do not imply that a continued provider session has forgotten previous inputs.
Provider account terms, usage, retention, and training settings still apply. Selecting an integration does not establish a zero-retention or no-training guarantee.
Invalid output, account errors, timeouts, or usage limits preserve recording and saved text. They may pause automatic reviews until an explicit retry. Generate insights remains an explicit action when automatic insights are off. There is no silent switch to another provider or paid API account.
An audio capture or processing overrun can end a session; completed text is preserved. Check the status, export what is available, and resolve the input or resource problem before starting again.

