Skip to content

Repository files navigation

Codex Inkbox Plugin

Codex, now with a phone



Give your Codex agent its own Inkbox identity:
a mailbox, iMessage, a phone number for calls and SMS, and an internet address.
Step away from the keyboard and keep working with it from anywhere.

Email · Calls · SMS / MMS · iMessage · Tunnel



Prerequisites

  • Codex installed and logged in. The bridge drives a real Codex session, so the codex CLI has to be on the machine and authenticated — install it (developers.openai.com/codex), then either sign in with a ChatGPT/Codex login or set OPENAI_API_KEY. inkbox-codex doctor checks for it.
  • Python 3.11+. The installer finds one and builds the bridge its own venv.
  • macOS or Linux. Boot persistence uses a systemd user unit on Linux and a launchd agent on macOS.
  • An Inkbox agent — nothing to set up in advance; the setup wizard self-signs up for you (or takes an existing API key).

Get started — one command

This finds a Python 3.11+, installs the bridge in its own venv, puts inkbox-codex on your PATH, and runs the setup wizard:

curl -fsSL https://raw.githubusercontent.com/inkbox-ai/codex-plugin/main/install.sh | bash

That's the whole setup. The wizard creates a fresh Inkbox agent for you (or takes an existing API key), provisions a phone number, connects iMessage, mints a webhook signing key, picks the project directory Codex works in, and offers to keep the bridge running on every boot. When it finishes, text/email/call your agent and it answers from a real Codex session.

The one thing to have ready: be logged into Codex — a ChatGPT/Codex login (via the Codex app/CLI) or OPENAI_API_KEY set. The installer checks this and warns if it's missing.

Flags: --start (launch the background gateway when done), --no-setup (install only). From a local checkout, run ./install.sh. Re-running is safe.

Bootstrap an existing identity without prompts

For unattended agent setup, install without opening the wizard and pass the API key through the environment (or standard input), never a command-line argument:

curl -fsSL https://raw.githubusercontent.com/inkbox-ai/codex-plugin/main/install.sh | bash -s -- --no-setup
export INKBOX_API_KEY="ApiKey_..."
inkbox-codex bootstrap --identity my-agent --project-dir "$PWD" \
  --voice-ai --rotate-signing-key --start-gateway
unset INKBOX_API_KEY

bootstrap validates that the key can access exactly the requested identity, scopes down an admin key before saving it, preserves existing Voice AI settings, enables native Inkbox tool approvals, and starts or restarts the detached gateway. Signing-key replacement is opt-in because it transfers verified webhook delivery away from any gateway using the previous key. The command prints a secret-redacted JSON result and is safe to resume.

Check it any time:

inkbox-codex doctor    # config, codex CLI/auth, identity reachability
inkbox-codex status    # is the background gateway up? where are the logs?

What it does

you (phone)  ── SMS / iMessage / email / call ──▶  Inkbox  ──▶  tunnel  ──▶  bridge
                                                                              │
                                                                              ▼
                                                                  Codex session
                                                                  (full tool access in
                                                                   your project dir)
  • Text, iMessage, email, or call your agent's Inkbox number. Each remote party gets one Codex session spanning every channel — text it on the walk home, then email it details, same conversation.

  • Codex runs with full tool access in CODEX_PROJECT_DIR. It reads, searches, and browses freely; anything risky (running commands, editing files) is escalated to you as a text:

    Codex wants to run the command: npm test

    Reply 1 (or YES) to allow once, 2 (or SESSION) to allow for this session, 3 (or NO) to deny. /stop cancels the task.

  • When Codex needs you to pick between options (the AskUserQuestion tool), you get a numbered poll on whatever channel you're on, and your reply is fed back as the answer.

  • Each message you send is tagged with its channel, so Codex knows whether it's on SMS, iMessage, email, or a call.

  • A channel prompt is appended to Codex's system prompt so replies fit a phone: plain text, no markdown, short, jargon kept to a minimum ("saved and published the change", not "pushed to origin/main").

  • Codex also gets Inkbox tools (inkbox_send_email, inkbox_send_sms, inkbox_send_imessage, …) so it can proactively reach you — "email me the full report" works.

Manual install

If you'd rather not run the installer (any Python 3.11+ environment):

pip install -e .

inkbox-codex setup    # interactive wizard — writes .env for you
set -a; source .env; set +a

inkbox-codex doctor
inkbox-codex run

inkbox-codex setup walks you through everything and writes .env: create a fresh Inkbox agent via self-signup (or bring an existing API key), pick or create the identity, attach the Codex avatar, provision a phone number, wait for your START opt-in, choose a phone-call voice stack, connect iMessage, mint a webhook signing key, choose the project directory, choose whether to trust Inkbox MCP tools without repeated allow prompts, and set up autostart. Realtime keys are validated before selection is saved. An admin key used to change Voice AI authority is held only for that setup run and is never written to .env. Rerun setup anytime to reconfigure.

Local Docker test image

The repository includes a manual-test image with Codex, the plugin, and the Inkbox SDK preinstalled. It reuses the host's native Codex login and keeps Inkbox plugin state in a named volume; neither is copied into the image:

docker build -t inkbox-codex-local .
docker run -dit --name inkbox-codex-local \
  -v "${CODEX_HOME:-$HOME/.codex}:/root/.codex" \
  -v "$PWD:/workspace" \
  -v inkbox-codex-state:/root/.inkbox-codex \
  inkbox-codex-local
docker exec -it inkbox-codex-local bash

# inside the container
inkbox-codex setup
inkbox-codex doctor
inkbox-codex run

# after leaving the container
docker rm -f inkbox-codex-local

The image contains no credentials. If you use API-key authentication instead of codex login, add -e OPENAI_API_KEY="$OPENAI_API_KEY" to docker run. Inkbox credentials are entered into setup or supplied only at runtime.

On startup the bridge opens an Inkbox tunnel, wires mail/text/iMessage webhook subscriptions and the incoming-call channel to it, and routes everything into Codex sessions.

Running it

inkbox-codex run        # foreground (Ctrl+C to stop) — good for first runs and debugging

Or run it as a background daemon (PID + log under ~/.inkbox-codex/):

inkbox-codex start      # detach and run in the background
inkbox-codex status     # is it running? where are the logs?
inkbox-codex restart    # restart it
inkbox-codex stop       # graceful stop (SIGTERM, then SIGKILL after 5s)

tail -f ~/.inkbox-codex/gateway.log

start auto-loads .env from the current directory, so you don't have to source it first. run is the foreground version a service manager (systemd, Docker) should supervise; start/stop are the self-contained background option.

Start on boot

The setup wizard offers to keep the bridge running for you — either just in the background for this session, or as a service that starts on every boot. On Linux it installs a systemd user unit (~/.config/systemd/user/inkbox-codex.service) and enables it; on macOS it installs a launchd agent. To keep a Linux service alive while you're logged out, enable lingering once:

sudo loginctl enable-linger "$USER"
systemctl --user status inkbox-codex   # restart | stop | status

Uninstall

inkbox-codex uninstall           # stop it, remove the boot service + launcher; keep config
inkbox-codex uninstall --purge   # also delete ~/.inkbox-codex (config, logs, sessions)

This is local-only — webhook subscriptions on the Inkbox side are left as-is; remove them in the Inkbox Console if you want.

Then, from your phone:

  1. Text START to the agent's number (first time only, carrier opt-in).
  2. Text it something like "clean up the TODOs in the auth module".
  3. Approve the permission texts as they arrive. Get the result as a text.

How escalation works

The bridge honors Codex's configured approval policy and relays its requests over your active Inkbox channel:

  • Commands, file changes, permission-profile changes, and request-user-input prompts block the agent mid-turn while the bridge texts you a one-line plain-language summary.
  • MCP tool approvals number the available choices consecutively; reply with the number shown or a word from the table below. A different instruction cancels the pending approval turn and becomes fresh work instead of being consumed as an approval answer.
  • Request-user-input prompts are formatted as numbered options; reply with the number or free text.
  • No reply within INKBOX_PERMISSION_TIMEOUT_S (default 10 min) cancels an MCP request; command/file approvals are denied.
Reply Decision
YES Allow once
SESSION Allow for this session, when offered by Codex
NO Deny this request
ALWAYS Remember approval across sessions, when offered by Codex
/stop / /cancel Cancel the entire task

For example, a tool offering only allow-once and deny displays 1 — Allow once and 2 — Deny. If session approval is also available, it displays 1 — Allow once, 2 — Allow for this session, and 3 — Deny. Always follow the numbers in the current prompt; words such as NO keep the same meaning.

Session and permanent approval are distinct: an unavailable scope is never silently changed to allow-once. Advertised MCP persistence scopes are forwarded through Codex's native protocol, not a bridge-wide allow list. Other tools still require their own approval when Inkbox tool auto-approval is enabled.

Clear replies such as “yes, proceed” are accepted. A natural cancellation such as “please cancel my request” stops the task. Companion answers and controls retain sender and mention gates. Structured MCP questions show their actual fields and return typed answers; unsupported forms require using Codex directly or canceling the interaction. URL-based requests require completing the displayed URL step before replying DONE.

Slack preview

Set INKBOX_SLACK_ENABLED=true to add Slack to the gateway. Inkbox SDK 0.7.11 includes the required Slack methods; the API must also support Slack. Connect a workspace to the same Inkbox identity to exchange messages; the receiver can start before installation. Run inkbox-codex setup to opt in interactively. The wizard offers Connect Slack now?, selects a saved provisioning workspace, waits for app preparation, prints an installation link to open in your browser, and polls until that workspace is connected. Keep the link private and complete authorization in the same browser. This flow requires an SDK exposing client.slack.list_provisioning_workspaces and identity-wide webhook subscriptions. Setup uses the bridge's existing claimed agent key for workspace credentials, app preparation, and installation; no extra admin key is needed. A newly self-signed-up identity must be claimed first. An organization admin key is also supported when already configured for the bridge.

If you need to add a provisioning workspace or renew its credentials, the wizard prompts for a masked Slack app-configuration access/refresh token pair from Your App Configuration Tokens. Inkbox verifies and stores the pair for your organization; the bridge does not save it locally. If the pair is rejected, setup explains the credential issue and offers another attempt without leaving the wizard. Other errors identify the failed step and HTTP status when available; they do not silently retry app creation or expose tokens. These are not bot tokens. Each identity's app is permanently bound to one workspace. Existing connections need no new installation; manage permission refreshes in the Inkbox console. After installation, add the bot to the channels where you want to use it.

Existing connections are shown on reruns. You can skip connecting, press Ctrl+C during either wait, or rerun after the five-minute wait expires. Rerunning setup and declining full reconfiguration still offers Slack onboarding. Declining Slack turns it off only in this bridge; it does not disconnect the workspace or disable Slack for other clients. On startup, the gateway registers the subscriptions below.

Work indicators. Accepted threaded Slack requests use the native agent loading indicator, without adding reaction bubbles. It stays while work is queued or running, switches to awaiting-input during questions or approvals, and returns to ready after completion, failure, or cancellation. Overlapping requests in the same thread share one indicator. Mention-mode context does not start it. Slack's native Stop button cancels work in the matching engaged thread, subject to the same allowed-sender policy. Ordinary main-DM messages receive replies in the main DM and share conversation context, without a native status indicator. Instead, the exact DM message receives 👀 while queued, running, or waiting for input. It is removed after success or cancellation; a failed turn or reply delivery replaces it with ❌. These reactions do not start a thread. An explicit native @mention starts a thread beneath that message; messages already inside a thread stay there. Threaded replies and their indicators use the same thread. Continue there for follow-ups, approval answers, and commands; each thread keeps its own saved context. Native status needs an app declared as a Slack Agent (including assistant:write), chat:write, SDK processing-status support, and workspace feature availability. Existing installations may need an app-configuration update and reauthorization. Unsupported or uncertain updates are logged without blocking replies or falling back to reactions. Restart cleanup clears unfinished native status, marks interrupted DM work with ❌, and removes pending indicators from the previous reaction version. Reaction failures do not block replies.

No separate Slack bot token is needed at runtime. Slack subscriptions use slack.dm_received, slack.group_dm_received, slack.mention_received, slack.thread_reply_received, and slack.session_stopped with no Slack-specific filter. The gateway also recognizes slack.channel_message_received, while its local attention rules keep unrelated channel chatter from waking Codex. Update the API and receiver together when upgrading from the older incoming-message preview event. Incoming messages can match several categories; the payload's message_kinds retains the full classification even when only one event type is selected. Replies remain deduplicated by the stable event ID, not the selected category.

  • By default, DMs, group DMs, and mentions wake Codex. New requests start a thread; follow-ups in an already-engaged thread continue its conversation without another mention. Unrelated channel messages and bot messages do not wake it.
  • With INKBOX_GROUP_REPLY_MODE=mention, Slack channels, group DMs, and their threads require a native @mention of this agent on each new message to generate a reply, even after the agent has joined the thread. Unmentioned follow-ups are context only: no reply, tool execution, or interruption of an active turn. Direct DMs are unchanged. Bridge commands and the requesting sender's answers to pending approval/question prompts still work without a mention.
  • Each workspace/conversation/thread gets a separate, resumable Codex session. When supplied by Inkbox, the sender's linked contact ID and Slack profile fields are included as context. Email and phone may be absent. The bridge does not match contacts itself, merge conversations, or grant permissions from profiles.
  • Replies and approval prompts stay on the originating Slack route. Only the requesting sender can answer a pending group-thread approval.
  • Messages are limited to 12,000 characters per send. The bridge rejects longer replies before sending; it does not split or automatically retry them.
  • Six tools list workspaces, list conversations, read messages/threads, search retained text, send explicitly requested messages, and inspect send outcomes. Ordinary final replies are sent automatically; do not send them again by tool.
  • Unconfirmed sends are logged with their action ID when available, never automatically resent. Reuse the exact idempotency key when retrying an explicit send. Webhook duplicate suppression follows the gateway's existing bounded, in-memory behavior.
  • Attachments are exposed as metadata in this first version. File transfer and arbitrary reaction tools are not included.

INKBOX_ALLOWED_USERS accepts Slack user IDs or workspace-qualified T_ID:U_ID entries. An empty list admits all human senders whose events reach the identity. Slack subscriptions are additive: starting this gateway does not remove another Slack receiver. Existing active subscriptions at this receiver's URL, including mixed-event subscriptions, are reused; only missing events are registered. Paused overlapping subscriptions require attention in the console before startup. Stop/remove any previous Slack receiver if it should no longer reply.

Isolated Slack harness

The harness uses the same gateway, sessions, signing verification, tools, and reply path. It registers only Slack and does not change email, phone, iMessage, or A2A subscriptions. Credentials are read into the process from the selected file, never copied into its state directory. Use a separate directory and port from your regular gateway. Run only one gateway on an identity’s tunnel at a time, or pass --public-url for a separately reachable receiver.

After installing this checkout and a Slack-capable SDK, run a read-only preflight:

python -m inkbox_codex.slack_harness \
  --credentials-file ~/.env --api-key-env INKBOX_API_KEY \
  --identity "$INKBOX_IDENTITY" --base-url "$INKBOX_BASE_URL" \
  --state-dir ~/.inkbox-codex-slack --project-dir "$PWD" \
  --signing-env-file /path/to/identity.env

Add --run to connect the identity’s existing tunnel and register the signed Slack receiver. The signing file must contain the identity's existing INKBOX_SIGNING_KEY; the harness never creates or rotates keys. --allow-user T_ID:U_ID can be repeated. Ctrl+C stops the local receiver; its subscription remains available for the next run. --group-reply-mode mention enables mention-only group/thread replies in the harness; it otherwise follows INKBOX_GROUP_REPLY_MODE, defaulting to auto. The webhook still includes thread messages so quiet context and approval answers can reach the session.

In mention mode, test a native @mention, an unmentioned follow-up in the same thread, and another @mention: expect reply, silence, reply, with the quiet message available as context for the last reply.

For manual acceptance: DM the agent, mention it in a channel, continue in that thread without a mention, ask for recent thread history or retained-text search, then exercise an approval and /status. Replies should stay in the right thread; unrelated channels and bot messages should remain quiet. The preflight does not send messages. Live delivery can only be verified after connecting the workspace and starting the harness.

Offline regression harness: python -m pytest -q tests/test_slack.py tests/test_slack_harness.py.

Sessions

Direct-message sessions are keyed by Inkbox contact, so one person = one conversation across channels. Group SMS messages share a session keyed by the group conversation, separate from direct messages and other groups, while retaining each sender's contact details. Codex session ids are persisted in ~/.inkbox-codex/sessions.json and resumed across bridge restarts — your conversation picks up where it left off. Replies go out on the channel you last used. If a voice call ends before Codex finishes a voice reply, that late voice reply is dropped instead of silently switching to SMS or email.

Group replies. The setup wizard offers Automatic (default) or Mention required for group SMS, iMessage, and Slack, saved as INKBOX_GROUP_REPLY_MODE=auto|mention. Automatic keeps the existing behavior: the agent decides whether to answer. For SMS/iMessage, mention mode starts a reply only when the new message itself includes @agent or @<agent-handle> as a whole mention, case-insensitively; older messages, links, and email addresses do not count. Slack requires a native @mention of the agent, including on follow-ups in an engaged thread. Other messages and group reactions are added to Codex's context without generating a reply or a typing indicator. While a turn is running, background messages wait until it finishes before being appended. Appended context persists with the Codex thread; messages still waiting in the bridge's queue are not persisted across a restart. Direct messages are unchanged. Commands such as /stop, and answers to the agent's pending questions from the sender it asked, do not require a mention.

Companion conversations (preview)

Verified SMS/MMS, iMessage and email webhooks may include a top-level companion object. Webhooks without it (or with null) retain normal routing. The bridge loads the full authorized initialization snapshot using the Inkbox SDK, including every page, and submits one combined input. The sponsor trigger is included once; historical messages, attachment descriptors, and history notices remain conversation data, never individual turns or slash commands.

The setup wizard offers two independent settings, including when rerun for an existing identity:

  • Companion responses: INKBOX_COMPANION_RESPONSE_MODE=safe|relaxed. Safe (default) allows only messages with sender_access="direct" to wake the agent. Sponsored messages, or messages with missing/unrecognized access, are appended to context without model generation, typing, or a reply. Relaxed allows any delivered message to wake the agent, including sponsored and unknown-access messages.
  • Group replies: INKBOX_GROUP_REPLY_MODE=auto|mention, also applied to Companion email. Automatic lets an eligible message start a turn; the agent decides whether a reply is warranted. Mention required additionally requires @agent or @<agent-handle> in the current message's own text. For Companion email, putting the agent's mailbox in the current message's To recipients also counts as a mention. Address matching is case-insensitive and supports display names; Cc/Bcc alone do not count. Quoted headers, historical mentions, notices, and attachment metadata do not count. Reply-all may retain the agent in To and therefore satisfy this gate again.

The signed webhook supplies sender_access on data.text_message for SMS/MMS or data.message for iMessage/email. direct means contact rules permitted the message without sponsorship, including allowed-by-default senders; sponsored means delivery was authorized through sponsorship. Neither value grants command permissions or establishes permanent trust. Access is not inferred from the sender's contact record, sponsor identity, or Companion phase.

Wake below means context plus a model turn; Context means context only. These rules apply equally to SMS/MMS, iMessage, and email Companion inputs.

Companion mode Group replies Direct, no mention Direct, mention Sponsored, no mention Sponsored, mention
Safe (default) Auto Wake Wake Context Context
Safe (default) Mention Context Wake Context Context
Relaxed Auto Wake Wake Wake Wake
Relaxed Mention Context Wake Context Wake

For email, the table's mention includes the agent being in To. This does not override sender access: sponsored and unknown-access messages remain context-only in Safe mode even when addressed To the agent. Auto is unchanged.

Context-only messages are retained for later eligible turns. For initialization, the entire snapshot is appended together; only the current received message can trigger generation. Webhooks without Companion metadata retain normal routing and are unaffected by the Companion response setting.

  • Sessions are isolated by API environment, identity, channel, server scope, and activation, separate from ordinary contact sessions. A new activation starts a fresh session with its authorized snapshot. Each new released snapshot must have a distinct activation ID; replaying an immutable activation is not a new batch.
  • Initialization completes before queued live messages run. A first-seen live event loads history as context, without a separate historical sponsor turn; only the live message can wake the agent. If that message is already in the snapshot, the combined input uses its access and mention instead of repeating it. Pending events run in sequence order; numeric gaps do not stall the conversation. Late unseen events older than an already submitted input require reconciliation instead of running out of order. phase="ordinary" uses a separate scoped session and never loads hidden history.
  • Replies use the original group conversation ID, or the SDK-approved email reply context's stored message UUID with canonical reply-all. They never fall back to privately messaging the latest author. Local sponsor restrictions remain in force. Live turns and replies use the signed conversation scope and saved sponsor message without additional activation lookups. Replies reuse the identity loaded at startup.
  • Only a live reply by the prompted, locally allowed sender that passes both response gates can answer an approval request. Historical or context-only approval-looking text cannot. Sponsor slash controls also require both gates. In Mention mode, address email To the agent or prefix answers and controls with @agent, for example @agent allow or @agent /stop; escalation prompts remind you of this.

SDK requirement: The bridge requires Inkbox SDK 0.7.11 or newer (below 1.0.0), including client.companion.load_initialization and activation_messages. Installation resolves the published SDK; CI tests the released inkbox==0.7.11 minimum across unit, real-host, and live-channel lanes. No preview source checkout is needed. An unsupported SDK produces an explicit webhook error; the bridge never falls back to submitting only the trigger.

Wake decisions use the signed current-message access and do not require extra SDK lookups. Older SDKs may omit per-entry access labels from rendered history; those historical entries remain context, never independent triggers. Safe mode on webhooks that omit sender_access is context-only; no missing field is treated as direct access.

Delivery and recovery: receipts and initialization checkpoints live under $INKBOX_CODEX_HOME/companion/ (default ~/.inkbox-codex/companion/), with private permissions and one receiver owner per environment/identity. Transient reads before submission retry with capped backoff while the gateway runs; other pre-submission failures retry up to five times. Pending inputs resume after restart or webhook redelivery. Failed Codex startup or resume attempts release the incomplete client and retry with the same saved thread; they do not count as submitted model turns. Closing or clearing a session also disposes its connecting client, so a late startup response cannot restore a cleared conversation. A retry may refresh its delivery timestamp or inline history preview without creating another input. Stable event IDs and acknowledged snapshot/live source-message IDs are deduplicated across restarts within their scope and activation. Completed answers are saved before delivery. Transient read failures while preparing delivery retry with capped backoff, including after restart, without rerunning the model. An interrupted host submission or uncertain reply send pauses that scope rather than risking another model turn or duplicate send. The log identifies the retained receipt. Inspect the scoped Codex thread and channel delivery before operator recovery; do not delete receipts or blindly replay uncertain events. This is at-most-once automatic retry behavior at ambiguous boundaries, not an exactly-once transport claim. Initialization above 8 MiB fails explicitly rather than truncating. The preview uses POSIX file locking (Linux/macOS).

Run inkbox-codex setup to change the choice without reconfiguring your identity, or edit .env, then restart the bridge to apply it. For background mode, use inkbox-codex restart; for a systemd installation, use systemctl --user restart inkbox-codex.service. Mention mode requires a Codex version supporting thread/inject_items.

Typing indicator. While Codex works on a turn, the bridge keeps a typing indicator alive on your iMessage thread (refreshed every few seconds, since it expires) so you can see it's busy. SMS, email, and voice have no typing indicator. Slack threads use the native agent status described above.

Delivery failures. Outbound messages can silently fail — a carrier filters an SMS, an iMessage is declined, an email bounces. Inkbox reports these asynchronously (text.delivery_failed, imessage.delivery_failed, message.bounced/message.failed). The bridge catches them and wakes the affected contact's session to tell Codex which message didn't land and why, so it can retry or reach you another way (a different channel, or a call) using its Inkbox tools. The notice runs as a side-effect turn — Codex acts via tools rather than replying on the channel that just failed — and repeat webhooks for the same message are de-duplicated so it can't loop. text.delivery_unconfirmed is different: it only means the carrier couldn't confirm delivery (the message usually landed), so it's logged for debugging without waking Codex — waking there would resend a message that was likely delivered.

Interrupt by texting again. By default, messaging the agent again while it's mid-turn works like pressing Esc in Codex and typing a new message: the running turn is interrupted, its partial answer is dropped, and Codex picks up your new message instead. (A reply while it's waiting on a permission/poll still answers that escalation — interrupting only applies while it's actively working.) The opt-in iMessage behavior below queues follow-ups instead.

Native iMessage replies (opt-in preview)

Set this in the running gateway's environment and restart its existing service:

INKBOX_IMESSAGE_THREADED_REPLIES=true

The default is false; there is no new setup-wizard step. Enabling requires the SDK's native reply/thread-read support (install Python SDK 0.7.13 or newer) and an API environment serving those endpoints. doctor distinguishes local SDK support from backend capability; a targeted reply also checks the native endpoint before sending. The bridge fails explicitly if support is missing, rather than silently sending an unthreaded reply. The existing SDK minimum is unchanged when disabled.

  • Several quick messages: ordinary iMessage inputs from the same sender, conversation, and reply context collect for 750 ms of quiet, up to 2 seconds. Codex receives their text and source IDs in order. Different senders or explicit native threads stay separate. Companion inputs retain their existing ordered, one-event-per-turn receipts and activation boundaries; they are not coalesced.
  • New requests during work: follow-ups queue behind the current turn; they do not interrupt it. This preview does not start parallel model runs or child agents. Timing does not change the reply target. An answer already saved for delivery keeps its original route; it is not resent or retargeted.
  • Where answers appear: every message-triggered answer replies to its source, including isolated standalone messages. A combined burst's default answer targets its first source. The bridge owns this decision; the model cannot select a different target or turn threading off for an answer. Same-conversation send tools use the same trigger. After an explicit send, Codex returns [SILENT] to avoid another automatic answer. A final answer exactly matching an observed targeted tool send is suppressed too.
  • Proactive messages: cron jobs, reminders, and other sends without a current message trigger start fresh, without a reply target. The bridge never picks the last message from conversation history. Scheduled jobs must run independently, without inheriting the active chat's tool-process environment.
  • No audience changes: targeted tools must use a visible source in the same conversation; in a running chat they must also belong to that active input. Opaque native thread IDs are not message IDs or new Codex sessions. Replies use the API's message IDs; missing parent/root IDs remain unknown, never inferred from a thread ID. A visible ancestor root can identify a reply even when its immediate parent is unavailable. Native thread-read tools return one bounded chronological page and do not broaden Companion activation history.
  • Fallback and failures: plain_reply_fallback=true delegates an eligible same-conversation fallback to the API. A timeout or arbitrary send failure never triggers a client-side unthreaded resend. Accepted sends are queued, not proof of delivery. Correlated late failures are retained and added as context to the original session without waking a retrying model turn; /health reports counts.
  • Restart behavior: ordinary inputs are saved before acknowledgement. Only known-unsubmitted inputs and saved, not-yet-sent answers resume automatically. Interrupted model work and uncertain sends remain retained for inspection, not replay. Do not delete receipts or resend blindly. /stop cancels this iMessage conversation's queued work without cancelling a voice consult on the same contact.

Set the variable back to false and restart to restore the previous behavior and tool schema. This does not erase retained receipts; inspect unfinished work before re-enabling. Identity, contact rules, approvals, other channels, and voice settings are unchanged.

Control commands. A handful of slash-commands steer the conversation itself and are handled by the bridge instead of being sent to Codex (works on any channel):

  • /clear (or /new) — start a fresh conversation: forgets the resumed session, tears down the client, and clears session-scoped permission grants.
  • /stop (or /cancel) — interrupt the current turn and drop anything queued, keeping your conversation context intact.
  • /resume — texts you back a numbered list of recent Codex conversations (each with a short summary and timestamp); reply with a number to reopen that one. Like /resume in the Codex CLI.
  • /status — reports what the bridge is doing for you right now (working, waiting on a reply, or idle) and whether you're in a fresh or ongoing conversation. Read-only; doesn't disturb a running turn.
  • /usage — reports Codex rate-limit windows and token summary from app-server account endpoints.
  • /health — reports identity reachability, tunnel connectivity, the latest Codex startup check, and Companion queue readiness.

These match only when the whole message is exactly the command, so "please /clear the cache" is still a normal turn.

Errors. Ordinary turns send a short plain-language failure notice. Companion failures are recovered automatically without replaying possibly accepted work; see below.

Automatic recovery and readiness

The bridge automatically retries inputs that failed before a model turn or message send could start, with backoff capped at one minute. Retry does not stop after a fixed number of failures, and a recovered Codex runtime wakes waiting conversations without requiring another message or a restart. Saved answers retry delivery without running the model again.

If the selected npm Codex launcher cannot find Node, the bridge uses the matching native executable already included in that same Codex installation. It does not change your configuration, download a replacement, or select an unrelated installation. Explicit native CODEX_BIN paths and working launchers are kept.

For an ambiguous turn outcome, the bridge stops the old session and checks saved Codex history for a completed turn matching its receipt. A positively matched answer can be delivered without repeating the model turn. When the outcome cannot be established—or delivery may already have happened—the original input is retained as unconfirmed and is not automatically replayed. Later messages continue in the conversation; the next reply can explain the earlier uncertainty. An interrupted initialization does not replay its old trigger. Payloads, deduplication records, and saved conversation threads are preserved across restart.

Outstanding approval prompts are canceled when their Codex connection closes or their host request is resolved. Only one question is displayed at a time, and canceling a task clears its queued questions. New instructions during a tool approval become fresh work. Trusted Inkbox tool auto-approval remains optional.

inkbox-codex doctor checks the effective Codex launcher with a bounded app-server initialization handshake, using the same saved environment as the gateway. Recognized startup failures include an exit code and a bounded diagnostic summary, without raw stderr. HTTP /health remains a liveness endpoint (ok: true, HTTP 200) with separate ready, codex, and companion fields. HTTP /ready returns 503 while startup cannot be verified or a conversation is still blocked.

Startup checks refresh in the background every minute; HTTP health requests do not start new processes. Queue diagnostics distinguish active blockage from retained unconfirmed outcomes (quarantined_count), which do not block new input. They also show unfinished counts and oldest receipt age (unknown for older receipts). These checks do not prove model generation or end-to-end delivery. Verify recovery with a fresh message and its reply in the same conversation.

Optional receipt inspection

Automatic recovery does not require an operator command. To inspect retained receipts without message content, use inkbox-codex inbox list. After inspecting an unconfirmed outcome, an operator can explicitly request retry or retirement with the gateway stopped:

inkbox-codex stop
# Choose one action for the receipt:
inkbox-codex inbox recover EVENT_ID --action retire --reason "Inspected; skip stale input"
# OR acknowledge that explicitly retrying uncertain work may duplicate effects:
inkbox-codex inbox recover EVENT_ID --action retry --reason "Inspected; retry intended" --acknowledge-duplicate-risk
inkbox-codex start

Each explicit decision records its reason, timestamp, and previous/new state. Retirement does not claim that processing or delivery succeeded; retry uses a saved answer when available. Recovery commands refuse to race a running gateway.

Voice

The setup wizard has a Phone call voice stack section with three choices:

  • Inkbox Voice AI: Inkbox handles the audio and conversation on Codex's behalf. Choose contact-scoped or YOLO authority during setup. Hosted outbound calls carry a task reason, inherit that saved authority by omitting a per-call override, and notify Codex through a signed call.ended event. The initiating turn ends after call placement; the separate call-ended turn owns the follow-up. Codex then fetches the authoritative transcript and executes any remaining post-call commitments in a side-effect-only turn; its plain model prose is never sent after hangup. Saved SMS receipts prevent a correction from repeating an accepted or uncertain send, even when Codex omits the tool result from its summary.

  • OpenAI Realtime (when configured): the bridge pre-opens an OpenAI Realtime session and accepts the call in raw-media mode, so a natural, low-latency voice handles the conversation. It runs the call itself and has these tools:

    • consult_agent — do real work now in the project; runs in the same contact-keyed session as your SMS/iMessage and its answer is spoken back.
    • register_post_call_action / edit_post_call_action / delete_post_call_action — queue, change, or cancel work to run after you hang up.
    • hang_up_call — two-step (say goodbye, then end the call).

    When the call ends, queued actions run in your session (and any plain "reflect on the call" follow-up if none were queued) — so "after we hang up, open a PR and text me" actually happens. Enable it in inkbox-codex setup (it validates your OpenAI key live) or via the INKBOX_REALTIME_* env vars below.

  • Inkbox STT/TTS (default): Inkbox auto-accepts the call and opens a WebSocket to the bridge; finalized transcripts become turns in your same session and Codex's replies are spoken back. Realtime may fall back to this if OpenAI cannot be reached (unless INKBOX_REALTIME_FALLBACK_TO_INKBOX_STT_TTS=false).

Two calling lines

Calls — inbound and outbound — can run over either of two lines, and the agent picks the one that matches the channel it's talking on:

  • The dedicated phone number. The agent's own number (the same line SMS uses). Outbound calls present this number; inbound calls to it ring the agent.
  • The shared Inkbox iMessage line. The agent can also place and receive voice calls with a person it's connected to over iMessage, over the same shared line that person already messages. The underlying number is never surfaced — Inkbox resolves it from the iMessage connection — and it only works for people already connected over iMessage (an unknown caller is rejected; an outbound call with no connection is refused).

Inbound answering is configured once per identity: Voice AI uses hosted_agent; local Realtime and TTS/STT use auto_accept plus the bridge WebSocket. Outbound, the agent sets origination on inkbox_place_call (dedicated_number / shared_imessage_number), or omits it: the bridge then uses the only available line, or — when both exist — the line matching the current conversation's channel.

External events

Besides Inkbox's own events, the webhook endpoint can inject events from outside systems (e.g. a CI failure) to wake the agent on its own external:<source> thread. Routing is by verified source, never by the body's claimed event type:

  • Registered providers (e.g. GitHub via X-Hub-Signature-256) are verified with their own secret from INKBOX_WEBHOOK_SECRET_<NAME>; registering the provider + setting its secret is the opt-in, and forged signatures are rejected outright.
  • Mock provider accepts a shared secret directly in X-Inkbox-Mock-Secret, matched against INKBOX_WEBHOOK_SECRET_MOCK. It is intended for manual curl probes and test systems that cannot calculate a body signature. A valid secret makes the event verified and wakes the agent even when generic external events are disabled.
  • Everything else (unknown sources, or Inkbox-signed payloads with no handler) is delivered only when INKBOX_EXTERNAL_EVENTS_ENABLED=true, and unverified events carry a cautious directive that forbids irreversible action on their say-so.

No human reads an external thread, so the agent is told to act via its tools rather than reply. Adding a source is drop-in: a new module in inkbox_codex/webhook_providers/ with a @register_provider class.

Send a mock event with:

curl --fail-with-body --request POST 'https://your-agent-host.example/webhook' \
  --header 'Content-Type: application/json' \
  --header "X-Inkbox-Mock-Secret: $INKBOX_WEBHOOK_SECRET_MOCK" \
  --data '{
    "id": "mock-run-123",
    "source": "mock-ci",
    "event": "workflow.failed",
    "title": "Mock workflow failed",
    "summary": "Inspect the repository and decide what action is appropriate.",
    "requested_action": "Investigate the failure and take any safe corrective action."
  }'

Email replies

Automatic email replies use reply-all on the original message. The reply goes to its Reply-To address (or sender), with the other visible To/CC recipients included in CC. The agent's own mailbox and BCC recipients are excluded, and the original email thread is preserved. No mention is required for email replies. Queued replies, approval prompts, and send-failure recovery retain the original message's reply target even if another message arrives in the same session.

If an inbound event lacks the original message ID, the bridge cannot send an automatic reply-all; it does not fall back to a sender-only email.

Live reply-all CI uses the existing CODEX_INKBOX_API_KEY and REMOTE_INKBOX_API_KEY to verify real reply delivery, sender deduplication, self-exclusion, and reply-thread headers. An optional REPLY_ALL_INKBOX_API_KEY for a third, non-auto-replying dedicated inbox enables separate CC delivery checks (including additional original To recipients). Those two additional cases are explicitly skipped when the third credential is absent; two-identity checks do not prove independent CC delivery. The live channel workflow's email_reply_all_only input runs just these checks, with both deterministic and real-model gateway runs, without SMS reset traffic.

Media

Inbound. When someone sends an MMS image, an iMessage attachment, or an email with files, the gateway downloads them to ~/.inkbox-codex/media/ (override with INKBOX_CODEX_MEDIA_DIR) and appends the local paths to the message, so Codex can open them with its Read tool — including viewing images. Media-only messages (no text) still wake the agent.

Outbound. Codex sends media with a single tool call per channel — it just passes local file paths, and the tool handles any upload-then-send round trip internally:

  • Email — inkbox_send_email(..., attachment_paths=[...]) (base64 inline, ~25 MB total).
  • iMessage — inkbox_send_imessage(..., media_path=...) (uploaded + sent, ≤10 MB).
  • SMS/MMS — inkbox_send_sms(..., media_paths=[...]) (uploaded + sent; media_urls also accepts already-hosted URLs).

Vault and 2FA codes

Codex can access credentials and generate current 2FA codes from Inkbox Vault without opening a browser for each request.

  1. Store the credential in your Inkbox Vault and grant the agent identity access. For 2FA, use a login secret with TOTP configured.
  2. Set INKBOX_CODEX_VAULT_KEY to your vault unlock key in the bridge's local environment or its .env file (normally ~/.inkbox-codex/.env). This is separate from INKBOX_API_KEY. Keep the key out of chat and source control.
  3. Restart the bridge with inkbox-codex restart, or restart its service if you run it through a service manager.
  4. Ask: "Get the current 2FA code for my Example login from Inkbox Vault."

The agent uses inkbox_list_vault_secrets to find the login, then calls inkbox_get_totp_code with its secret_id. The result includes code, period_start, period_end (Unix timestamps), and seconds_remaining. Request a fresh code if it expires. This tool returns neither the password nor the TOTP seed.

inkbox_get_vault_secret retrieves one credential when the task needs it. Login results include has_totp instead of the TOTP seed. Listing returns metadata only and works without an unlock key. The tools follow the API key's access scope, so use the agent-scoped key configured by setup.

The bridge unlocks the vault only when a credential or 2FA tool needs it. An incorrect bridge vault key or failed unlock returns a tool error; messaging and metadata listing remain available. Correct the key locally and restart, or retry after a temporary service failure.

Use the bridge-specific INKBOX_CODEX_VAULT_KEY rather than the SDK-wide INKBOX_VAULT_KEY or vault_key in ~/.inkbox/config. Those SDK-wide settings trigger eager unlocking for every SDK client, including the messaging gateway. If you followed earlier instructions using INKBOX_VAULT_KEY, rename that setting and remove any SDK-wide vault key used only by this bridge before restarting.

If the requested secret is missing, check its access grant to this agent. A login without TOTP must have it configured before the agent can generate codes.

Config reference

Env var Required Default Description
INKBOX_API_KEY yes - Agent-scoped Inkbox API key.
INKBOX_IDENTITY yes - Inkbox agent identity handle.
INKBOX_SIGNING_KEY inbound - Webhook HMAC secret for signed inbound events.
INKBOX_CODEX_VAULT_KEY vault reads / 2FA - Vault unlock key, supplied locally and used only when a credential or 2FA tool needs it. See Vault and 2FA codes.
CODEX_PROJECT_DIR yes cwd Directory Codex works in.
CODEX_MODEL no CLI default Model override for bridged sessions.
INKBOX_REQUIRE_SIGNATURE no true Refuse unsigned inbound webhooks unless false.
INKBOX_SKIP_WEBHOOK_RECONCILE no false Leave webhook subscriptions untouched on start. For deployments that provision them ahead of time, where the destination is fixed or this API key may not change it. They must already point at this bridge's webhook URL, or nothing arrives.
INKBOX_CONTACT_MEMORIES_ENABLED no true Add memories supplied with the matched webhook contact as background context.
INKBOX_BASE_URL no SDK default Override the Inkbox API base URL.
INKBOX_PUBLIC_URL no - Public bridge URL. Omit to use an Inkbox tunnel.
INKBOX_TUNNEL_NAME no identity handle Tunnel name override.
INKBOX_ALLOWED_USERS no - Local allowlist (emails / E.164 numbers). Usually leave empty and use Inkbox contact rules.
INKBOX_ALLOW_ALL_USERS no false Allow all senders admitted by Inkbox contact rules.
INKBOX_BRIDGE_PORT no 8767 Local webhook server port.
INKBOX_PERMISSION_TIMEOUT_S no 600 Seconds to wait for a permission/poll reply.
INKBOX_GROUP_REPLY_MODE no auto Group SMS/iMessage, Slack, and Companion email replies: auto lets the agent decide; mention requires a native Slack @mention, SMS/iMessage @agent or @<agent-handle>, or Companion email addressed To the agent. Other messages become context without starting a turn. Also configurable in setup.
INKBOX_IMESSAGE_THREADED_REPLIES no false Preview native iMessage replies, ordinary burst collection, and queued follow-ups. Requires a compatible SDK and API; see Native iMessage replies. Environment-only; no wizard change.
INKBOX_COMPANION_RESPONSE_MODE no safe Companion SMS/MMS, iMessage, and email: safe wakes only for direct access; relaxed permits any delivered sender. Both honor Auto/Mention. Sponsored and unknown access stays context-only in Safe mode. Also configurable in setup.
INKBOX_CODEX_AUTO_APPROVE_INKBOX_TOOLS no false Auto-accept Codex MCP prompts for Inkbox tools only. The setup wizard writes true when you trust the agent to send through Inkbox without per-call approval.
INKBOX_A2A_PROGRESS_INTERVAL_SECONDS no 180 Seconds between progress updates for active inbound A2A tasks. Set to 0 to disable periodic updates.
INKBOX_VOICE_STACK no inkbox_tts_stt inkbox_voice_ai, openai_realtime, or inkbox_tts_stt. When absent, legacy Realtime settings remain compatible.
INKBOX_VOICE_AI_AUTHORITY_MODE Voice AI contact_scoped Saved Voice AI authority selected during setup: contact_scoped or yolo.
INKBOX_VOICEMAIL_DETECTION no enabled Outbound-call voicemail policy: enabled or disabled. Live CI uses disabled.
CODEX_BIN no codex Codex CLI executable to run.
CODEX_SANDBOX no workspace-write App-server thread sandbox (read-only, workspace-write, danger-full-access).
CODEX_APPROVAL_POLICY no on-request Codex approval policy for bridged turns.
INKBOX_REALTIME_ENABLED no false Use OpenAI Realtime for calls. Needs a key; off → Inkbox STT/TTS.
INKBOX_REALTIME_API_KEY realtime OPENAI_API_KEY OpenAI key with /v1/realtime access.
INKBOX_REALTIME_MODEL no gpt-realtime-2 Realtime model id.
INKBOX_REALTIME_VOICE no cedar Realtime voice name.
INKBOX_REALTIME_FALLBACK_TO_INKBOX_STT_TTS no true Fall back to Inkbox STT/TTS if OpenAI connect fails.
INKBOX_EXTERNAL_EVENTS_ENABLED no false Wake the agent on unrecognised (external) webhooks — see External events.
INKBOX_WEBHOOK_SECRET_<NAME> per provider - Verification secret for a registered third-party webhook provider (e.g. INKBOX_WEBHOOK_SECRET_GITHUB).
INKBOX_WEBHOOK_SECRET_MOCK mock provider - Expected plaintext value of the X-Inkbox-Mock-Secret header. Use a random secret.

Tools exposed to Codex

The agent reaches you (or third parties) through an in-process MCP server:

  • inkbox_whoami — its own identity: handle, mailbox, iMessage status, and its two calling lines (dedicated number vs shared iMessage line).
  • inkbox_place_call — place an outbound voice call over either line (origination: dedicated_number / shared_imessage_number) — see Two calling lines.
  • inkbox_list_calls · inkbox_get_call_transcript — browse call history and transcripts.
  • inkbox_send_email — send email; attach local files with attachment_paths.
  • inkbox_send_sms — send SMS/MMS; attach local files with media_paths (or hosted media_urls).
  • inkbox_send_imessage — send into an iMessage conversation; attach a local file with media_path.
  • inkbox_get_imessage_thread · inkbox_get_imessage_conversation_thread — bounded native-thread pages, available only with INKBOX_IMESSAGE_THREADED_REPLIES=true; reply targets and fallback policy are bridge-owned, while the send tool also accepts idempotency_key.
  • inkbox_list_text_conversations · inkbox_get_text_conversation — browse SMS threads and history.
  • inkbox_list_imessage_conversations · inkbox_get_imessage_conversation — browse iMessage threads and history (find the conversation_id to send into).
  • inkbox_lookup_contact · inkbox_list_contacts · inkbox_get_contact — resolve and read address-book contacts (reverse-lookup by email/phone, free-text search, or full record by id).
  • inkbox_create_contact · inkbox_update_contact · inkbox_delete_contact — save, edit, and remove organization-wide contacts. Changes affect the shared address book. vCard export/import is not exposed.
  • inkbox_list_vault_secrets · inkbox_get_vault_secret · inkbox_get_totp_code: find vault credentials, retrieve one credential, or generate a current 2FA code. See Vault and 2FA codes.
  • inkbox_a2a_call · inkbox_a2a_check · inkbox_a2a_reply — delegate work to another agent and follow its task.
  • inkbox_list_a2a_tasks · inkbox_list_a2a_messages — page and search this identity's inbound and outbound A2A history, with participant, task, context, role, state, and timestamp filters.
  • inkbox_a2a_complete · inkbox_a2a_ask_caller · inkbox_a2a_fail — commit the outcome of a verified inbound A2A task. These tools are rejected outside that task's isolated session.

Inbound A2A tasks acknowledge pickup immediately. While a task remains active, the worker sends a short progress update about every three minutes by default; these updates are visible in task history without starting a requester turn.

The bridge requires Inkbox SDK 0.7.11 or newer (below 1.0.0).

On a live call, the OpenAI Realtime voice agent additionally gets consult_agent, register_post_call_action / edit_post_call_action / delete_post_call_action, and hang_up_call — see Voice.

Smoke test

  1. inkbox-codex doctor — everything green.
  2. Text START, then text the agent; verify it replies in the same thread.
  3. Ask it to do something requiring a command (e.g. "run the tests") and verify you get a permission text; reply 1 and verify the result comes back.
  4. Ask it something open-ended enough to trigger a poll; reply with a number.
  5. Email the agent; verify the reply lands as an email on the same thread.
  6. Call the number, ask what it's working on, hang up mid-answer, and verify the late voice tail is not silently sent as SMS or email.

Development

python -m pytest

Architecture notes

  • Tunnel-first inbound: with a signing key, the gateway opens an Inkbox tunnel, reconciles mail/text/iMessage plus call.ended subscriptions, and sets the identity's incoming-call action from the selected stack — hosted_agent for Voice AI or auto_accept plus the call WebSocket for local stacks.
  • Session routing: direct messages use the resolved contact id across channels, falling back to the channel conversation or raw address/number. Group SMS uses the group conversation id regardless of sender.
  • Escalation over the active channel: a pending permission/poll captures the contact's next inbound message as its answer, on whichever text channel they're using.
  • Codex app-server: each contact session owns one codex app-server subprocess, one Codex thread, app-server approval request handling over Inkbox, and a local stdio MCP server for the Inkbox tools.

HD realtime voice

Realtime calls request 16 kHz mono PCM16 call audio. The bridge continuously resamples to and from the realtime session's 24 kHz PCM format, preserving audio across WebSocket frame boundaries. Older call streams that advertise 8 kHz μ-law (or omit their audio descriptor) remain supported. Call audio quality also depends on the remote connection. Hosted voice and managed speech modes are unchanged.

About

Give your Codex agent its own Inkbox identity: a mailbox, iMessage, a phone number for calls and SMS, and an internet address. Talk to it from anywhere.

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages