Give your Codex agent its own Inkbox identity:
a mailbox, iMessage, a phone number for calls and SMS, and an internet address.
Step away from the keyboard and keep working with it from anywhere.
Email · Calls · SMS / MMS · iMessage · Tunnel
- Codex installed and logged in. The bridge drives a real Codex session, so the
codexCLI has to be on the machine and authenticated — install it (developers.openai.com/codex), then either sign in with a ChatGPT/Codex login or setOPENAI_API_KEY.inkbox-codex doctorchecks for it. - Python 3.11+. The installer finds one and builds the bridge its own venv.
- macOS or Linux. Boot persistence uses a systemd user unit on Linux and a launchd agent on macOS.
- An Inkbox agent — nothing to set up in advance; the setup wizard self-signs up for you (or takes an existing API key).
This finds a Python 3.11+, installs the bridge in its own venv, puts inkbox-codex on your PATH, and runs the setup wizard:
curl -fsSL https://raw.githubusercontent.com/inkbox-ai/codex-plugin/main/install.sh | bashThat's the whole setup. The wizard creates a fresh Inkbox agent for you (or takes an existing API key), provisions a phone number, connects iMessage, mints a webhook signing key, picks the project directory Codex works in, and offers to keep the bridge running on every boot. When it finishes, text/email/call your agent and it answers from a real Codex session.
The one thing to have ready: be logged into Codex — a ChatGPT/Codex login (via the Codex app/CLI) or OPENAI_API_KEY set. The installer checks this and warns if it's missing.
Flags: --start (launch the background gateway when done), --no-setup (install only). From a local checkout, run ./install.sh. Re-running is safe.
For unattended agent setup, install without opening the wizard and pass the API key through the environment (or standard input), never a command-line argument:
curl -fsSL https://raw.githubusercontent.com/inkbox-ai/codex-plugin/main/install.sh | bash -s -- --no-setup
export INKBOX_API_KEY="ApiKey_..."
inkbox-codex bootstrap --identity my-agent --project-dir "$PWD" \
--voice-ai --rotate-signing-key --start-gateway
unset INKBOX_API_KEYbootstrap validates that the key can access exactly the requested identity, scopes down an admin key before saving it, preserves existing Voice AI settings, enables native Inkbox tool approvals, and starts or restarts the detached gateway. Signing-key replacement is opt-in because it transfers verified webhook delivery away from any gateway using the previous key. The command prints a secret-redacted JSON result and is safe to resume.
Check it any time:
inkbox-codex doctor # config, codex CLI/auth, identity reachability
inkbox-codex status # is the background gateway up? where are the logs?you (phone) ── SMS / iMessage / email / call ──▶ Inkbox ──▶ tunnel ──▶ bridge
│
▼
Codex session
(full tool access in
your project dir)
-
Text, iMessage, email, or call your agent's Inkbox number. Each remote party gets one Codex session spanning every channel — text it on the walk home, then email it details, same conversation.
-
Codex runs with full tool access in
CODEX_PROJECT_DIR. It reads, searches, and browses freely; anything risky (running commands, editing files) is escalated to you as a text:Codex wants to run the command: npm test
Reply 1 (or YES) to allow once, 2 (or SESSION) to allow for this session, 3 (or NO) to deny. /stop cancels the task.
-
When Codex needs you to pick between options (the
AskUserQuestiontool), you get a numbered poll on whatever channel you're on, and your reply is fed back as the answer. -
Each message you send is tagged with its channel, so Codex knows whether it's on SMS, iMessage, email, or a call.
-
A channel prompt is appended to Codex's system prompt so replies fit a phone: plain text, no markdown, short, jargon kept to a minimum ("saved and published the change", not "pushed to origin/main").
-
Codex also gets Inkbox tools (
inkbox_send_email,inkbox_send_sms,inkbox_send_imessage, …) so it can proactively reach you — "email me the full report" works.
If you'd rather not run the installer (any Python 3.11+ environment):
pip install -e .
inkbox-codex setup # interactive wizard — writes .env for you
set -a; source .env; set +a
inkbox-codex doctor
inkbox-codex runinkbox-codex setup walks you through everything and writes .env: create a fresh Inkbox agent via self-signup (or bring an existing API key), pick or create the identity, attach the Codex avatar, provision a phone number, wait for your START opt-in, choose a phone-call voice stack, connect iMessage, mint a webhook signing key, choose the project directory, choose whether to trust Inkbox MCP tools without repeated allow prompts, and set up autostart. Realtime keys are validated before selection is saved. An admin key used to change Voice AI authority is held only for that setup run and is never written to .env. Rerun setup anytime to reconfigure.
The repository includes a manual-test image with Codex, the plugin, and the Inkbox SDK preinstalled. It reuses the host's native Codex login and keeps Inkbox plugin state in a named volume; neither is copied into the image:
docker build -t inkbox-codex-local .
docker run -dit --name inkbox-codex-local \
-v "${CODEX_HOME:-$HOME/.codex}:/root/.codex" \
-v "$PWD:/workspace" \
-v inkbox-codex-state:/root/.inkbox-codex \
inkbox-codex-local
docker exec -it inkbox-codex-local bash
# inside the container
inkbox-codex setup
inkbox-codex doctor
inkbox-codex run
# after leaving the container
docker rm -f inkbox-codex-localThe image contains no credentials. If you use API-key authentication instead
of codex login, add -e OPENAI_API_KEY="$OPENAI_API_KEY" to docker run.
Inkbox credentials are entered into setup or supplied only at runtime.
On startup the bridge opens an Inkbox tunnel, wires mail/text/iMessage webhook subscriptions and the incoming-call channel to it, and routes everything into Codex sessions.
inkbox-codex run # foreground (Ctrl+C to stop) — good for first runs and debuggingOr run it as a background daemon (PID + log under ~/.inkbox-codex/):
inkbox-codex start # detach and run in the background
inkbox-codex status # is it running? where are the logs?
inkbox-codex restart # restart it
inkbox-codex stop # graceful stop (SIGTERM, then SIGKILL after 5s)
tail -f ~/.inkbox-codex/gateway.logstart auto-loads .env from the current directory, so you don't have to source it first. run is the foreground version a service manager (systemd, Docker) should supervise; start/stop are the self-contained background option.
The setup wizard offers to keep the bridge running for you — either just in the background for this session, or as a service that starts on every boot. On Linux it installs a systemd user unit (~/.config/systemd/user/inkbox-codex.service) and enables it; on macOS it installs a launchd agent. To keep a Linux service alive while you're logged out, enable lingering once:
sudo loginctl enable-linger "$USER"
systemctl --user status inkbox-codex # restart | stop | statusinkbox-codex uninstall # stop it, remove the boot service + launcher; keep config
inkbox-codex uninstall --purge # also delete ~/.inkbox-codex (config, logs, sessions)This is local-only — webhook subscriptions on the Inkbox side are left as-is; remove them in the Inkbox Console if you want.
Then, from your phone:
- Text
STARTto the agent's number (first time only, carrier opt-in). - Text it something like "clean up the TODOs in the auth module".
- Approve the permission texts as they arrive. Get the result as a text.
The bridge honors Codex's configured approval policy and relays its requests over your active Inkbox channel:
- Commands, file changes, permission-profile changes, and request-user-input prompts block the agent mid-turn while the bridge texts you a one-line plain-language summary.
- MCP tool approvals number the available choices consecutively; reply with the number shown or a word from the table below. A different instruction cancels the pending approval turn and becomes fresh work instead of being consumed as an approval answer.
- Request-user-input prompts are formatted as numbered options; reply with the number or free text.
- No reply within
INKBOX_PERMISSION_TIMEOUT_S(default 10 min) cancels an MCP request; command/file approvals are denied.
| Reply | Decision |
|---|---|
YES |
Allow once |
SESSION |
Allow for this session, when offered by Codex |
NO |
Deny this request |
ALWAYS |
Remember approval across sessions, when offered by Codex |
/stop / /cancel |
Cancel the entire task |
For example, a tool offering only allow-once and deny displays 1 — Allow once
and 2 — Deny. If session approval is also available, it displays 1 — Allow
once, 2 — Allow for this session, and 3 — Deny. Always follow the numbers
in the current prompt; words such as NO keep the same meaning.
Session and permanent approval are distinct: an unavailable scope is never silently changed to allow-once. Advertised MCP persistence scopes are forwarded through Codex's native protocol, not a bridge-wide allow list. Other tools still require their own approval when Inkbox tool auto-approval is enabled.
Clear replies such as “yes, proceed” are accepted. A natural cancellation such as
“please cancel my request” stops the task. Companion answers and controls retain
sender and mention gates. Structured MCP questions show their actual fields and
return typed answers; unsupported forms require using Codex directly or canceling
the interaction. URL-based requests require completing the displayed URL step
before replying DONE.
Set INKBOX_SLACK_ENABLED=true to add Slack to the gateway. Inkbox SDK 0.7.11
includes the required Slack methods; the API must also support Slack. Connect a
workspace to the same Inkbox identity to exchange messages; the receiver can
start before installation.
Run inkbox-codex setup to opt in interactively. The wizard offers Connect Slack
now?, selects a saved provisioning workspace, waits for app preparation, prints
an installation link to open in your browser, and polls until that workspace is
connected. Keep the link private and complete authorization in the same browser.
This flow requires an SDK exposing client.slack.list_provisioning_workspaces
and identity-wide webhook subscriptions. Setup uses the bridge's existing claimed
agent key for workspace credentials, app preparation, and installation; no extra
admin key is needed. A newly self-signed-up identity must be claimed first. An
organization admin key is also supported when already configured for the bridge.
If you need to add a provisioning workspace or renew its credentials, the wizard prompts for a masked Slack app-configuration access/refresh token pair from Your App Configuration Tokens. Inkbox verifies and stores the pair for your organization; the bridge does not save it locally. If the pair is rejected, setup explains the credential issue and offers another attempt without leaving the wizard. Other errors identify the failed step and HTTP status when available; they do not silently retry app creation or expose tokens. These are not bot tokens. Each identity's app is permanently bound to one workspace. Existing connections need no new installation; manage permission refreshes in the Inkbox console. After installation, add the bot to the channels where you want to use it.
Existing connections are shown on reruns. You can skip connecting, press Ctrl+C during either wait, or rerun after the five-minute wait expires. Rerunning setup and declining full reconfiguration still offers Slack onboarding. Declining Slack turns it off only in this bridge; it does not disconnect the workspace or disable Slack for other clients. On startup, the gateway registers the subscriptions below.
Work indicators. Accepted threaded Slack requests use the native agent
loading indicator, without adding reaction bubbles. It stays while work is queued
or running, switches to awaiting-input during questions or approvals, and returns
to ready after completion, failure, or cancellation. Overlapping requests in the
same thread share one indicator. Mention-mode context does not start it.
Slack's native Stop button cancels work in the matching engaged thread, subject to
the same allowed-sender policy. Ordinary main-DM messages receive replies in the
main DM and share conversation context, without a native status indicator.
Instead, the exact DM message receives 👀 while queued, running, or waiting for
input. It is removed after success or cancellation; a failed turn or reply delivery
replaces it with ❌. These reactions do not start a thread.
An explicit native @mention starts a thread beneath that message; messages already
inside a thread stay there. Threaded replies and their indicators use the same
thread. Continue there for follow-ups, approval answers, and commands; each thread
keeps its own saved context.
Native status needs an app declared as a Slack Agent (including assistant:write),
chat:write, SDK processing-status support, and workspace feature availability.
Existing installations may need an app-configuration update and reauthorization.
Unsupported or uncertain updates are logged without
blocking replies or falling back to reactions. Restart cleanup clears unfinished
native status, marks interrupted DM work with ❌, and removes pending indicators
from the previous reaction version. Reaction failures do not block replies.
No separate Slack bot token is needed at runtime. Slack subscriptions use
slack.dm_received, slack.group_dm_received, slack.mention_received,
slack.thread_reply_received, and slack.session_stopped with no Slack-specific
filter. The gateway also
recognizes slack.channel_message_received, while its local attention rules keep
unrelated channel chatter from waking Codex. Update the API and receiver together
when upgrading from the older incoming-message preview event.
Incoming messages can match several categories; the payload's message_kinds
retains the full classification even when only one event type is selected.
Replies remain deduplicated by the stable event ID, not the selected category.
- By default, DMs, group DMs, and mentions wake Codex. New requests start a thread; follow-ups in an already-engaged thread continue its conversation without another mention. Unrelated channel messages and bot messages do not wake it.
- With
INKBOX_GROUP_REPLY_MODE=mention, Slack channels, group DMs, and their threads require a native @mention of this agent on each new message to generate a reply, even after the agent has joined the thread. Unmentioned follow-ups are context only: no reply, tool execution, or interruption of an active turn. Direct DMs are unchanged. Bridge commands and the requesting sender's answers to pending approval/question prompts still work without a mention. - Each workspace/conversation/thread gets a separate, resumable Codex session. When supplied by Inkbox, the sender's linked contact ID and Slack profile fields are included as context. Email and phone may be absent. The bridge does not match contacts itself, merge conversations, or grant permissions from profiles.
- Replies and approval prompts stay on the originating Slack route. Only the requesting sender can answer a pending group-thread approval.
- Messages are limited to 12,000 characters per send. The bridge rejects longer replies before sending; it does not split or automatically retry them.
- Six tools list workspaces, list conversations, read messages/threads, search retained text, send explicitly requested messages, and inspect send outcomes. Ordinary final replies are sent automatically; do not send them again by tool.
- Unconfirmed sends are logged with their action ID when available, never automatically resent. Reuse the exact idempotency key when retrying an explicit send. Webhook duplicate suppression follows the gateway's existing bounded, in-memory behavior.
- Attachments are exposed as metadata in this first version. File transfer and arbitrary reaction tools are not included.
INKBOX_ALLOWED_USERS accepts Slack user IDs or workspace-qualified T_ID:U_ID
entries. An empty list admits all human senders whose events reach the identity.
Slack subscriptions are additive: starting this gateway does not remove another
Slack receiver. Existing active subscriptions at this receiver's URL, including
mixed-event subscriptions, are reused; only missing events are registered.
Paused overlapping subscriptions require attention in the console before startup.
Stop/remove any previous Slack receiver if it should no longer reply.
The harness uses the same gateway, sessions, signing verification, tools, and
reply path. It registers only Slack and does not change email, phone,
iMessage, or A2A subscriptions. Credentials are read into the process from the
selected file, never copied into its state directory. Use a separate directory
and port from your regular gateway. Run only one gateway on an identity’s tunnel
at a time, or pass --public-url for a separately reachable receiver.
After installing this checkout and a Slack-capable SDK, run a read-only preflight:
python -m inkbox_codex.slack_harness \
--credentials-file ~/.env --api-key-env INKBOX_API_KEY \
--identity "$INKBOX_IDENTITY" --base-url "$INKBOX_BASE_URL" \
--state-dir ~/.inkbox-codex-slack --project-dir "$PWD" \
--signing-env-file /path/to/identity.envAdd --run to connect the identity’s existing tunnel and register the signed Slack receiver. The signing
file must contain the identity's existing INKBOX_SIGNING_KEY; the harness never
creates or rotates keys. --allow-user T_ID:U_ID can be repeated. Ctrl+C stops
the local receiver; its subscription remains available for the next run.
--group-reply-mode mention enables mention-only group/thread replies in the
harness; it otherwise follows INKBOX_GROUP_REPLY_MODE, defaulting to auto.
The webhook still includes thread messages so quiet context and approval answers
can reach the session.
In mention mode, test a native @mention, an unmentioned follow-up in the same thread, and another @mention: expect reply, silence, reply, with the quiet message available as context for the last reply.
For manual acceptance: DM the agent, mention it in a channel, continue in that
thread without a mention, ask for recent thread history or retained-text search,
then exercise an approval and /status. Replies should stay in the right thread;
unrelated channels and bot messages should remain quiet. The preflight does not
send messages. Live delivery can only be verified after connecting the workspace
and starting the harness.
Offline regression harness: python -m pytest -q tests/test_slack.py tests/test_slack_harness.py.
Direct-message sessions are keyed by Inkbox contact, so one person = one conversation across channels. Group SMS messages share a session keyed by the group conversation, separate from direct messages and other groups, while retaining each sender's contact details. Codex session ids are persisted in ~/.inkbox-codex/sessions.json and resumed across bridge restarts — your conversation picks up where it left off. Replies go out on the channel you last used. If a voice call ends before Codex finishes a voice reply, that late voice reply is dropped instead of silently switching to SMS or email.
Group replies. The setup wizard offers Automatic (default) or Mention required for group SMS, iMessage, and Slack, saved as INKBOX_GROUP_REPLY_MODE=auto|mention. Automatic keeps the existing behavior: the agent decides whether to answer. For SMS/iMessage, mention mode starts a reply only when the new message itself includes @agent or @<agent-handle> as a whole mention, case-insensitively; older messages, links, and email addresses do not count. Slack requires a native @mention of the agent, including on follow-ups in an engaged thread. Other messages and group reactions are added to Codex's context without generating a reply or a typing indicator. While a turn is running, background messages wait until it finishes before being appended. Appended context persists with the Codex thread; messages still waiting in the bridge's queue are not persisted across a restart. Direct messages are unchanged. Commands such as /stop, and answers to the agent's pending questions from the sender it asked, do not require a mention.
Verified SMS/MMS, iMessage and email webhooks may include a top-level
companion object. Webhooks without it (or with null) retain normal routing.
The bridge loads the full authorized initialization snapshot using the Inkbox
SDK, including every page, and submits one combined input. The sponsor trigger
is included once; historical messages, attachment descriptors, and history notices remain
conversation data, never individual turns or slash commands.
The setup wizard offers two independent settings, including when rerun for an existing identity:
- Companion responses:
INKBOX_COMPANION_RESPONSE_MODE=safe|relaxed. Safe (default) allows only messages withsender_access="direct"to wake the agent. Sponsored messages, or messages with missing/unrecognized access, are appended to context without model generation, typing, or a reply. Relaxed allows any delivered message to wake the agent, including sponsored and unknown-access messages. - Group replies:
INKBOX_GROUP_REPLY_MODE=auto|mention, also applied to Companion email. Automatic lets an eligible message start a turn; the agent decides whether a reply is warranted. Mention required additionally requires@agentor@<agent-handle>in the current message's own text. For Companion email, putting the agent's mailbox in the current message's To recipients also counts as a mention. Address matching is case-insensitive and supports display names; Cc/Bcc alone do not count. Quoted headers, historical mentions, notices, and attachment metadata do not count. Reply-all may retain the agent in To and therefore satisfy this gate again.
The signed webhook supplies sender_access on data.text_message for SMS/MMS
or data.message for iMessage/email. direct means contact rules permitted the
message without sponsorship, including allowed-by-default senders; sponsored
means delivery was authorized through sponsorship. Neither value grants command
permissions or establishes permanent trust. Access is not inferred from the
sender's contact record, sponsor identity, or Companion phase.
Wake below means context plus a model turn; Context means context only. These rules apply equally to SMS/MMS, iMessage, and email Companion inputs.
| Companion mode | Group replies | Direct, no mention | Direct, mention | Sponsored, no mention | Sponsored, mention |
|---|---|---|---|---|---|
| Safe (default) | Auto | Wake | Wake | Context | Context |
| Safe (default) | Mention | Context | Wake | Context | Context |
| Relaxed | Auto | Wake | Wake | Wake | Wake |
| Relaxed | Mention | Context | Wake | Context | Wake |
For email, the table's mention includes the agent being in To. This does not override sender access: sponsored and unknown-access messages remain context-only in Safe mode even when addressed To the agent. Auto is unchanged.
Context-only messages are retained for later eligible turns. For initialization, the entire snapshot is appended together; only the current received message can trigger generation. Webhooks without Companion metadata retain normal routing and are unaffected by the Companion response setting.
- Sessions are isolated by API environment, identity, channel, server scope, and activation, separate from ordinary contact sessions. A new activation starts a fresh session with its authorized snapshot. Each new released snapshot must have a distinct activation ID; replaying an immutable activation is not a new batch.
- Initialization completes before queued live messages run. A first-seen live
event loads history as context, without a separate historical sponsor turn;
only the live message can wake the agent. If that message is already in the
snapshot, the combined input uses its access and mention instead of repeating it.
Pending events run in sequence order; numeric
gaps do not stall the conversation. Late unseen events older than an already
submitted input require reconciliation instead of running out of order.
phase="ordinary"uses a separate scoped session and never loads hidden history. - Replies use the original group conversation ID, or the SDK-approved email reply context's stored message UUID with canonical reply-all. They never fall back to privately messaging the latest author. Local sponsor restrictions remain in force. Live turns and replies use the signed conversation scope and saved sponsor message without additional activation lookups. Replies reuse the identity loaded at startup.
- Only a live reply by the prompted, locally allowed sender that passes both
response gates can answer an approval request. Historical or context-only
approval-looking text cannot. Sponsor slash controls also require both gates.
In Mention mode, address email To the agent or prefix answers and controls with
@agent, for example@agent allowor@agent /stop; escalation prompts remind you of this.
SDK requirement: The bridge requires Inkbox SDK 0.7.11 or newer (below 1.0.0), including
client.companion.load_initialization and activation_messages. Installation
resolves the published SDK; CI tests the released inkbox==0.7.11 minimum across
unit, real-host, and live-channel lanes. No preview source checkout is needed.
An unsupported SDK produces an explicit webhook error; the bridge never falls
back to submitting only the trigger.
Wake decisions use the signed current-message access and do not require extra
SDK lookups. Older SDKs may omit per-entry access labels from rendered history;
those historical entries remain context, never independent triggers. Safe mode
on webhooks that omit sender_access is context-only; no missing field is treated
as direct access.
Delivery and recovery: receipts and initialization checkpoints live under
$INKBOX_CODEX_HOME/companion/ (default ~/.inkbox-codex/companion/), with private
permissions and one receiver owner per environment/identity. Transient reads before
submission retry with capped backoff while the gateway runs; other pre-submission
failures retry up to five times. Pending inputs resume after restart or webhook
redelivery. Failed Codex startup or resume attempts release the incomplete client
and retry with the same saved thread; they do not count as submitted model turns.
Closing or clearing a session also disposes its connecting client, so a late
startup response cannot restore a cleared conversation.
A retry may refresh its delivery timestamp or inline history preview
without creating another input. Stable event IDs and acknowledged snapshot/live source-message IDs
are deduplicated across restarts within their scope and activation. Completed answers
are saved before delivery.
Transient read failures while preparing delivery retry with capped backoff, including
after restart, without rerunning the model. An interrupted host
submission or uncertain reply send pauses that scope rather than risking another
model turn or duplicate send. The log identifies the retained receipt. Inspect the
scoped Codex thread and channel delivery before operator recovery; do not delete
receipts or blindly replay uncertain events. This is at-most-once automatic
retry behavior at ambiguous boundaries, not an exactly-once transport claim.
Initialization above 8 MiB fails explicitly rather than truncating. The preview
uses POSIX file locking (Linux/macOS).
Run inkbox-codex setup to change the choice without reconfiguring your identity, or edit .env, then restart the bridge to apply it. For background mode, use inkbox-codex restart; for a systemd installation, use systemctl --user restart inkbox-codex.service. Mention mode requires a Codex version supporting thread/inject_items.
Typing indicator. While Codex works on a turn, the bridge keeps a typing indicator alive on your iMessage thread (refreshed every few seconds, since it expires) so you can see it's busy. SMS, email, and voice have no typing indicator. Slack threads use the native agent status described above.
Delivery failures. Outbound messages can silently fail — a carrier filters an SMS, an iMessage is declined, an email bounces. Inkbox reports these asynchronously (text.delivery_failed, imessage.delivery_failed, message.bounced/message.failed). The bridge catches them and wakes the affected contact's session to tell Codex which message didn't land and why, so it can retry or reach you another way (a different channel, or a call) using its Inkbox tools. The notice runs as a side-effect turn — Codex acts via tools rather than replying on the channel that just failed — and repeat webhooks for the same message are de-duplicated so it can't loop. text.delivery_unconfirmed is different: it only means the carrier couldn't confirm delivery (the message usually landed), so it's logged for debugging without waking Codex — waking there would resend a message that was likely delivered.
Interrupt by texting again. By default, messaging the agent again while it's mid-turn works like pressing Esc in Codex and typing a new message: the running turn is interrupted, its partial answer is dropped, and Codex picks up your new message instead. (A reply while it's waiting on a permission/poll still answers that escalation — interrupting only applies while it's actively working.) The opt-in iMessage behavior below queues follow-ups instead.
Set this in the running gateway's environment and restart its existing service:
INKBOX_IMESSAGE_THREADED_REPLIES=trueThe default is false; there is no new setup-wizard step. Enabling requires the
SDK's native reply/thread-read support (install Python SDK 0.7.13 or newer)
and an API environment serving those endpoints. doctor distinguishes local SDK support
from backend capability; a targeted reply also checks the native endpoint before
sending. The bridge fails explicitly if support is missing, rather than silently
sending an unthreaded reply. The existing SDK minimum is unchanged when disabled.
- Several quick messages: ordinary iMessage inputs from the same sender, conversation, and reply context collect for 750 ms of quiet, up to 2 seconds. Codex receives their text and source IDs in order. Different senders or explicit native threads stay separate. Companion inputs retain their existing ordered, one-event-per-turn receipts and activation boundaries; they are not coalesced.
- New requests during work: follow-ups queue behind the current turn; they do not interrupt it. This preview does not start parallel model runs or child agents. Timing does not change the reply target. An answer already saved for delivery keeps its original route; it is not resent or retargeted.
- Where answers appear: every message-triggered answer replies to its source,
including isolated standalone messages. A combined burst's default answer
targets its first source. The bridge owns this decision; the model cannot
select a different target or turn threading off for an answer. Same-conversation
send tools use the same trigger. After an explicit send, Codex returns
[SILENT]to avoid another automatic answer. A final answer exactly matching an observed targeted tool send is suppressed too. - Proactive messages: cron jobs, reminders, and other sends without a current message trigger start fresh, without a reply target. The bridge never picks the last message from conversation history. Scheduled jobs must run independently, without inheriting the active chat's tool-process environment.
- No audience changes: targeted tools must use a visible source in the same conversation; in a running chat they must also belong to that active input. Opaque native thread IDs are not message IDs or new Codex sessions. Replies use the API's message IDs; missing parent/root IDs remain unknown, never inferred from a thread ID. A visible ancestor root can identify a reply even when its immediate parent is unavailable. Native thread-read tools return one bounded chronological page and do not broaden Companion activation history.
- Fallback and failures:
plain_reply_fallback=truedelegates an eligible same-conversation fallback to the API. A timeout or arbitrary send failure never triggers a client-side unthreaded resend. Accepted sends are queued, not proof of delivery. Correlated late failures are retained and added as context to the original session without waking a retrying model turn;/healthreports counts. - Restart behavior: ordinary inputs are saved before acknowledgement. Only
known-unsubmitted inputs and saved, not-yet-sent answers resume automatically.
Interrupted model work and uncertain sends remain retained for inspection, not
replay. Do not delete receipts or resend blindly.
/stopcancels this iMessage conversation's queued work without cancelling a voice consult on the same contact.
Set the variable back to false and restart to restore the previous behavior and
tool schema. This does not erase retained receipts; inspect unfinished work before
re-enabling. Identity, contact rules, approvals, other channels, and voice settings
are unchanged.
Control commands. A handful of slash-commands steer the conversation itself and are handled by the bridge instead of being sent to Codex (works on any channel):
/clear(or/new) — start a fresh conversation: forgets the resumed session, tears down the client, and clears session-scoped permission grants./stop(or/cancel) — interrupt the current turn and drop anything queued, keeping your conversation context intact./resume— texts you back a numbered list of recent Codex conversations (each with a short summary and timestamp); reply with a number to reopen that one. Like/resumein the Codex CLI./status— reports what the bridge is doing for you right now (working, waiting on a reply, or idle) and whether you're in a fresh or ongoing conversation. Read-only; doesn't disturb a running turn./usage— reports Codex rate-limit windows and token summary from app-server account endpoints./health— reports identity reachability, tunnel connectivity, the latest Codex startup check, and Companion queue readiness.
These match only when the whole message is exactly the command, so "please /clear the cache" is still a normal turn.
Errors. Ordinary turns send a short plain-language failure notice. Companion failures are recovered automatically without replaying possibly accepted work; see below.
The bridge automatically retries inputs that failed before a model turn or message send could start, with backoff capped at one minute. Retry does not stop after a fixed number of failures, and a recovered Codex runtime wakes waiting conversations without requiring another message or a restart. Saved answers retry delivery without running the model again.
If the selected npm Codex launcher cannot find Node, the bridge uses the matching
native executable already included in that same Codex installation. It does not
change your configuration, download a replacement, or select an unrelated
installation. Explicit native CODEX_BIN paths and working launchers are kept.
For an ambiguous turn outcome, the bridge stops the old session and checks saved Codex history for a completed turn matching its receipt. A positively matched answer can be delivered without repeating the model turn. When the outcome cannot be established—or delivery may already have happened—the original input is retained as unconfirmed and is not automatically replayed. Later messages continue in the conversation; the next reply can explain the earlier uncertainty. An interrupted initialization does not replay its old trigger. Payloads, deduplication records, and saved conversation threads are preserved across restart.
Outstanding approval prompts are canceled when their Codex connection closes or their host request is resolved. Only one question is displayed at a time, and canceling a task clears its queued questions. New instructions during a tool approval become fresh work. Trusted Inkbox tool auto-approval remains optional.
inkbox-codex doctor checks the effective Codex launcher with a bounded app-server
initialization handshake, using the same saved environment as the gateway.
Recognized startup failures include an exit code and a bounded diagnostic summary,
without raw stderr. HTTP /health remains a liveness endpoint (ok: true, HTTP
200) with separate ready, codex, and companion fields. HTTP /ready returns
503 while startup cannot be verified or a conversation is still blocked.
Startup checks refresh in the background every minute; HTTP health requests do
not start new processes. Queue diagnostics distinguish active blockage from
retained unconfirmed outcomes (quarantined_count), which do not block new input.
They also show unfinished counts and oldest receipt age (unknown for older
receipts). These checks do not prove model generation or end-to-end delivery.
Verify recovery with a fresh message and its reply in the same conversation.
Automatic recovery does not require an operator command. To inspect retained
receipts without message content, use inkbox-codex inbox list. After inspecting
an unconfirmed outcome, an operator can explicitly request retry or retirement
with the gateway stopped:
inkbox-codex stop
# Choose one action for the receipt:
inkbox-codex inbox recover EVENT_ID --action retire --reason "Inspected; skip stale input"
# OR acknowledge that explicitly retrying uncertain work may duplicate effects:
inkbox-codex inbox recover EVENT_ID --action retry --reason "Inspected; retry intended" --acknowledge-duplicate-risk
inkbox-codex startEach explicit decision records its reason, timestamp, and previous/new state. Retirement does not claim that processing or delivery succeeded; retry uses a saved answer when available. Recovery commands refuse to race a running gateway.
The setup wizard has a Phone call voice stack section with three choices:
-
Inkbox Voice AI: Inkbox handles the audio and conversation on Codex's behalf. Choose contact-scoped or YOLO authority during setup. Hosted outbound calls carry a task reason, inherit that saved authority by omitting a per-call override, and notify Codex through a signed
call.endedevent. The initiating turn ends after call placement; the separate call-ended turn owns the follow-up. Codex then fetches the authoritative transcript and executes any remaining post-call commitments in a side-effect-only turn; its plain model prose is never sent after hangup. Saved SMS receipts prevent a correction from repeating an accepted or uncertain send, even when Codex omits the tool result from its summary. -
OpenAI Realtime (when configured): the bridge pre-opens an OpenAI Realtime session and accepts the call in raw-media mode, so a natural, low-latency voice handles the conversation. It runs the call itself and has these tools:
consult_agent— do real work now in the project; runs in the same contact-keyed session as your SMS/iMessage and its answer is spoken back.register_post_call_action/edit_post_call_action/delete_post_call_action— queue, change, or cancel work to run after you hang up.hang_up_call— two-step (say goodbye, then end the call).
When the call ends, queued actions run in your session (and any plain "reflect on the call" follow-up if none were queued) — so "after we hang up, open a PR and text me" actually happens. Enable it in
inkbox-codex setup(it validates your OpenAI key live) or via theINKBOX_REALTIME_*env vars below. -
Inkbox STT/TTS (default): Inkbox auto-accepts the call and opens a WebSocket to the bridge; finalized transcripts become turns in your same session and Codex's replies are spoken back. Realtime may fall back to this if OpenAI cannot be reached (unless
INKBOX_REALTIME_FALLBACK_TO_INKBOX_STT_TTS=false).
Calls — inbound and outbound — can run over either of two lines, and the agent picks the one that matches the channel it's talking on:
- The dedicated phone number. The agent's own number (the same line SMS uses). Outbound calls present this number; inbound calls to it ring the agent.
- The shared Inkbox iMessage line. The agent can also place and receive voice calls with a person it's connected to over iMessage, over the same shared line that person already messages. The underlying number is never surfaced — Inkbox resolves it from the iMessage connection — and it only works for people already connected over iMessage (an unknown caller is rejected; an outbound call with no connection is refused).
Inbound answering is configured once per identity: Voice AI uses hosted_agent; local Realtime and TTS/STT use auto_accept plus the bridge WebSocket. Outbound, the agent sets origination on inkbox_place_call (dedicated_number / shared_imessage_number), or omits it: the bridge then uses the only available line, or — when both exist — the line matching the current conversation's channel.
Besides Inkbox's own events, the webhook endpoint can inject events from outside systems (e.g. a CI failure) to wake the agent on its own external:<source> thread. Routing is by verified source, never by the body's claimed event type:
- Registered providers (e.g. GitHub via
X-Hub-Signature-256) are verified with their own secret fromINKBOX_WEBHOOK_SECRET_<NAME>; registering the provider + setting its secret is the opt-in, and forged signatures are rejected outright. - Mock provider accepts a shared secret directly in
X-Inkbox-Mock-Secret, matched againstINKBOX_WEBHOOK_SECRET_MOCK. It is intended for manualcurlprobes and test systems that cannot calculate a body signature. A valid secret makes the event verified and wakes the agent even when generic external events are disabled. - Everything else (unknown sources, or Inkbox-signed payloads with no handler) is delivered only when
INKBOX_EXTERNAL_EVENTS_ENABLED=true, and unverified events carry a cautious directive that forbids irreversible action on their say-so.
No human reads an external thread, so the agent is told to act via its tools rather than reply. Adding a source is drop-in: a new module in inkbox_codex/webhook_providers/ with a @register_provider class.
Send a mock event with:
curl --fail-with-body --request POST 'https://your-agent-host.example/webhook' \
--header 'Content-Type: application/json' \
--header "X-Inkbox-Mock-Secret: $INKBOX_WEBHOOK_SECRET_MOCK" \
--data '{
"id": "mock-run-123",
"source": "mock-ci",
"event": "workflow.failed",
"title": "Mock workflow failed",
"summary": "Inspect the repository and decide what action is appropriate.",
"requested_action": "Investigate the failure and take any safe corrective action."
}'Automatic email replies use reply-all on the original message. The reply goes
to its Reply-To address (or sender), with the other visible To/CC recipients
included in CC. The agent's own mailbox and BCC recipients are excluded, and the
original email thread is preserved. No mention is required for email replies.
Queued replies, approval prompts, and send-failure recovery retain the original
message's reply target even if another message arrives in the same session.
If an inbound event lacks the original message ID, the bridge cannot send an automatic reply-all; it does not fall back to a sender-only email.
Live reply-all CI uses the existing CODEX_INKBOX_API_KEY and
REMOTE_INKBOX_API_KEY to verify real reply delivery, sender deduplication,
self-exclusion, and reply-thread headers. An optional REPLY_ALL_INKBOX_API_KEY
for a third, non-auto-replying dedicated inbox enables separate CC delivery
checks (including additional original To recipients). Those two additional
cases are explicitly skipped when the third credential is absent; two-identity
checks do not prove independent CC delivery.
The live channel workflow's email_reply_all_only input runs just these checks,
with both deterministic and real-model gateway runs, without SMS reset traffic.
Inbound. When someone sends an MMS image, an iMessage attachment, or an email with files, the gateway downloads them to ~/.inkbox-codex/media/ (override with INKBOX_CODEX_MEDIA_DIR) and appends the local paths to the message, so Codex can open them with its Read tool — including viewing images. Media-only messages (no text) still wake the agent.
Outbound. Codex sends media with a single tool call per channel — it just passes local file paths, and the tool handles any upload-then-send round trip internally:
- Email —
inkbox_send_email(..., attachment_paths=[...])(base64 inline, ~25 MB total). - iMessage —
inkbox_send_imessage(..., media_path=...)(uploaded + sent, ≤10 MB). - SMS/MMS —
inkbox_send_sms(..., media_paths=[...])(uploaded + sent;media_urlsalso accepts already-hosted URLs).
Codex can access credentials and generate current 2FA codes from Inkbox Vault without opening a browser for each request.
- Store the credential in your Inkbox Vault and grant the agent identity access. For 2FA, use a login secret with TOTP configured.
- Set
INKBOX_CODEX_VAULT_KEYto your vault unlock key in the bridge's local environment or its.envfile (normally~/.inkbox-codex/.env). This is separate fromINKBOX_API_KEY. Keep the key out of chat and source control. - Restart the bridge with
inkbox-codex restart, or restart its service if you run it through a service manager. - Ask: "Get the current 2FA code for my Example login from Inkbox Vault."
The agent uses inkbox_list_vault_secrets to find the login, then calls
inkbox_get_totp_code with its secret_id. The result includes code,
period_start, period_end (Unix timestamps), and seconds_remaining. Request
a fresh code if it expires. This tool returns neither the password nor the TOTP seed.
inkbox_get_vault_secret retrieves one credential when the task needs it. Login
results include has_totp instead of the TOTP seed. Listing returns metadata only
and works without an unlock key. The tools follow the API key's access scope, so
use the agent-scoped key configured by setup.
The bridge unlocks the vault only when a credential or 2FA tool needs it. An incorrect bridge vault key or failed unlock returns a tool error; messaging and metadata listing remain available. Correct the key locally and restart, or retry after a temporary service failure.
Use the bridge-specific INKBOX_CODEX_VAULT_KEY rather than the SDK-wide
INKBOX_VAULT_KEY or vault_key in ~/.inkbox/config. Those SDK-wide settings
trigger eager unlocking for every SDK client, including the messaging gateway.
If you followed earlier instructions using INKBOX_VAULT_KEY, rename that setting
and remove any SDK-wide vault key used only by this bridge before restarting.
If the requested secret is missing, check its access grant to this agent. A login without TOTP must have it configured before the agent can generate codes.
| Env var | Required | Default | Description |
|---|---|---|---|
INKBOX_API_KEY |
yes | - | Agent-scoped Inkbox API key. |
INKBOX_IDENTITY |
yes | - | Inkbox agent identity handle. |
INKBOX_SIGNING_KEY |
inbound | - | Webhook HMAC secret for signed inbound events. |
INKBOX_CODEX_VAULT_KEY |
vault reads / 2FA | - | Vault unlock key, supplied locally and used only when a credential or 2FA tool needs it. See Vault and 2FA codes. |
CODEX_PROJECT_DIR |
yes | cwd | Directory Codex works in. |
CODEX_MODEL |
no | CLI default | Model override for bridged sessions. |
INKBOX_REQUIRE_SIGNATURE |
no | true |
Refuse unsigned inbound webhooks unless false. |
INKBOX_SKIP_WEBHOOK_RECONCILE |
no | false |
Leave webhook subscriptions untouched on start. For deployments that provision them ahead of time, where the destination is fixed or this API key may not change it. They must already point at this bridge's webhook URL, or nothing arrives. |
INKBOX_CONTACT_MEMORIES_ENABLED |
no | true |
Add memories supplied with the matched webhook contact as background context. |
INKBOX_BASE_URL |
no | SDK default | Override the Inkbox API base URL. |
INKBOX_PUBLIC_URL |
no | - | Public bridge URL. Omit to use an Inkbox tunnel. |
INKBOX_TUNNEL_NAME |
no | identity handle | Tunnel name override. |
INKBOX_ALLOWED_USERS |
no | - | Local allowlist (emails / E.164 numbers). Usually leave empty and use Inkbox contact rules. |
INKBOX_ALLOW_ALL_USERS |
no | false |
Allow all senders admitted by Inkbox contact rules. |
INKBOX_BRIDGE_PORT |
no | 8767 |
Local webhook server port. |
INKBOX_PERMISSION_TIMEOUT_S |
no | 600 |
Seconds to wait for a permission/poll reply. |
INKBOX_GROUP_REPLY_MODE |
no | auto |
Group SMS/iMessage, Slack, and Companion email replies: auto lets the agent decide; mention requires a native Slack @mention, SMS/iMessage @agent or @<agent-handle>, or Companion email addressed To the agent. Other messages become context without starting a turn. Also configurable in setup. |
INKBOX_IMESSAGE_THREADED_REPLIES |
no | false |
Preview native iMessage replies, ordinary burst collection, and queued follow-ups. Requires a compatible SDK and API; see Native iMessage replies. Environment-only; no wizard change. |
INKBOX_COMPANION_RESPONSE_MODE |
no | safe |
Companion SMS/MMS, iMessage, and email: safe wakes only for direct access; relaxed permits any delivered sender. Both honor Auto/Mention. Sponsored and unknown access stays context-only in Safe mode. Also configurable in setup. |
INKBOX_CODEX_AUTO_APPROVE_INKBOX_TOOLS |
no | false |
Auto-accept Codex MCP prompts for Inkbox tools only. The setup wizard writes true when you trust the agent to send through Inkbox without per-call approval. |
INKBOX_A2A_PROGRESS_INTERVAL_SECONDS |
no | 180 |
Seconds between progress updates for active inbound A2A tasks. Set to 0 to disable periodic updates. |
INKBOX_VOICE_STACK |
no | inkbox_tts_stt |
inkbox_voice_ai, openai_realtime, or inkbox_tts_stt. When absent, legacy Realtime settings remain compatible. |
INKBOX_VOICE_AI_AUTHORITY_MODE |
Voice AI | contact_scoped |
Saved Voice AI authority selected during setup: contact_scoped or yolo. |
INKBOX_VOICEMAIL_DETECTION |
no | enabled |
Outbound-call voicemail policy: enabled or disabled. Live CI uses disabled. |
CODEX_BIN |
no | codex |
Codex CLI executable to run. |
CODEX_SANDBOX |
no | workspace-write |
App-server thread sandbox (read-only, workspace-write, danger-full-access). |
CODEX_APPROVAL_POLICY |
no | on-request |
Codex approval policy for bridged turns. |
INKBOX_REALTIME_ENABLED |
no | false |
Use OpenAI Realtime for calls. Needs a key; off → Inkbox STT/TTS. |
INKBOX_REALTIME_API_KEY |
realtime | OPENAI_API_KEY |
OpenAI key with /v1/realtime access. |
INKBOX_REALTIME_MODEL |
no | gpt-realtime-2 |
Realtime model id. |
INKBOX_REALTIME_VOICE |
no | cedar |
Realtime voice name. |
INKBOX_REALTIME_FALLBACK_TO_INKBOX_STT_TTS |
no | true |
Fall back to Inkbox STT/TTS if OpenAI connect fails. |
INKBOX_EXTERNAL_EVENTS_ENABLED |
no | false |
Wake the agent on unrecognised (external) webhooks — see External events. |
INKBOX_WEBHOOK_SECRET_<NAME> |
per provider | - | Verification secret for a registered third-party webhook provider (e.g. INKBOX_WEBHOOK_SECRET_GITHUB). |
INKBOX_WEBHOOK_SECRET_MOCK |
mock provider | - | Expected plaintext value of the X-Inkbox-Mock-Secret header. Use a random secret. |
The agent reaches you (or third parties) through an in-process MCP server:
inkbox_whoami— its own identity: handle, mailbox, iMessage status, and its two calling lines (dedicated number vs shared iMessage line).inkbox_place_call— place an outbound voice call over either line (origination:dedicated_number/shared_imessage_number) — see Two calling lines.inkbox_list_calls·inkbox_get_call_transcript— browse call history and transcripts.inkbox_send_email— send email; attach local files withattachment_paths.inkbox_send_sms— send SMS/MMS; attach local files withmedia_paths(or hostedmedia_urls).inkbox_send_imessage— send into an iMessage conversation; attach a local file withmedia_path.inkbox_get_imessage_thread·inkbox_get_imessage_conversation_thread— bounded native-thread pages, available only withINKBOX_IMESSAGE_THREADED_REPLIES=true; reply targets and fallback policy are bridge-owned, while the send tool also acceptsidempotency_key.inkbox_list_text_conversations·inkbox_get_text_conversation— browse SMS threads and history.inkbox_list_imessage_conversations·inkbox_get_imessage_conversation— browse iMessage threads and history (find theconversation_idto send into).inkbox_lookup_contact·inkbox_list_contacts·inkbox_get_contact— resolve and read address-book contacts (reverse-lookup by email/phone, free-text search, or full record by id).inkbox_create_contact·inkbox_update_contact·inkbox_delete_contact— save, edit, and remove organization-wide contacts. Changes affect the shared address book. vCard export/import is not exposed.inkbox_list_vault_secrets·inkbox_get_vault_secret·inkbox_get_totp_code: find vault credentials, retrieve one credential, or generate a current 2FA code. See Vault and 2FA codes.inkbox_a2a_call·inkbox_a2a_check·inkbox_a2a_reply— delegate work to another agent and follow its task.inkbox_list_a2a_tasks·inkbox_list_a2a_messages— page and search this identity's inbound and outbound A2A history, with participant, task, context, role, state, and timestamp filters.inkbox_a2a_complete·inkbox_a2a_ask_caller·inkbox_a2a_fail— commit the outcome of a verified inbound A2A task. These tools are rejected outside that task's isolated session.
Inbound A2A tasks acknowledge pickup immediately. While a task remains active, the worker sends a short progress update about every three minutes by default; these updates are visible in task history without starting a requester turn.
The bridge requires Inkbox SDK 0.7.11 or newer (below 1.0.0).
On a live call, the OpenAI Realtime voice agent additionally gets consult_agent, register_post_call_action / edit_post_call_action / delete_post_call_action, and hang_up_call — see Voice.
inkbox-codex doctor— everything green.- Text
START, then text the agent; verify it replies in the same thread. - Ask it to do something requiring a command (e.g. "run the tests") and verify you get a permission text; reply
1and verify the result comes back. - Ask it something open-ended enough to trigger a poll; reply with a number.
- Email the agent; verify the reply lands as an email on the same thread.
- Call the number, ask what it's working on, hang up mid-answer, and verify the late voice tail is not silently sent as SMS or email.
python -m pytest- Tunnel-first inbound: with a signing key, the gateway opens an Inkbox tunnel, reconciles mail/text/iMessage plus
call.endedsubscriptions, and sets the identity's incoming-call action from the selected stack —hosted_agentfor Voice AI orauto_acceptplus the call WebSocket for local stacks. - Session routing: direct messages use the resolved contact id across channels, falling back to the channel conversation or raw address/number. Group SMS uses the group conversation id regardless of sender.
- Escalation over the active channel: a pending permission/poll captures the contact's next inbound message as its answer, on whichever text channel they're using.
- Codex app-server: each contact session owns one
codex app-serversubprocess, one Codex thread, app-server approval request handling over Inkbox, and a local stdio MCP server for the Inkbox tools.
Realtime calls request 16 kHz mono PCM16 call audio. The bridge continuously resamples to and from the realtime session's 24 kHz PCM format, preserving audio across WebSocket frame boundaries. Older call streams that advertise 8 kHz μ-law (or omit their audio descriptor) remain supported. Call audio quality also depends on the remote connection. Hosted voice and managed speech modes are unchanged.