Skip to content
TheCryptoDonkeyPublic

About

Lightning-paid AI inference - monetise any OpenAI-compatible endpoint in 30 seconds

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

satgate

MIT licence Nostr TypeScript Node

Your GPU is burning money. Make it earn money.

satgate sits in front of Ollama, vLLM, llama.cpp or any other OpenAI-compatible backend (with UPSTREAM_API_KEY for one that needs a key) and turns it into a pay-per-token API. No accounts. No API keys. No Stripe. Clients pay per token, you earn sats before the response finishes streaming.

satgate demo

Quick start

LIGHTNING_KEY=<phoenixd password> npx satgate --upstream http://localhost:11434 --lightning phoenixd

satgate auto-detects your models, charges per token over Lightning (here through a local phoenixd; lnbits, lnd, cln and nwc work too), and proxies paid inference requests.

Without --lightning (or Cashu mints, or an allowlist) satgate runs in open mode: no payment and no authentication. Open mode stays on localhost; it only gets a public tunnel if you pass --tunnel.


Try it live

A public instance is running at satgate.forgesworn.dev. Open it in a browser for the chat playground, or use curl:

# 250 sats of free usage per day per IP — after that you'll get a 402 + invoice
curl -s -w '\n%{http_code}\n' https://satgate.forgesworn.dev/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3:0.6b","messages":[{"role":"user","content":"What is Bitcoin?"}]}'

# Check pricing
curl -s https://satgate.forgesworn.dev/.well-known/l402 | jq .

# Machine-readable description
curl -s https://satgate.forgesworn.dev/llms.txt

The old way vs satgate

The old way With satgate
Sell GPU time Sign up for a marketplace (OpenRouter, Together). They set the price, take a cut, own the customer. npx satgate --upstream http://localhost:11434 --lightning phoenixd. You set the price. You keep 100%.
Handle billing Stripe account, KYC, usage tracking, invoices, chargebacks Payments settle before the response finishes streaming. No accounts, no disputes.
Serve AI agents OAuth flows, API key management, billing portals — none of which machines can use Agents discover your endpoint, pay per token from their own wallet, no human in the loop.
Price fairly Flat rate per request, regardless of whether it's 10 tokens or 10,000 Billed from the token counts the upstream reports. Unused holds credited back.

Built for machines

satgate doesn't just serve humans with curl. It's designed for AI agents that pay for their own resources.

Every satgate instance exposes three discovery endpoints — no auth required:

Endpoint Who reads it
/.well-known/l402 Machines — pricing, models, payment methods as structured JSON
/llms.txt AI agents — plain-text description of what you're selling
/openapi.json Code generators — full OpenAPI spec

Pair with 402-mcp and an AI agent can autonomously discover your endpoint, check your prices, pay from its own wallet, and start prompting — no human involved.

sequenceDiagram
    participant A as AI Agent
    participant M as 402-mcp
    participant T as satgate
    participant G as Your GPU

    A->>M: "Use this inference endpoint"
    M->>T: GET /.well-known/l402
    T-->>M: Pricing, models, payment methods
    M->>T: POST /v1/chat/completions
    T-->>M: 402 + Lightning invoice
    M->>M: Pay invoice from wallet
    M->>T: Retry with L402 credential
    T->>G: Proxy request
    G-->>T: Stream response
    T-->>M: Stream completion
    M-->>A: Response
Loading

The secret

Everything you just saw — the payment gating, the multi-rail support, the credit system, the free tier, the macaroon credentials — that's not satgate. That's toll-booth.

satgate is a thin layer on top of toll-booth. It adds the AI-specific bits: token counting, model pricing, streaming reconciliation, capacity management. Everything else comes from the middleware.

You could build your own satgate for your domain in an afternoon.

Monetise a routing API. Gate a translation service. Sell weather data per request. toll-booth handles the payments — you just write the product logic.

→ See toll-booth

graph TB
    subgraph "satgate"
        TC[Token counting]
        MP[Model pricing]
        SR[Streaming reconciliation]
        CM[Capacity management]
        AD[Agent discovery]
    end
    subgraph "toll-booth"
        L402[L402 protocol]
        CR[Credit system]
        FT[Free tier]
        PR[Payment rails]
        MA[Macaroon auth]
    end
    TC --> L402
    MP --> CR
    SR --> CR
    CM --> L402
    AD --> L402
Loading

What satgate adds

  • Pay-per-token: billed from the prompt and completion token counts the upstream reports, streaming and buffered, reasoning tokens included. If an upstream reports no usage, satgate falls back to a chunk count with a byte-based floor.
  • Model-specific pricing — 1 sat/1k for Llama, 5 sats/1k for DeepSeek. You set the rates.
  • Streaming reconciliation: the worst case for the request (its size plus max_tokens) is held up front and settled to actual usage after. The unused hold goes back to the credit balance.
  • Capacity management — limit concurrent inference requests to protect your GPU.
  • Auto-detect models — queries your upstream on startup. No manual model list.
  • Four server-side payment rails — Lightning, Cashu ecash, LNURLcash bearer notes, and x402 stablecoins. The Lightning rail can run on phoenixd, LNbits, LND, CLN or any NWC wallet (--lightning nwc, URI read from a file). Callers can pay from their own NWC wallet through 402-mcp without disclosing it to satgate.
  • Privacy by design: no accounts and no cookies, and client IPs are not logged. IPs are used for free-tier and invoice limits and stored only as hashes keyed to the date.
  • Instant public URL — when payment or an allowlist is required, satgate spawns a Cloudflare quick tunnel (if cloudflared is installed), so your GPU is reachable from the internet in seconds. Open mode never tunnels unless you pass --tunnel.

How it works

sequenceDiagram
    participant C as Client
    participant T as satgate
    participant G as Your GPU

    C->>T: POST /v1/chat/completions
    T-->>C: 402 + Lightning invoice (estimated cost)
    C->>C: Pay invoice
    C->>T: Retry with L402 credential
    T->>G: Proxy request
    G-->>T: Stream response
    T->>T: Count actual tokens
    T-->>C: Stream completion
    T->>T: Reconcile: credit back overpayment
Loading

Before the upstream is called, satgate holds the most the request could cost: every byte of the body counted as a prompt token, plus max_tokens (clamped to --max-tokens, default 2048) for each choice. A balance that cannot cover that gets a 402. After the response, the charge is settled to the tokens the upstream reported, rounded up to the next sat, and the rest of the hold goes back to the client's balance.

Per-request IETF Payment charges and the free tier cannot be topped up or refunded, so for those max_tokens is shrunk to fit what was paid, and the per-request price (estimatedCostSats) is sized by default to cover a 4 KiB request plus a full max_tokens completion.


Paying with a bearer note

--lnurlcash-mints mint.example.com turns on the LUD-25 rail. A 402 then carries an X-LNURLcash challenge naming the price and the mints this booth will take a note from, and a caller pays by putting the note's URL in an X-LNURLcash header on the retry.

Settlement is one rotate at the mint. That single call proves the note is live, moves it to this server, and burns the secret the caller sent - so a replayed note fails at the mint rather than at a replay table here, and nothing has to be remembered to make that true.

With --lightning configured, the note is then melted to your own node for the whole of what it was worth: the mint pays the invoice out of the note and covers routing from its own fee. Without a Lightning backend the rail still works and satgate says so loudly at startup - notes are accepted and then discarded, which is a choice, not a default.

satgate --upstream http://localhost:11434 \
        --lightning phoenixd \
        --lnurlcash-mints mint.example.com

Accepted mints are matched by HOST, so a bare host, a host:port and any URL on that host all name the same mint. They are published on /.well-known/l402 under payment.lnurlcash, so a caller can find out what is accepted without provoking a 402 first.

On the production deploy the setting is LNURLCASH_MINTS in /opt/satgate/deploy.conf, read and validated by deploy/production.sh.


Configuration

Zero config works (just --upstream). For production, create satgate.yaml:

upstream: http://localhost:11434
port: 3000
pricing:
  default: 1          # 1 sat per 1k tokens
  models:
    llama3: 1
    deepseek-r1: 5
freeTier:
  creditsPerDay: 250
capacity:
  maxConcurrent: 4    # default 8; 0 = unlimited
maxTokens: 2048       # cap on completion tokens per request (default 2048)
trustProxy: true      # behind a reverse proxy: read client IPs from X-Forwarded-For
trustedProxies:
  - 127.0.0.1
maxPendingInvoicesPerIp: 20   # unpaid invoices per client before a 429 (default 20)
lnurlcash:
  mints:
    - mint.example.com

With a payment rail configured, storage defaults to SQLite (./satgate.db) and a generated macaroon root key is kept beside it in satgate.root-key, so balances and credentials survive a restart. Set ROOT_KEY to manage the key yourself.

CLI flags > environment variables > config file > defaults.


Examples

The examples/ directory contains runnable scripts and config templates:

Release deployment

deploy/production.sh is the canonical server-side deploy script. It accepts only an exact release tag through DEPLOY_REF, builds that tag in an isolated Git worktree, refuses missing Lightning credentials, persists the Nostr announcement identity, rolls back failed containers, and records the deployed tag and commit.

The server-local /opt/satgate/deploy.conf contains non-secret runtime identity and the models provisioned by the operator:

PUBLIC_URL=https://service.example.com
ANNOUNCE_RELAYS=wss://relay.example.com,wss://relay2.example.com
OLLAMA_MODELS=qwen3:0.6b,gemma3:4b

The script deliberately refuses to install mutable model images or use a hard-coded credential fallback during an application deployment.

Production acceptance

The public deployment completed a controlled mainnet L402 acceptance on 14 August 2026: one 10-sat invoice, a 1-sat payer routing fee, cryptographically verified settlement, and an HTTP 200 paid inference response. The receiver recorded the same 10 sats and the Lightning channel remained Normal.

See the full acceptance record, including the safety envelope and the limits of what this single run proves.


Get started

# Monetise your local Ollama
LIGHTNING_KEY=<phoenixd password> npx satgate --upstream http://localhost:11434 --lightning phoenixd

# Or point at another OpenAI-compatible backend
LIGHTNING_KEY=<phoenixd password> npx satgate --upstream http://your-vllm-server:8000 --lightning phoenixd

→ toll-booth — the middleware that powers all of this. Build your own. → 402-mcp — give AI agents a wallet. Let them pay for your GPU.


Built by @TheCryptoDonkey.

  • Lightning tips: profusemeat89@walletofsatoshi.com
  • Nostr: npub1mgvlrnf5hm9yf0n5mf9nqmvarhvxkc6remu5ec3vf8r0txqkuk7su0e7q2

MIT

About

Lightning-paid AI inference - monetise any OpenAI-compatible endpoint in 30 seconds

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages