-
Notifications
You must be signed in to change notification settings - Fork 3.8k
feat(self-hosting): report on-prem usage to a Sim instance for valuation #8349
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
myxamediyar
wants to merge
1
commit into
simstudioai:main
Choose a base branch
from
myxamediyar:feat/onprem-usage-telemetry
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
106 changes: 106 additions & 0 deletions
106
apps/docs/content/docs/platform/self-hosting/usage-telemetry.mdx
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,106 @@ | ||
| --- | ||
| title: Usage Telemetry | ||
| description: Report daily workflow and credit usage from a self-hosted deployment to a Sim instance | ||
| --- | ||
|
|
||
| import { Callout } from 'fumadocs-ui/components/callout' | ||
|
|
||
| A self-hosted deployment can report a daily summary of its usage — workflow runs, credits, and model token volume — to a Sim instance, so that instance can see how the deployment is used over time and value that usage at a rate agreed for the deployment. | ||
|
|
||
| The feature is **off by default**. Nothing is collected or sent until the four variables below are set, and the workflow execution path never touches it: reporting runs from a background job that reads the deployment's own database after the fact. | ||
|
|
||
| ## What is sent | ||
|
|
||
| One record per UTC calendar day, re-sent for a trailing window on every run so a day converges on its final figures once it has elapsed. Each record carries: | ||
|
|
||
| | Field | Source | | ||
| |-------|--------| | ||
| | Workflow executions, failures, total duration | Execution logs | | ||
| | Credits | The usage ledger's dollar sum for the day × 200 | | ||
| | Input and output tokens, per model | The usage ledger's token metadata | | ||
| | Events and credits per usage source | The usage ledger (`workflow`, `knowledge-base`, `wand`, …) | | ||
|
|
||
| Nothing else crosses the wire: no workflow inputs or outputs, no credentials, no workflow or user identifiers, and no tool or block names. Aggregation happens inside the database query, so per-execution rows never reach the reporting code. | ||
|
|
||
| <Callout type="info"> | ||
| On a self-hosted deployment, models usually run on your own API keys, so their ledger rows carry **tokens with zero credits**. Credits on such a deployment mostly reflect the per-run execution charge (1 credit per workflow run); token volume is reported alongside so the receiving instance can see model usage that Sim never billed. | ||
| </Callout> | ||
|
|
||
| ## Enable reporting | ||
|
|
||
| Register the deployment on the receiving instance (see [Receiving reports](#receiving-reports) below), then set on the self-hosted deployment: | ||
|
|
||
| ```bash | ||
| ONPREM_TELEMETRY_ENABLED=true | ||
| ONPREM_TELEMETRY_ENDPOINT=https://sim.example.com # the receiving instance | ||
| ONPREM_TELEMETRY_DEPLOYMENT_ID=acme-prod # issued at registration | ||
| ONPREM_TELEMETRY_API_KEY=simot_… # issued at registration, shown once | ||
| ``` | ||
|
|
||
| Optional: | ||
|
|
||
| | Variable | Default | Description | | ||
| |----------|---------|-------------| | ||
| | `ONPREM_TELEMETRY_LOOKBACK_DAYS` | `7` | How many trailing days each run re-sends (max 90). Bounds how long the receiving instance can be unreachable before a day is missed. | | ||
|
|
||
| The report runs from the `onprem-usage-report` background job every six hours (Docker Compose: `docker/crontab`; Kubernetes: `cronjobs.jobs.onpremUsageReport`). To send immediately: | ||
|
|
||
| ```bash | ||
| curl -H "Authorization: Bearer $CRON_SECRET" https://your-deployment/api/cron/onprem-usage-report | ||
| ``` | ||
|
|
||
| The response reports `delivered`, `failed` (with the reason), or `disabled`. | ||
|
|
||
| ## Disable reporting | ||
|
|
||
| Unset `ONPREM_TELEMETRY_ENABLED` (or set it to `false`). The next job run exits before running any query or opening any connection, and stays that way until re-enabled. No restart is needed. This is the right setting for airgapped deployments; the job itself is harmless to leave scheduled. | ||
|
|
||
| A deployment that is enabled but cannot reach the endpoint keeps running workflows normally. The job logs the failure and re-sends the same window next time. | ||
|
|
||
| ## Receiving reports | ||
|
|
||
| The receiving instance is any Sim deployment with the [admin API](/platform/self-hosting/environment-variables) enabled. Register a deployment and (optionally) its rate in one call: | ||
|
|
||
| ```bash | ||
| curl -X POST https://sim.example.com/api/v1/admin/onprem-telemetry/deployments \ | ||
| -H "x-admin-key: $ADMIN_API_KEY" -H "Content-Type: application/json" \ | ||
| -d '{"id": "acme-prod", "name": "Acme production", "usdPerCredit": 0.005}' | ||
| ``` | ||
|
|
||
| The response includes `apiKey` exactly once; only its hash is stored. | ||
|
|
||
| ### Conversion rates | ||
|
|
||
| Each deployment has its own append-only history of credit → dollar rates: | ||
|
|
||
| ```bash | ||
| # Set a new rate, effective now | ||
| curl -X POST https://sim.example.com/api/v1/admin/onprem-telemetry/deployments/acme-prod/rates \ | ||
| -H "x-admin-key: $ADMIN_API_KEY" -H "Content-Type: application/json" \ | ||
| -d '{"usdPerCredit": 0.004}' | ||
|
|
||
| # Backdate a correction | ||
| curl -X POST … -d '{"usdPerCredit": 0.0045, "effectiveFrom": "2026-09-01T00:00:00Z"}' | ||
| ``` | ||
|
|
||
| Rates are never edited or deleted. Credits are the stored fact; dollars are computed when usage is read, from the rate whose `effectiveFrom` most recently precedes each day's start. That defines what a rate change does to history: | ||
|
|
||
| - A rate effective **now** changes the value of days from now on. Earlier days keep the rate that covered them. | ||
| - A rate effective at an **earlier instant** re-values the days from that instant forward the next time they are read. No credit figure changes, and the rate table shows what applied when. | ||
| - Days before a deployment's first rate have no dollar value; the usage response counts their credits as `unvaluedCredits` rather than pricing them silently. | ||
|
|
||
| ### Reading usage | ||
|
|
||
| ```bash | ||
| curl "https://sim.example.com/api/v1/admin/onprem-telemetry/deployments/acme-prod/usage?from=2026-09-01T00:00:00Z&to=2026-10-01T00:00:00Z" \ | ||
| -H "x-admin-key: $ADMIN_API_KEY" | ||
| ``` | ||
|
|
||
| Every row states its `credits`, the `rate` it was valued with (`usdPerCredit` and `effectiveFrom`), and the resulting `usd`, plus the per-source and per-model breakdown the deployment sent. Totals cover the range. `GET …/deployments` lists deployments with their current rate and last report time; `GET …/deployments/{id}/rates` shows the rate history. | ||
|
|
||
| ## Limitations | ||
|
|
||
| - **Cooperative, not enforced.** The deployment controls its own database and can disable reporting; nothing here proves completeness. It is a usage picture, not a metering system. | ||
| - **Chat usage is not included.** With billing disabled, cost callbacks from the Sim agent service are acknowledged without being recorded, so Chat model usage never reaches the ledger on a self-hosted deployment. | ||
| - **The current day is partial** until the first run after UTC midnight; `reportedAt` on each row says when it was last sent. | ||
| - **Ledger scan.** Each run aggregates the trailing window from `usage_log` by `created_at`. On a very large deployment consider adding an index on `usage_log (created_at)`. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,41 @@ | ||
| import { createMockRequest } from '@sim/testing' | ||
| import { authInternalMock, authInternalMockFns } from '@sim/testing/mocks/auth-internal.mock' | ||
| import { createMockFetch } from '@sim/testing/mocks/fetch.mock' | ||
| import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest' | ||
| import { env } from '@/lib/core/config/env' | ||
|
|
||
| vi.mock('@/lib/auth/internal', () => authInternalMock) | ||
|
|
||
| import { GET } from '@/app/api/cron/onprem-usage-report/route' | ||
|
|
||
| describe('GET /api/cron/onprem-usage-report', () => { | ||
| let fetchMock: ReturnType<typeof createMockFetch> | ||
|
|
||
| beforeEach(() => { | ||
| env.ONPREM_TELEMETRY_ENABLED = undefined | ||
| fetchMock = createMockFetch({ json: { accepted: 0 } }) | ||
| vi.stubGlobal('fetch', fetchMock) | ||
| authInternalMockFns.mockVerifyCronAuth.mockReturnValue(null) | ||
| }) | ||
|
|
||
| afterEach(() => vi.unstubAllGlobals()) | ||
|
|
||
| it('requires cron authentication', async () => { | ||
| authInternalMockFns.mockVerifyCronAuth.mockReturnValueOnce( | ||
| new Response(JSON.stringify({ error: 'Unauthorized' }), { status: 401 }) | ||
| ) | ||
| const response = await GET(createMockRequest('GET')) | ||
| expect(response.status).toBe(401) | ||
| }) | ||
|
|
||
| it('answers disabled without any outbound request when telemetry is off', async () => { | ||
| const response = await GET(createMockRequest('GET')) | ||
| expect(response.status).toBe(200) | ||
| await expect(response.json()).resolves.toEqual({ | ||
| success: true, | ||
| status: 'disabled', | ||
| reason: 'ONPREM_TELEMETRY_ENABLED is not set', | ||
| }) | ||
| expect(fetchMock).not.toHaveBeenCalled() | ||
| }) | ||
| }) | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,34 @@ | ||
| import { createLogger } from '@sim/logger' | ||
| import { type NextRequest, NextResponse } from 'next/server' | ||
| import { verifyCronAuth } from '@/lib/auth/internal' | ||
| import { withRouteHandler } from '@/lib/core/utils/with-route-handler' | ||
| import { runOnPremUsageReport } from '@/lib/onprem-telemetry/report' | ||
|
|
||
| const logger = createLogger('OnPremUsageReportCron') | ||
|
|
||
| export const dynamic = 'force-dynamic' | ||
|
|
||
| /** | ||
| * Cron endpoint that reports this deployment's usage to the configured Sim | ||
| * instance. Scheduled in helm/sim/values.yaml (`cronjobs.jobs.onpremUsageReport`) | ||
| * and docker/crontab. | ||
| * | ||
| * The whole feature lives behind this endpoint: it is the only caller of the | ||
| * reporter, and nothing on the workflow execution path imports it. When | ||
| * `ONPREM_TELEMETRY_ENABLED` is unset the reporter returns before touching the | ||
| * database or the network. Re-reports are idempotent on the receiver, so no | ||
| * lock is needed to keep overlapping runs or multiple replicas correct. | ||
| */ | ||
| export const GET = withRouteHandler(async (request: NextRequest) => { | ||
| const authError = verifyCronAuth(request, 'On-prem usage report') | ||
| if (authError) return authError | ||
|
|
||
| const result = await runOnPremUsageReport() | ||
|
|
||
| if (result.status === 'failed') { | ||
| logger.warn('On-prem usage report run failed', result) | ||
| return NextResponse.json({ success: false, ...result }, { status: 502 }) | ||
| } | ||
|
|
||
| return NextResponse.json({ success: true, ...result }) | ||
| }) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,97 @@ | ||
| import { sha256Hex } from '@sim/security/hash' | ||
| import { | ||
| createMockRequest, | ||
| dbChainMockFns, | ||
| queueTableRows, | ||
| resetDbChainMock, | ||
| schemaMock, | ||
| } from '@sim/testing' | ||
| import { beforeEach, describe, expect, it } from 'vitest' | ||
| import { POST } from '@/app/api/onprem-telemetry/report/route' | ||
|
|
||
| const API_KEY = 'simot_test_key' | ||
| const URL = 'http://localhost:3000/api/onprem-telemetry/report' | ||
|
|
||
| const bucket = { | ||
| periodStart: '2026-09-26T00:00:00.000Z', | ||
| periodEnd: '2026-09-27T00:00:00.000Z', | ||
| workflowExecutions: 4, | ||
| workflowExecutionsFailed: 1, | ||
| workflowDurationMs: 1000, | ||
| credits: 4, | ||
| inputTokens: 10, | ||
| outputTokens: 5, | ||
| sources: [{ source: 'workflow', category: 'fixed', events: 4, credits: 4 }], | ||
| models: [], | ||
| } | ||
|
|
||
| function report(body: unknown, headers: Record<string, string> = {}) { | ||
| return createMockRequest('POST', body, { Authorization: `Bearer ${API_KEY}`, ...headers }, URL) | ||
| } | ||
|
|
||
| function validBody(overrides: Record<string, unknown> = {}) { | ||
| return { | ||
| schemaVersion: 1, | ||
| deploymentId: 'acme-prod', | ||
| reportedAt: '2026-09-27T06:15:00.000Z', | ||
| buckets: [bucket], | ||
| ...overrides, | ||
| } | ||
| } | ||
|
|
||
| describe('POST /api/onprem-telemetry/report', () => { | ||
| beforeEach(() => resetDbChainMock()) | ||
|
|
||
| it('rejects a request without a bearer token', async () => { | ||
| const response = await POST(createMockRequest('POST', validBody(), {}, URL)) | ||
| expect(response.status).toBe(401) | ||
| expect(dbChainMockFns.insert).not.toHaveBeenCalled() | ||
| }) | ||
|
|
||
| it('rejects an unknown key', async () => { | ||
| const response = await POST(report(validBody())) | ||
| expect(response.status).toBe(401) | ||
| expect(dbChainMockFns.insert).not.toHaveBeenCalled() | ||
| }) | ||
|
|
||
| it('looks the deployment up by the hash of the key, never the key', async () => { | ||
| queueTableRows(schemaMock.onpremDeployment, [{ id: 'acme-prod' }]) | ||
| await POST(report(validBody())) | ||
| const boundValues = dbChainMockFns.where.mock.calls.flat().map((c) => JSON.stringify(c)) | ||
| expect(boundValues.join()).toContain(sha256Hex(API_KEY)) | ||
| expect(boundValues.join()).not.toContain(API_KEY) | ||
| }) | ||
|
|
||
| it('refuses a body claiming a different deployment', async () => { | ||
| queueTableRows(schemaMock.onpremDeployment, [{ id: 'acme-prod' }]) | ||
| const response = await POST(report(validBody({ deploymentId: 'someone-else' }))) | ||
| expect(response.status).toBe(403) | ||
| expect(dbChainMockFns.insert).not.toHaveBeenCalled() | ||
| }) | ||
|
|
||
| it('rejects a payload that does not match the contract', async () => { | ||
| queueTableRows(schemaMock.onpremDeployment, [{ id: 'acme-prod' }]) | ||
| const response = await POST(report(validBody({ schemaVersion: 2 }))) | ||
| expect(response.status).toBe(400) | ||
| expect(dbChainMockFns.insert).not.toHaveBeenCalled() | ||
| }) | ||
|
|
||
| it('upserts one row per day and reports how many were accepted', async () => { | ||
| queueTableRows(schemaMock.onpremDeployment, [{ id: 'acme-prod' }]) | ||
| const duplicateDay = { ...bucket, workflowExecutions: 9 } | ||
| const response = await POST(report(validBody({ buckets: [bucket, duplicateDay] }))) | ||
|
|
||
| expect(response.status).toBe(200) | ||
| await expect(response.json()).resolves.toEqual({ accepted: 1 }) | ||
| const [rows] = dbChainMockFns.values.mock.calls[0] as [Array<Record<string, unknown>>] | ||
| expect(rows).toHaveLength(1) | ||
| expect(rows[0]).toMatchObject({ | ||
| deploymentId: 'acme-prod', | ||
| periodStart: new Date(bucket.periodStart), | ||
| workflowExecutions: 9, | ||
| credits: '4', | ||
| breakdown: { sources: bucket.sources, models: [] }, | ||
| schemaVersion: 1, | ||
| }) | ||
| }) | ||
| }) |
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
fetchMockwas called, rather than checking an observable outcome. The repository's testing directive says never to write tests that assert mock calls; the same pattern appears in the receiver and reporter tests. Replace these assertions with boundary-level checks before merging to satisfy that requirement.Context Used: CLAUDE.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!