Self-Hosting Grantex
This guide covers running your own Grantex auth service — from a quick local spin-up to a production-grade Kubernetes deployment.1. Quick Start (Dev)
The fastest way to run the full stack locally:
Verify it’s running:
Note: The dev compose exposes database and Redis ports and uses hardcoded credentials. Never use it in production.
2. Generating a Production Signing Key
Grantex signs grant tokens with RS256 by default, or with ES256 whenJWT_SIGNING_ALG=ES256.
Generate the private key once, in PKCS#8 form, and store it securely:
\n between each PEM line:
-----BEGIN PRIVATE KEY-----\n...) and use it as
RSA_PRIVATE_KEY (or EC_PRIVATE_KEY for the EC key).
Instead of supplying keys, SIGNING_KEY_STORE=postgres lets the service generate the key on
first start and store it in platform_signing_keys, encrypted with VAULT_ENCRYPTION_KEY
(see Section 7).
Keepprivate.pemout of source control. The JWKS endpoint (GET /.well-known/jwks.json) exposes only the public key, so tokens remain verifiable after key rotation.
3. Production Docker Compose
Prerequisites
- Docker 24+ with Compose v2
- A domain name with DNS pointing to your server
- TLS certificate (self-signed for testing; Let’s Encrypt for production)
Step 1 — Copy and fill in the env file
.env.prod and replace every change-me-* placeholder with strong randomly generated
values. Set RSA_PRIVATE_KEY to the collapsed PEM from Section 2, and JWT_ISSUER to your
public base URL (e.g. https://auth.example.com).
Step 2 — Provide TLS certificates
Place your certificate and private key at:Step 3 — Start the stack
Architecture
4. Kubernetes / Helm
Prerequisites
- Kubernetes 1.26+
- Helm 3.x
- A managed PostgreSQL instance (RDS, Cloud SQL, Neon, etc.)
- A managed Redis instance (ElastiCache, Upstash, etc.)
- An RSA private key (see Section 2)
Install
Enable Ingress
Use an existing Secret
If you manage secrets externally (Vault, Sealed Secrets, External Secrets Operator):Upgrading
Rollback
5. Environment Variable Reference
This table is a quick-start subset, not an exhaustive schema. Consultapps/auth-service/src/config.ts and .env.example from the exact release you deploy for all feature-specific settings and validation rules.
6. Database Migrations
Migrations run automatically on every startup, and each file is applied once per database. The built-in runner (src/db/migrate.ts) reads all *.sql files from the migrations/ directory in
alphabetical order, applies the ones this database has not seen, and records them in the
schema_migrations ledger (filename, checksum, applied-at). A start that has nothing to apply
touches no table at all.
That matters during a rolling deploy. Postgres takes an ACCESS EXCLUSIVE lock before it
evaluates ADD COLUMN IF NOT EXISTS, so a no-op ALTER TABLE grants … still queues behind
whatever transaction is touching grants — and every reader arriving after it waits behind that
queued request, including /v1/authorize, token exchange and delegation on the instance that is
still serving traffic. With the ledger a repeat start issues no DDL, so it cannot stall anything.
While applying, the runner sets lock_timeout (MIGRATION_LOCK_TIMEOUT, default 2 s) and retries
a few times, so a migration that cannot take its lock fails the boot loudly instead of stalling
the table. The setting is reset before the connection returns to the pool, so no application
statement inherits it. A file whose content changed after it was applied is reported as a warning
and never re-applied — ship a new migration instead. That is a warning and not a failure on
purpose: the edit has already had no effect on this database, and refusing to boot over it would
take the service down for nothing. The same applies to a ledger row whose file is no longer on
disk: it is warned about, because a renamed migration counts as a new pending file and its
statements run again.
Adopting a database that is already at head
A database migrated by a release before the ledger existed has the full schema and noschema_migrations table, so the first start after the upgrade treats all files as pending and
re-executes them. That is safe — every file is idempotent — and against ordinary traffic it takes
a few seconds. But if a single transaction is holding a row in grants for longer than
MIGRATION_LOCK_TIMEOUT, the ALTER TABLE grants files cannot take their lock and the boot
fails. Nothing is corrupted and no traffic is affected (migrations run before the server
listens, so the new instance never becomes ready and the old one keeps serving), but the deploy
is broken and has to be retried.
Baselining removes that risk. It records every file as applied without executing any of them,
so the upgrade’s first start is a no-op like every start after it:
The command runs from the built service, so run it inside the image you are about to deploy
rather than from a source checkout (dist/ does not exist until npm run build). It needs the
same DATABASE_URL as the service and nothing else:
--dry-run writes nothing at all — not even the ledger table — and prints the verdict:
--dry-run to record. Then deploy: the new instance logs applied 0 and takes no lock on
any table.
It checks the precondition itself. Before recording anything it reads every
CREATE TABLE IF NOT EXISTS and ALTER TABLE … ADD COLUMN IF NOT EXISTS out of the migration
files and confirms each object exists in the database. A database that is behind is refused, with
the missing objects named:
schema_migrations by hand. --dry-run exits non-zero on such a
database, so it can be used as a pre-deploy check in a script.
Rules:
- Run it only against a database whose schema is already at head — and let the command confirm that rather than taking it on trust.
- It is not needed for a new database. Start the service and it applies everything itself.
- It is safe to repeat: files already in the ledger are left alone.
- If you skip it, the upgrade still works — retry the deploy at a quieter moment, or during a short maintenance window.
migrations/ directory of the release you are deploying is the only authoritative list of what will be applied; a number copied into this page goes stale on the next merge, so there is none here. The files cover core authorization, webhooks, policy, enterprise identity, credentials, budgets, offline operation, trust registry, DPDP, commerce, MCP certification-state integrity, query-performance indexes, agent prepaid wallets, layered wallet spend controls, the event bridge and the revocation feed. Index builds use CREATE INDEX CONCURRENTLY, and the runner serializes migrations across service instances with a PostgreSQL advisory lock. An index a cancelled concurrent build left INVALID is dropped before the file that creates it is retried, because CREATE INDEX CONCURRENTLY IF NOT EXISTS matches such an index by name and would otherwise skip it forever.
Upgrade procedure — just restart the service:
7. Key Rotation
GET /.well-known/jwks.json publishes every platform signing key with kid, alg and
use: "sig". Verifiers select the key by kid and refuse a key whose type does not match the
token’s algorithm. No step below invalidates an outstanding token: a key leaves the JWK Set only
after the tokens it signed have expired.
Key ids
A key’skid is its RFC 7638 thumbprint, grantex-rs256-… or grantex-es256-…, so every
instance publishes the same kid for the same key whenever it started.
Before 0.6 the RS256 kid was grantex-YYYY-MM of the month the process started. Tokens carrying
such a kid keep verifying:
- the auth service verifies an RS256 token whose
kidisgrantex-YYYY-MM, or that has nokid, with the legacy key —RSA_PRIVATE_KEY, or the key named byJWT_LEGACY_KID_KEY; - the JWK Set also publishes the legacy key under
grantex-YYYY-MMfor the current month and the previousJWT_LEGACY_KID_MONTHS - 1months, so SDK verifiers find it; - for
SIGNING_KEY_ACTIVATION_DELAY_SECONDSafter start, an instance still signs with the legacy kid, so resource servers holding a JWK Set fetched from a pre-0.6 instance keep accepting new tokens until they refresh it.
JWT_LEGACY_KID_KEY if you replace
it, until pre-0.6 tokens have expired.
Postgres key store
Switching from the env store. SetSIGNING_KEY_STORE=postgres and keep the existing key
settings for the first start. Every instance imports them: the env signing key becomes the stored
active key (same kid, so nothing changes for verifiers), and the other configured keys are stored
as retired public keys, keeping the legacy kid marker. Once the table holds them, the private key
settings can be removed.
Rotation is publish-then-sign:
SIGNING_KEY_ACTIVATION_DELAY_SECONDS the next reload makes it the signing
key and retires the previous key, erasing its stored private key. The retired public key stays in
the JWK Set for SIGNING_KEY_RETIRED_GRACE_SECONDS, and the legacy kid key for the legacy alias
window. A second rotation is refused while one is pending. The stored active key is authoritative:
instances with a different JWT_SIGNING_ALG keep using it.
Set MAX_GRANT_LIFETIME_SECONDS; start-up refuses a grace window shorter than it, and without it a
warning says grants may outlive their key.
Erasure limits. Retiring a key sets its encrypted private key to NULL. The ciphertext can
remain in dead tuples until vacuum, in WAL and replicas, and in backups for their retention. It is
encrypted with VAULT_ENCRYPTION_KEY and bound to its kid, so it is useless without that key.
If a private key may have been exposed, rotate at once, and rotate VAULT_ENCRYPTION_KEY as part
of the response.
Env key store
To replace a key (RSA to RSA, EC to EC, or a change of algorithm):- Publish the new key. Add its public JWK, with its thumbprint
kidandalg, toJWT_VERIFICATION_PUBLIC_KEYS(for a change of algorithm you can instead set the other private key setting, for exampleEC_PRIVATE_KEYwhileJWT_SIGNING_ALG=RS256). Restart and wait at leastSIGNING_KEY_ACTIVATION_DELAY_SECONDSso verifiers see it. - Sign with it. Set the new private key (and
JWT_SIGNING_ALGif it changes). Keep the old key verifiable: add the old public JWK toJWT_VERIFICATION_PUBLIC_KEYS(the entry for the new key may stay; the same key listed twice is one key). If the old key is an RSA key that signed pre-0.6 tokens, setJWT_LEGACY_KID_KEYto its thumbprintkid. Restart. - Clean up only after every token the old key signed has expired: remove its entry from
JWT_VERIFICATION_PUBLIC_KEYS, and unsetJWT_LEGACY_KID_KEYonce pre-0.6 tokens have expired.
8. Health Checks & Monitoring
Health endpoint
200 when the service is up and connected. The Docker Compose healthcheck and
Kubernetes liveness/readiness probes both use this endpoint.
Structured logging
All logs are emitted as JSON to stdout, compatible with Datadog, Loki, and CloudWatch Logs. No configuration needed — just forward stdout from your container runtime.Prometheus metrics
WhenMETRICS_ENABLED=true (the default), the auth service exposes Prometheus text at
GET /metrics. The endpoint is unauthenticated and limited to 10 requests per minute per IP;
restrict it at your network boundary if metrics must remain private.
9. Backup & Recovery
PostgreSQL
Back up withpg_dump:
Redis
Redis holds ephemeral token metadata and rate-limiting state — not primary data. For durability enable AOF persistence:503 RATE_LIMIT_UNAVAILABLE
because their per-developer rate limit cannot be counted — once the Redis client gives up on
the command, which against a stopped Redis took more than a minute. Revoking a grant, token,
passport or consent bundle and the emergency stop are the exception: after at most 500 ms
they are counted in each instance’s memory against the containment ceiling instead, and the
revocation is committed, so an incident can be contained during a Redis outage. The response
can still wait on the best-effort cache write that follows the commit.
grantex_rate_limit_decisions_total{bucket="containment",outcome=~"local_.*"} shows it
happening.
10. Production Readiness Checklist
Before going live, verify each item:-
RSA_PRIVATE_KEYis a real 2048-bit (minimum) RSA key — notAUTO_GENERATE_KEYS=true -
POSTGRES_PASSWORDandREDIS_PASSWORDare strong, randomly generated values (e.g.openssl rand -hex 32) -
SEED_API_KEYandSEED_SANDBOX_KEYare not set in production - TLS is enabled end-to-end — nginx terminates HTTPS; internal services are on a private network with no exposed ports
- Database and Redis ports are not exposed to the public internet
-
JWT_ISSUERmatches your public base URL exactly — clients validate this claim during token verification - Automated database backups are scheduled and have been tested with a restore
- Health checks are wired into your load balancer or uptime monitor
- CPU and memory limits are set to prevent runaway containers
- Log forwarding is configured (stdout → your observability stack)
11. Emergency Stop (Runbook)
One call halts every agent under a grant, an agent, a principal or a whole developer. Use it when an agent is doing damage, a provider credential has leaked, or a tenant must be stopped now and questions asked afterwards. Withlockout: true the same call also freezes issuance under that scope until the
freeze is lifted.
It is off unless EMERGENCY_STOP_ENABLED=true. Grants stopped this way are
revoked, not paused: there is no undo, and the principals involved have to
authorise again.
A sweep, and a lockout only when you ask for one
Withoutlockout, it is a sweep, not a lockout. It revokes what exists —
repeatedly, until the scope comes back empty, so a grant delegated while it
runs is caught by a later sweep — and then it is finished. It does not
prevent new grants from being issued a second later. Anyone still holding the
developer’s API key can call POST /v1/authorize and mint another one, and
POST /v1/agents to register another agent. The response says
"lockout": false.
With "lockout": true, it also freezes issuance. The stop records a freeze
over its scope before it sweeps, in the same transaction as its own record,
and until the freeze is lifted nothing is issued under that scope. These are
refused with 403 ISSUANCE_FROZEN (403 access_denied on the OAuth
endpoints):
POST /v1/authorize, andPOST /v1/tokenfor a code approved before the stop (the code is not consumed, so it works again once the freeze is lifted);POST /v1/token/refreshandPOST /v1/grants/delegate;- the OAuth profile’s
POST /oauth/parandPOST /oauth/token(authorization code, refresh token and token exchange); POST /v1/consent-bundles,POST /v1/consent-bundles/:id/refreshandPOST /v1/passport/issue.
agent or principal lockout does not stop the key registering a new
agent and asking for grants for it. A developer lockout covers every path
listed above for every agent and principal of the tenant, including agents
registered after it. Commerce passports and decision grants are outside any
lockout (“What a lockout does not refuse”, below). The response says
"lockout": true and gives the freezeId.
So an incident that starts with a leaked credential:
- Stop with a lockout, as the platform operator if you can
(
POST /v1/admin/emergency-stop). Only the operator can lift a lockout the operator placed. A lockout the tenant places with its own key can be lifted by any key of that tenant, including the leaked one. - Rotate or disable the leaked credential:
POST /v1/keys/rotatefor a developer API key, or remove the agent (DELETE /v1/agents/:id). The platform operator can also disable the developer. - Lift the lockout once the credential is safe (“Lifting a lockout”, below).
status tells you how the sweep ended:
completed;incomplete: grants were still appearing after five sweeps. Something is still issuing them; add a lockout or go back to step 2;failed: a batch did not finish. The row records what was revoked before it stopped, and the call is safe to repeat.
EMERGENCY_STOP_ENABLED is not true, no issuance path reads the freeze
state. Turning the flag off while a freeze is in force therefore stops
enforcing it. The freeze stays recorded and is enforced again when the flag
comes back. Lift freezes before turning the stop off. If the freeze state
cannot be read, issuance fails closed: every path answers
503 FREEZE_STATE_UNAVAILABLE (503 temporarily_unavailable on the OAuth
endpoints) and logs alert: "issuance_freeze_unavailable".
Before the incident
- Keep the revocation feed on (
REVOCATION_FEED_ENABLEDunset ortrue; it is on by default from the next release) and make sure the agents you need to stop check revocation:revocationCheck: 'online'(the default from the next SDK release) or'feed'. An agent checking neither ('offline', or a client of@grantex/sdk0.7.0 orgrantex0.6.0 and earlier, which do not check revocation at all) keeps working with the token it already holds until that token expires — the stop revokes the grant, but nothing tells that agent. - Keep grant lifetimes short enough that the tokens of an agent you cannot reach expire in a time you can live with.
- Rehearse it:
scripts/revocation-release-test.shruns agents under a grant tree, stops them and measures how long each kept working. Every production release should have rehearsed it (PRD section 10).
Stopping
Work out the blast radius first — same call,dryRun: true. Nothing is
revoked; the rehearsal itself is recorded, as a row with dryRun: true,
so GET /v1/emergency-stops shows who has been measuring the blast radius of
a tenant and when:
dryRun. confirm must be exactly
stop <type>:<id> — stop agent:ag_01... for the call above. Anything else
is refused with 412 CONFIRMATION_REQUIRED. The expected phrase is not
echoed back: the point of the confirmation is that the caller knows what they
are stopping, which is lost if the endpoint hands them the answer to paste.
To freeze issuance as well, add
"lockout": true to the real call. It must be
a boolean; anything else is refused with 400 before anything is recorded. A
dry run never freezes anything, and says lockout: false.
The response names the stop (stopId), its status, how many sweeps it took,
how many grants matched and were revoked, which agents were stopped, and
lockout: true with the freezeId when the stop placed a lockout, or
reaffirmed one already in force over the same scope, and false otherwise.
agentsStopped lists at most 100 ids; when more were stopped,
agentsStoppedTruncated is true and agentsStoppedTotal gives the real
number.
Suspended grants are swept up too. A suspension is reversible; a stop is
not. Every grant the scope covers with status active or suspended is
revoked permanently, and the suspension bookkeeping that POST /v1/grants/:id/resume needs goes with it. If a subtree is suspended pending
an investigation and you stop its scope, that investigation’s subject cannot
be resumed afterwards — the principals must authorise again.
As the platform operator, use POST /v1/admin/emergency-stop with
ADMIN_API_KEY and the same body plus developerId (not needed for a
developer scope, where the scope names it). A developer API key can only ever
stop, or freeze, its own grants.
What happens
Withlockout: true, the freeze goes in first, in one transaction with the
stop’s emergency_stops row and a grantex.issuance_frozen entry on the
developer’s audit hash chain. From then on, nothing new is issued under the
scope. An issuance that had already passed its check when the freeze arrived
is waited for, and the sweep finds what it wrote: a new grant, which it
revokes, or a passport’s credential, whose status bit it sets. The same holds
for the verifiable credential a code exchange or a delegation issues after its
grant is committed (credentialFormat: "vc-jwt" or "both"): it is written in
a transaction of its own that reads the grant and the freeze again under the
same lock. If a lockout committed in between, the call is refused with
403 ISSUANCE_FROZEN and no credential is written; the grant it had just
created is revoked by the stop’s sweep. If the grant was revoked in between
without a lockout, the call returns the grant token without the credential, as
a failed best-effort issuance always has. Then, with or without a lockout:
- Every matched grant and everything delegated beneath it is revoked in one transaction per batch, with wallet reservations released and credential revocation started. The scope is then read again and swept until it comes back empty, so a grant delegated mid-stop is caught.
- One audit entry per grant (
grantex.grant.revoked, causeemergency_stop) plus a summary entry (grantex.emergency_stop) go on the developer’s audit hash chain, and a row goes intoemergency_stops. - Each revocation reaches the revocation feed in the same transaction, so SDKs in feed mode deny the agents’ next calls — measured in well under a second on a local stack, with two seconds as the requirement.
- The auth service logs
alert: "emergency_stop", andgrantex_emergency_stops_total{scope,outcome}andgrantex_grant_revocations_total{cause="emergency_stop"}move. A lockout also movesgrantex_issuance_freeze_changes_total{action,scope}. Every refusal it causes movesgrantex_issuance_refusals_total{path,reason}and logsalert: "issuance_frozen".
Lifting a lockout
A freeze stays in force until someone lifts it. There is no expiry, and repeating the stop does not lift it. Lift it once the credential it was containing is rotated or disabled:confirmmust be exactlyunfreeze <type>:<id>. It is deliberately not the stop’s phrase, so pasting the stop’s confirmation lifts nothing. It is not echoed back either. A mismatch is412 CONFIRMATION_REQUIRED.404 NOT_FROZEN: no freeze is in force for exactly that scope. A freeze is lifted by the scope it was placed with; adeveloperfreeze is not lifted by unfreezing one agent.403 FREEZE_HELD_BY_OPERATOR: the platform operator placed it, or reaffirmed it with a stop of its own, so only the operator can lift it. The operator usesPOST /v1/admin/emergency-stop/unfreezewithADMIN_API_KEYand the same body, plusdeveloperIdfor agrant,agentorprincipalscope. The operator can lift any freeze.- The response is the lifted freeze:
freezeId,scope, thestopIdthat placed it,placedBy,frozenAt,clearedAt,clearedByandclearReason. The row is kept, so the table is also the history of every lockout. Agrantex.issuance_unfrozenentry goes on the audit chain in the same transaction. GET /v1/emergency-stopslists the freezes still in force underfreezes, beside the stops, oldest first, 50 to a page.freezesTotalis how many are in force in all; when it is larger than the page, ask for the next one with?page=2, or for up to 200 at a time with?pageSize=200, as on the other paged lists. A stop’s ownlockoutfield says whether it asked for one, not whether that freeze is still in force.- A second stop with a lockout over the same scope reaffirms the freeze in force rather than stacking another. If the operator does it, the freeze becomes the operator’s.
Afterwards
- Check
statusinGET /v1/emergency-stops: anything other thancompletedmeans the sweep did not finish cleanly, and the row says how far it got. - Check that agents stopped:
grantex_revocation_feed_entries_totaland the agents’ own denial logs (grant_revoked). - Any agent still running is one that is not watching the feed. Rotate or block its credentials, or wait out the token lifetime.
- Check
freezesin the same response, and every page of it whenfreezesTotalis larger than the page. A lockout you placed is still refusing issuance until you lift it. - To restore service, lift any lockout, and then the principals authorise again; the revoked grants cannot come back.
- Keep the
stopId: the audit entries, theemergency_stopsrow, the freeze and the feed entries all carry it.
What the stop cannot see
It revokes grants. Anything already handed out and cached elsewhere is outside its reach:- Decision grants and passports already issued keep verifying until they
expire; they are signed artefacts. Whatever consumes them has to check
revocation itself. For an agent passport (
POST /v1/passport/issue) that check works: the sweep sets its status-list bit when it revokes the grant behind it. Decision grants and commerce passports are not touched by the sweep. - Work already in flight — a tool call the agent has already made, a payment already authorised downstream — is not recalled. The stop denies the next call.
- An agent that checks neither the feed nor the status endpoint keeps using the token it holds until that token expires.
- Anything below the depth your rehearsal covered. The release test exercises three levels of delegation; deeper chains are handled by the same recursive query, but they are not measured.
- Agents beyond the hundredth: the response and the audit summary name at
most 100. The summary also carries
agents_stopped_total, so the true number is written down, but the list is not. - What a lockout does not refuse. It refuses new issuance only; tokens
already held keep working until the sweep revokes their grants. It does not
stop the key registering new agents: under an
agentorprincipallockout, a new agent can still be given grants. It does not refuse decision grants, which are issued when an approver signs in and approves, not by the tenant’s key. It does not refuse resuming a suspended grant that a sweep which did not finish left behind; repeat the stop, which revokes it. It does not refuse commerce passports (POST /v1/commerce/passports/exchange), which a commerce tenant’s agent mints from a consent the shopper approved rather than from a grant, and the stop does not sweep them either. Contain those with the commerce tenant’s own control: disable the tenant (PATCH /v1/commerce/tenants/:tenant_idwith{"status": "disabled"}, as the platform operator or the tenant’s owner), which refuses new ones with403 tenant_disabled, and revoke any already issued withPOST /v1/commerce/passports/revoke.
If the stop itself fails
403 FEATURE_DISABLED/404:EMERGENCY_STOP_ENABLEDis nottrueon the instance you reached.412 CONFIRMATION_REQUIRED: theconfirmphrase does not match.429: the stop’s own limits — 20 calls a minute from one address, and the developer’s containment budget of 2,000 revocation calls a minute, which ordinary traffic does not use up. Wait outRetry-After; the stop is idempotent. A Redis outage does not refuse the stop.- A 5xx: the stop is idempotent — run it again. Grants already revoked are left alone, and a partly finished stop finishes on the retry. A lockout the failed call had already placed is still in force; the retry reaffirms it rather than adding another.
- If the API cannot be reached at all, revoke at the database
(
UPDATE grants SET status = 'revoked', revoked_at = NOW() WHERE …): the feed triggers fire on that too, so agents still find out. The audit chain will not record it, so write it up. A row inserted straight intoissuance_freezesis enforced as a lockout too. It bypasses the audit chain and the lock that keeps a concurrent issuance from slipping past a freeze, so repeat the stop withlockout: trueas soon as the API is back.