Operating a node
This is the operator’s page: what runs where, how the node is reached, how it is
backed up, and what to do when something is lost. Design rationale lives in
SPEC.md; this page only tells you what to run.
Layout
Section titled “Layout”Path (in data_dir, /data in the container) |
What |
|---|---|
hdtp.db |
the SQLite store: accounts, contacts, messages, integrations, audit chain |
blobs/ |
content-addressed media |
keyring.key |
the keyring master key (0600). Every account’s private key is sealed under it. |
hdtp.lock |
held while serve runs; offline commands refuse to run while it is held |
admin.sock |
the admin socket the CLI talks to while the node is running |
Configuration: env HDTP_* > config.json > defaults (hdtp-gateway doctor prints the
resolved result). The deployment mode is derived from the tunnel adapter, never
declared (SPEC §10.1).
Reachability
Section titled “Reachability”| You have | Use | Mode |
|---|---|---|
| a public port / static IP / VPS | tunnel: direct (default) |
direct |
| a Tailscale account | tunnel: tailscale — tsnet Funnel, free, ports 443/8443/10000 |
direct |
an frps you run on any VPS |
tunnel: frp — raw SNI/tcp passthrough |
direct |
| a paid ngrok plan | tunnel: ngrok — tls:// endpoint |
direct |
| a Cloudflare account | cloudflare edge adapter (P4-05) — Cloudflare terminates TLS |
edge |
| a domain of your own | run a second hdtp-gateway in the ingress role on a VPS (below) | direct (passthrough) or edge (terminate) |
| nothing inbound at all | one of the tunnels above, or a provider hosting the identity under a leaf you issue (HDTP §9). There is no relay role: it went on 2026-09-18, because a store-and-forward gateway sees every sender, recipient and timestamp | — |
Direct mode keeps mTLS end to end: callers’ client certificates reach the node. Edge
mode cannot — a third party terminates TLS — so the node forces seal=required
and client_cert=off, and callers are identified by their sealed-envelope
signatures. hdtp-gateway doctor reports the derived mode and probes your
advertised endpoint; a wrong_cert verdict means something between the caller and
the node is terminating TLS that should not be.
Every carrier still sees metadata (who talks to whom, sizes, timing). Where that is
itself sensitive, use direct on infrastructure you control (SPEC §13).
Ingress role (own domain)
Section titled “Ingress role (own domain)”On the VPS run the ingress; on the node pair with a one-time token:
- ingress:
hdtp-gateway ingress serve --domain example.com(P5-05 wires the command; the library isinternal/ingress) and mint a token from its portal. - node: portal → Settings → Ingress — paste the token, pick the subdomain and mode:
- passthrough — the ingress routes on SNI and forwards raw TLS; your node’s own certificate is what callers see (direct mode).
- terminate — the ingress holds an ACME certificate, terminates the public TLS session, and re-originates a mutually-pinned mTLS leg to your node (edge mode for your node). Both keys are pinned at pairing; nothing else can deliver traffic.
- The node connects outbound (embedded frp client) — no inbound port at home.
Call budgets
Section titled “Call budgets”Every call counts against a budget of the account it is addressed to, sized by
how many contacts that account may hold (HDTP §12). The budgets are not the
node’s to decide: the limits sidecar, hdtp-limitd, holds their numbers and
their counters and decides each call with hdtp-identity’s hdtp-limits crate,
the decision the hosted cloud makes (SPEC §5.7). The node asks it over a unix
socket, limits_socket (HDTP_LIMITS_SOCKET, default <data_dir>/limits.sock),
on one connection it keeps open; nothing but the charge — the account, the
caller’s root and tier, its address, the contact cap — crosses it.
Its numbers are its configuration file, the rules document of hdtp-identity’s
contract, shipped as deploy/limitd/limits.json (in the image at
/etc/hdtp-limitd/limits.json):
| Member | Budget | Keyed by |
|---|---|---|
contact_calls_per_second, contact_burst |
a contact | account, contact root |
identity_capacity_per_second |
every contact together: limit.contacts × the contact rate, burst one second of it, at most this |
account |
guest_calls_per_hour |
a guest (a proven root that is not a contact) | account, root, address |
guest_source_calls_per_hour |
an address alone (nothing proven) | account, address |
stranger_calls_out_per_hour |
calls OUT to strangers (request_contact, redeem_invite, the answers to a request) |
account |
integration_calls_per_hour |
one integration’s tools, per contact | account, integration, contact |
pending_in_cap |
requests waiting on the owner that strangers may write | account |
guest_total_calls_per_hour |
every caller the open does not prove a contact, together (below) | account |
guest_total_calls_per_hour is checked before the envelope is opened (the owner’s decision of
2026-09-29). Nothing unopened says who sent a sealed call, so the node asks the sidecar first,
before it reads a key: while the account’s total holds a call, every sealed call goes on to be
opened; once it holds none, only a call from a known address does — one that carried an active or
pending_out contact’s call to the account in the last hour, which the sidecar remembers — and
everything else is refused rate_limited, in the clear, with the total’s retry_after. Asking
spends nothing. The total is spent after the open by every call that is opened and does not prove
an active or pending_out contact: a guest, a blocked or superseded root, a root whose only row is
the request it left, a small form naming a leaf nobody pinned, and an envelope_invalid or
certificate_renewed the open found. A proven contact never spends it; a plaintext call opens
nothing and never spends it.
The known cost: during a flood, a contact calling from an address it has not used in the last hour
is refused before the open, as a stranger is, with retry_after, until the total refills. The known
addresses live in the sidecar’s memory beside the counters, an hour each and a bounded number an
account (the oldest going first: KNOWN_SOURCE_TTL_MS and KNOWN_SOURCES_CAP,
cmd/hdtp-limitd/src/lib.rs), so a sidecar restart forgets them too.
A known address is an address: every caller arriving from it shares its standing. Behind a carrier
that delivers every caller from one address of its own — frp, ngrok and tailscale from the
node’s host, a terminate-mode ingress from its own, none of which names the client’s address — one
contact’s call makes that one address known, and the check before the open lets every stranger
through to be opened (and refused after it). The check does its work where each caller arrives with
an address of its own: behind Envoy (below), or the cloudflare adapter. (TestAStrangerFloodDrainsTheTotalThenIsRefusedBeforeTheOpenAndAKnownContactGetsThrough
shows a stranger at the contact’s known address opened and refused.)
Calls out to a contact spend that contact’s rate and the account’s aggregate, in
buckets of their own. A refusal is a rate_limited tool error carrying
retry_after — the whole seconds until the bucket holds a call again — and an
audited rate_limited row naming the bucket, not a dropped connection, so the
caller’s agent can read it and back off. A call out that is refused never leaves
the node. A request past pending_in_cap is answered unavailable: no wait
empties a list only the owner can.
The shipped identity_capacity_per_second is measured, and it is below what 500 contacts ask
for. One node on an Apple M2 Max served sealed send_message from 500 contacts at up to 280
calls a second in every run, and broke between 300 and 450 a second from run to run
(TestMeasureAccountCapacity, internal/node; the method and the numbers are in its file), and
the shipped figure is the lowest knee with a margin. The figure is the node’s, and every identity
on the node shares it.
Changing a number is editing the file and restarting the sidecar, which reads it once and
refuses one it cannot enforce (a rate of zero, a burst under one call, a contact bucket that takes
longer than the hour an idle row is kept to refill), saying which member and why. The node needs no
restart: get_card asks the sidecar for the numbers it advertises on every call.
When the sidecar is down, the node refuses. Every sealed call, call out, request and
integration call is answered unavailable until it answers again, which the node notices by
itself. /healthz answers 503 and names the socket, so the container’s healthcheck fails;
hdtp-gateway doctor prints FAIL limits with the reason, and the serve banner says limits: NOT ANSWERING. Start the sidecar (hdtp-limitd -config <file>; the compose file runs it) and
the node serves again with no restart.
The counters live in the sidecar’s memory, one set for every node process on the host. A restart of the sidecar refills every bucket, which is the trade-off for not writing to a database on every call — the budget is there to blunt abuse, not to meter usage, and an attacker who can restart your sidecar has already won. Memory is bounded: a bucket idle for an hour is full whatever it budgets, so it is dropped (checked once a minute), and a caller cycling addresses or fingerprints cannot grow the table past who called in the last hour.
The one budget setting that is the node’s is how many contacts each identity may hold, which sizes the aggregate:
| Setting | Environment | Default | Meaning |
|---|---|---|---|
limit.contacts |
HDTP_LIMIT_CONTACTS |
500 | contacts each identity may hold — active contacts plus the requests it sent — and the size of its call budget |
It is an ordinary knob: the portal’s Settings page edits it under Security, the
config file carries it as limit_contacts, and the environment pins it above
both (SPEC §12.2), in which case the page shows it locked and says why. It is
read per use, so a change takes effect with no restart. Empty, zero or
unparseable restores 500 rather than removing the cap. The cap is enforced
wherever a contact is added — approving a request, unblocking a contact,
accepting somebody’s invite, sending a request, a peer redeeming an auto-accept
invite, approving a contact at a new address, and an import — and nothing already
held is removed when it is lowered.
Connection bounds
Section titled “Connection bounds”The public listener holds at most 1,024 connections open (SPEC §5.7); one more is closed before
its TLS handshake, and the audit trail gets one listener_full row a minute while that continues,
with how many there were. Requests have 10 s for their headers, 60 s in all, 75 s for the answer,
and an idle connection is kept 120 s. Rate limits per address or for the whole node are not the
node’s: put them where the traffic arrives — the edge, or a proxy in front of the node, such as
the one below.
Behind Envoy
Section titled “Behind Envoy”deploy/envoy/ is the node behind a proxy of its own, the first of the two layers of its rate
limits (the second is the sidecar above): docker compose -f deploy/envoy/compose.yaml up -d runs
Envoy, the node and the limits sidecar, and only Envoy publishes a port. Before it: make identity-proxy limitd-vendor (the image build), a certificate for the node’s public name at
deploy/envoy/tls/cert.pem and key.pem, and HDTP_PUBLIC_URL, that name, in the environment.
What Envoy does (deploy/envoy/envoy.yaml, whose numbers are its own and nowhere else):
- terminates the caller’s TLS with that certificate, which is what a caller now sees — WebPKI for the node’s name, as behind a terminating edge — and asks for the caller’s certificate chain, accepting any, since there is no authority above the person;
- limits each source address per path — the MCP endpoints, the invite landing, everything else, each with a bucket of that address’s own — and answers 429 past it, before the node sees the request; and holds the listener’s connection cap and timeouts;
- forwards the caller’s chain in
X-Forwarded-Client-Certand the address its socket saw inX-HDTP-Client-Address, replacing whatever the caller sent in either.
The node reads those two headers only from Envoy’s address, proxy_address
(HDTP_PROXY_ADDRESS, an IP; the compose file gives Envoy a fixed one on its network and the node
that one). From anywhere else they are a caller’s own claim and prove nothing, and with no
proxy_address they are never read. The node still opens every envelope: Envoy sees the MCP
requests, never what a sealed one carries. internal/integrationtest/envoy_test.go holds the two
files to what the node relies on, on every commit, and the harness’s S23 runs them with the image
and floods them, with a control.
The store, when it is large
Section titled “The store, when it is large”Nothing here needs setting. It is written down so that what the node does to its database is not a surprise, and so that the one knob that exists is findable.
- SQLite (the default) is opened in WAL mode with
synchronous=FULL, write transactions that take their lock atBEGIN, and a pool of four connections. Each was measured against a store of a million messages and the reasons are beside the code (internal/core/store/sqlite.go).synchronous=FULLis the one that costs speed on purpose: a commit here is a message a peer was told was delivered, and it is not given up to a power cut for a faster write. - Postgres (
store_engine: postgres) uses pgx’s pool with its default size, the larger of four and the number of CPUs. The DSN is where it changes: appendpool_max_conns=16(andpool_min_conns,pool_max_conn_lifetime) topostgres_dsn. - Every hour the node removes what has outlived its own window, whatever retention an account has set: idempotency records of sealed calls, which a node would otherwise keep one of for every call it ever took, owner sessions nobody came back to, and change-log rows older than a week. With a retention window set it also removes the messages, threads and media past it.
- Every statement the store can run is checked at build time to have an index on any table
that grows (
TestEveryQueryHasAPlan, both engines). To see the numbers on your own hardware:
HDTP_SCALE_DB=/tmp/hdtp-scale.db go test ./internal/core/store/ -run '^$' -bench '^BenchmarkScale' -benchtime 20xThe first run seeds the file — 10,000 contacts, a million messages, a million audit rows — and
takes about a minute. make scale runs these benchmarks over a scratch file, and times an import
of 40,000 threads against one of 10,000 (HDTP_EXPORT_SCALE), three rounds interleaved, failing
if the best of three takes more than 6 times as long for 4 times the threads, in reading or in
writing (linear is 4); the pre-push hook runs make scale last, after every other step.
More than one process
Section titled “More than one process”Several serve processes can share one store, each with its own internal_bind and public_bind
behind whatever balances between them (SPEC §11.1):
- SQLite: on one host, all with the same
data_dir. Not across hosts, and not on a network filesystem: SQLite’s locking does not hold there. One limit of this profile:Scrub, the checkpoint that clears the write-ahead log after a leaf key is destroyed, cannot finish while another process is reading the file. It says so — a warning on an install or a signing request, an error on a retirement or a leave — and the next scrub that finishes clears the log; until then the destroyed key’s bytes may remain in it. Postgres has no scrub at all (SPEC §3.9). - Postgres: on any hosts, each with a
data_dirof its own and the samepostgres_dsn. Give every host the same master key inHDTP_MASTER_KEY(a host that generates akeyring.keyof its own cannot open a key another sealed), and pointblob_dir(HDTP_BLOB_DIR) at storage every host mounts, or media stored through one host is missing on the others.
The first serve on an idle data dir migrates; the others check that the schema is the one they
were built for and refuse to start if it is not. To migrate, stop every process on the data dir
(on Postgres, every process) and start the new binary. migrate, export, import and the
audit commands refuse to run while any serve holds the data dir. One process serves the admin
socket and serve prints admin: ... is served by another hdtp-gateway process on the others;
the setup URL a first run prints works on the portal of the process that printed it. The outbound
retries and the hourly retention pass run on one process at a time: the one holding the work’s
lease in the store, renewed every 10 s; a holder that stops lets it go at once, and one that
crashes is replaced within 30 s.
The call budgets are shared by every process on a host: they are the limits sidecar’s, and every
process names the same limits_socket. A unix socket does not cross hosts, so on Postgres each
host runs a sidecar of its own, and each host grants the whole budget.
What each process still keeps to itself, and so what is not yet shared between them:
- integrations: every process connects each one itself, so a stdio integration runs a child per process. An OAuth token is one for them all: the store holds it, and an expired one is refreshed by the one process holding that integration’s refresh lease while the others wait for it.
Export and import
Section titled “Export and import”hdtp-gateway export --config config.json --slug alina --out alina.ziphdtp-gateway import alina.zip --config config.json --slug alina # review: writes nothinghdtp-gateway import alina.zip --config config.json --slug alina --yes # writes itBoth are offline commands: stop the node first (they refuse while hdtp.lock is held).
An export is one identity’s contacts, chats and files, and nothing else (SPEC §3.10, HDTP
§9.2): one unencrypted zip of manifest.json, contacts.csv, threads.csv, messages.jsonl and
media/<sha256>, the same format the cloud and the hdtp CLI read and write. export says, before
it writes, that the file is not encrypted: anyone who gets it can read the contact list and every
conversation and file, though it holds no key and cannot be used to speak as anyone. The file is
created 0600 and never replaces an existing one. A stranger’s request that was never accepted,
and its conversation, stay behind.
What is not in it, because it is the host’s and not the person’s: every key, saved settings, integration credentials, owners, passkeys, sessions, tokens, invites, the audit chain, the ledger of leaves, and a contact’s preset, trust flag and card. There is no flag that adds any of them.
import checks the WHOLE file before it writes anything, and refuses it whole at the first fault,
naming the member, row or line. Without --yes it shows what it would write and stops. Into a slug
that is not here it creates the identity keyless — its root and nothing more; into the slug the
file belongs to it merges, keeping every pin this host already holds. The rows go in under one
transaction, the files after it. An undelivered outbound message arrives as failed: delivering
it was the old host’s job, under the old host’s leaf.
Every import ends the same way: a request for a new leaf, which the import mints itself — move
for an identity that arrived, renew for one that was here — and prints with how to complete it
(the portal’s /identity/<slug>/wallet for the web wallet, or the request itself for the CLI
wallet). Installing the leaf sends every imported contact this host’s handshake (account announce
reports it). With no public_url set there is no address to ask for, and it names account csr
instead.
This is not a backup of the node, and the node has none: what a host accumulates beyond contacts
and chats is rebuilt, not restored. After a lost machine: import, the setup wizard for a passkey,
reconnect integrations, one certificate per identity.
Recovery
Section titled “Recovery”| Lost | Consequence | Do |
|---|---|---|
| the node, with an export per identity | identities are not served until re-certified; settings, integrations, passkeys and the audit history are not in an export | import each, serve, the setup wizard; then one account csr / install-leaf per identity. Contacts keep their pins |
keyring.key only |
sealed leaf keys, saved settings and integration credentials are unreadable. No identity is lost — the root is in the wallet. If nothing on the node opens under the master key it was given, serve refuses to start and says why; if anything does, it starts and prints NOT SERVED for each account that does not |
put the key back if it was kept anywhere, and nothing is lost. If it is gone for good: export -slug each identity, then import each into a fresh data directory — an export never needed the master key, because it never held anything sealed under it. For a single NOT SERVED account on a running node, a renewal alone does it: the install retires the key it cannot open and says so |
| a leaf simply ran out (nobody renewed it) | the account stops being served within the hour and its key is destroyed; contacts keep their pins, and doctor warns before it happens |
account csr --slug me -purpose renew, the wallet signs, account install-leaf |
| a card or certificate on file that the identity core no longer reads (its reading got stricter with a release of the identity library: the identity core 0.4.2 refuses a character outside base64url in a card’s certificate that 0.4.1 skipped, and what intake accepted then is still on file) | the core refuses it where it reads it, and the node says so only then: a contact whose card does not read cannot be written to (node.PeerOf reads the card first and refuses — a refresh of that contact included), and a call from a contact whose pinned leaf does not parse, in either form, is this node’s unreadable state (identity_state_unreadable on the trail, envelope_invalid to the caller). The serve banner’s store: line and hdtp-gateway check store read every such field first and name each that does not read — table, row, field, reason — and change nothing; check store exits 1 on one, and runs beside a serving node |
for a contact’s card: this node cannot ask for one; the contact’s own next update_contact writes one that reads, or remove the contact and re-add it from a new card of theirs. For a pin’s leaf: remove the contact and re-add it, since nothing it sends can be decided against that pin. Run hdtp-gateway check store before and after a binary that bumps the identity library |
a contact in a state neither a pin nor a request has (the schema admits only active, pending_in, pending_out and blocked — both engines — so such a row is a hand-edited store’s or a later binary’s) |
the node hands the identity core no pin for it, where the core would refuse the state as unreadable: a call from its holder is decided as a stranger’s — the small form chain_required, the chain form a guest’s — and no effect reaches the row. The serve banner’s store: line counts such rows and a NO PIN line names each by account, root and status; hdtp-gateway check store prints the same and exits 1 |
remove the contact (remove_contact ends a relationship in any state) and re-add it from a new card of theirs |
| one account’s leaf key (compromised) | the thief speaks as that host until the leaf expires or is outranked | hdtp-gateway account csr --slug me -purpose renew, have the wallet sign it, account install-leaf: the newer leaf outranks the stolen one with every contact it reaches (HDTP §14.3) |
| nothing: the person moved an identity to another host | this node goes on serving it, with its key, until told | hdtp-gateway account leave -slug me -yes once the new host has told the contacts (without -yes it shows what it would erase): every record and leaf key of the identity erased, its address reserved until its last leaf expires (SPEC.md §3.11) |
| the wallet’s root | the identity itself; this node cannot help | the wallet’s own recovery, if it has one (HDTP §9, §14.5) |
| the audit chain shows a break | someone altered history | audit verify names the first bad row; treat the store as untrusted from there |
No telemetry leaves the node, ever; the audit chain is yours alone.