Federation Pairing
Pair two Syncropel instances so they can sync threads, exchange records, and converge on shared state. Covers when to pair, the pair lifecycle, debug recipes for the common failure modes, how to read a pair's health, and capacity planning.
Audience
This page is for the operator running two (or more) Syncropel instances that need to talk to each other. Single-instance operators don't need pairs — pairs exist specifically to bridge separate instances under separate identities, with consent and credentials negotiated explicitly.
If you only run spl serve on one machine, skip this page. If you run two and they need to share threads, read on.
When to pair
A pair is the right primitive when:
- You have two laptops and want one set of threads visible from both, with edits flowing in either direction.
- You're bridging a hosted instance and a local one — say,
<label>.syncropel.appfor the agent stack and a laptop instance for capture-and-emit. - You're sharing a workspace across collaborators — each collaborator has their own instance, their own identity, their own trust ledger; pairs let consented threads cross the boundary.
- You're running a fleet of mostly-identical instances behind the same operator and want intra-fleet sync without re-running the discovery dance every time.
A pair is not the right primitive for:
- Multi-tenancy in one instance (use namespaces — see
spl namespace --help). - Backup or migration (use
spl thread snapshot/restore). - Read-only embedding into an external app (use the SDK).
The handshake, conceptually
Pair establishment is a one-shot CLI handshake that exchanges DIDs, signs a manifest, and mints federation-scoped bearer tokens reciprocally. The stored state is a record on th_federation_pairs; the wire is the existing federation transport (HTTP + signed manifest at /.well-known/syncropel).
The diagram above is the canonical handshake — discovery, signed manifest, reciprocal token mint, persistent records on both sides.
Prerequisites
Before either instance runs spl federation pair:
- Both instances have a cryptographic identity. Run
spl identity generateon each side ifspl identity showreturns nothing. Without an identity, the federation manifest is not auto-published and the handshake has nothing to sign. - Both instances publish a federation manifest.
curl https://<peer>/.well-known/syncropelshould return JSON withdid,pair_endpoint, and a signature. The manifest auto-publishes once the instance has an identity. - Each side has a token with the
adminscope (you'll be writing pair records and minting federation-scoped SAs). - The pair_endpoint is reachable. Port-forward, tunnel, or DNS — whatever it takes for
curlfrom one instance's host to the other'spair_endpoint.
2. Pair lifecycle
Create a pair
spl federation pair https://bob.example/This is the headline command — discover, verify, handshake, persist, all in one. Output:
Pairing with did:sync:instance:bob (https://bob.example/) ...
✓ manifest fetched + signature verified
✓ POST /v1/federation/pair → 200 OK
✓ peer token persisted (federation:sync, federation:subscribe)
✓ recorded on th_federation_pairs (pair_id: pair_a1b2c3d4)
✓ initial sync cursor advanced (12 records)
Pair active.If the responder runs in manual-approval mode, the initiator sees 202 Accepted and the pair sits in establishing state until Bob's operator approves it via spl decisions approve <pair_id>. Approval triggers token mint on Bob's side and the initiator's next refresh advances the state to active.
Useful flags on the initiator:
| Flag | Meaning |
|---|---|
--map local_ns:peer_ns | Propose a cross-namespace mapping. Default is intra-namespace only. Repeatable. |
--no-strict | Skip strict manifest expiry validation. Default is strict; use only when peer's clock is off. |
--auto-generate-identity | Create a local identity if none exists. Default refuses with a hint to run spl identity generate first. |
List pairs
spl federation listReturns a table of pair_id, peer DID, state, mode, last sync timestamp. States:
| State | Meaning |
|---|---|
establishing | Handshake in flight; waiting on peer approval or response. |
active | Pair is live; sync is happening per the configured mode. |
paused | Operator paused sync but credentials remain valid. |
degraded | Sync errors are recurring (manifest changed, peer offline, auth refused). |
revoked | Terminal; tokens invalidated; pair record retained for audit. |
Show pair detail
spl federation show <pair_id>Reveals the peer manifest snapshot, last refresh timestamp, last sync cursor position, consent grants attached to this pair, and token IDs only (plaintext bearers are never displayed once issued). If you need a token's plaintext to embed in a different tool, mint a new token with spl token create --scopes federation:sync,federation:subscribe and revoke the original.
Pause and resume
spl federation pause <pair_id>
spl federation resume <pair_id>Pause is the right move when you're upgrading the peer or making sensitive config changes — sync stops, credentials remain valid, no records are lost. Resume picks up from the saved cursor.
Refresh (re-fetch peer manifest)
spl federation refresh <pair_id>The instance refreshes manifests on a schedule, but refresh forces it now. Use this after the peer rotated identity or changed pair_endpoint. The pair transitions through peer_manifest_changed and lands back in active if the new manifest verifies.
Revoke (terminal)
spl federation revoke <pair_id>Best-effort calls the peer's notify-revoke endpoint, then emits pair.revoke.v1 locally. Tokens are invalidated immediately on the local side; the peer should see the notification and invalidate symmetrically. If the peer is offline, the local revocation still takes effect; the peer's tokens for you remain valid until their TTL expires (90 days default).
Consent grants (cross-namespace sharing)
By default a pair only carries records within matching namespace pairs. To allow cross-namespace sharing, attach a consent grant:
spl federation grant did:sync:instance:bob \
--namespace music \
--maps music,projects/music \
--hash-level L1This composes with the L0-sharing consent rules — L0 sharing across namespaces still requires explicit two-sided consent records on th_consent. The default hash level for grants is L1 (structural); only raise to L0 (exact) when you genuinely intend to share content-identical records.
Sync mode
Pairs default to polling at 5-minute intervals. For active threads (a shared task list, a shared dispatch queue) bump up:
spl federation set-mode <pair_id> continuous
spl federation set-poll-interval <pair_id> 30s # only for polling modecontinuous opens an SSE subscription to the peer; latency drops from minutes to seconds, at the cost of an open TCP connection. on-demand is the third option — sync only when explicitly requested via spl sync <pair_id>. Use it for archival peers that don't need real-time updates.
3. Debugging pair issues
The four failure modes that account for most pair problems, in roughly the order they happen:
Manifest fetch fails
curl -v https://bob.example/.well-known/syncropelThe manifest must return 200 with a JSON body that includes did and a signature. Common causes for failure:
- Bob's instance hasn't generated an identity. The manifest auto-publishes only when an identity exists. Fix on Bob's side:
spl identity generate. pair_endpointis reachable from the public internet but the/.well-known/syncropelroute isn't. Reverse proxies sometimes strip dot-prefixed paths. Test from Bob's host:curl http://localhost:9100/.well-known/syncropelshould return identical JSON.- TLS chain doesn't validate.
curl --insecuresucceeds butspl federation pairrefuses. Fix the TLS chain —spl federation pairdoes not accept self-signed certs.
Manifest signature verification fails
spl federation pair returns: manifest signature does not verify against advertised DID.
Causes:
- The peer rotated identity but the manifest was cached. Force-refresh:
curl -H 'cache-control: no-cache' https://bob.example/.well-known/syncropeland compare DIDs. - The DID document at the published endpoint doesn't match the DID inside the manifest. Bob must run
spl identity generateagain or bring the DID document in line with the instance's identity.
POST /v1/federation/pair returns 401
peer rejected handshake: 401 UnauthorizedThe handshake is authenticated by the manifest signature, not by a bearer — but the responder's pair handler still runs the standard auth middleware. Check on Bob's side: auth.required = true is fine, but the /v1/federation/pair route must be reachable without a pre-existing bearer (the handshake is the bearer-mint event). In the canonical config this Just Works; if you've layered a custom permission rule on top, audit it:
spl config list-permission-rules | grep federationA permission rule that requires record_write for the federation:pair route will lock out pairing. Either widen the rule or delete it.
Sync cursor doesn't advance
spl federation show <pair_id>
# last_sync_cursor: clock=42, lagging by 138 recordsCauses, in order of likelihood:
auth.required = trueon the peer + the federation token's scopes don't cover what's being read. Checkspl federation show <pair_id>for the peer-issued token's scopes; you need at minimumfederation:syncandfederation:subscribe. Mint a new token with the right scopes if needed.- The peer paused the pair and didn't tell you.
spl federation refresh <pair_id>will resync the manifest and surface state changes. - Network drop between cursor advances. The cursor is monotonic — re-running
spl syncis safe, idempotent, and resumes from the last good position. - Same-clock cursor skip (rare, historical bug). Older peers'
(clock, id_prefix)cursor implementation could permanently skip a same-clock record arriving out of order. Symptom: lag never decreases even after restart. Workaround: upgrade the peer, then pause + force-resync viaspl federation pause+resume.
Whole-pair recovery procedure
When a pair goes wrong in a way the per-mode debug doesn't fix:
# 1. Pause to stop further drift.
spl federation pause <pair_id>
# 2. Snapshot both sides for forensics.
spl thread snapshot th_federation_pairs > /tmp/local-pairs.snap.jsonl
ssh bob 'spl thread snapshot th_federation_pairs > /tmp/bob-pairs.snap.jsonl'
# 3. Compare the records — usually one side has a transition the other doesn't.
spl debug thread-diff <local-pair-record-id> <bob-pair-record-id>
# 4. If the divergence is irreparable, revoke + re-pair.
spl federation revoke <pair_id>
spl federation pair https://bob.example/Re-pairing is cheap. The historical cursor on the new pair starts at the current peer clock; you don't get the missed records back unless you also restore from snapshots, but you get a clean slate.
4. Reading a pair's health
There is no bundled load generator for pairs. What the daemon does give you is a live view of every sync pair it is running: the pair's lifecycle state, how many records it has pulled, how long each poll takes, how far behind it is, and whether its delivery latency has drifted. Three surfaces expose it, from coarsest to finest.
The state column
spl federation list
spl federation show <peer_did>list folds th_federation_pairs and prints one row per pair with its state (not_paired, establishing, active, paused, degraded, revoked). show adds the last clock the pair reached, when it last changed state, and the reason for the last transition. This is the record-level view: it tells you what the pair is, not how fast it is moving.
The sync health doors
The sync loop keeps in-memory metrics per pair over a rolling one-hour window. They are not records: they reset when the daemon restarts, and they are meant to be scraped, not archived. Both doors need a bearer carrying the federation:manage scope.
# Every pair the daemon is syncing, in one report.
curl -s -H "Authorization: Bearer $SPL_TOKEN" http://localhost:9100/v1/sync/health
# One pair in detail.
curl -s -H "Authorization: Bearer $SPL_TOKEN" http://localhost:9100/v1/sync/pairs/<pair_id>/statsGET /v1/sync/health returns:
| Field | Meaning |
|---|---|
pairs_total, pairs_by_state | How many sync loops are registered, bucketed by loop state: starting, running, paused, failing, stopped. |
records_pulled_last_hour, records_pulled_last_minute | Throughput across every pair. |
mean_delivery_latency_ms, p50_delivery_latency_ms, p95_delivery_latency_ms | How long a record takes to arrive after the peer wrote it, across every pair in the window. |
slowest_pairs | The five pairs furthest behind, each with lag_secs (seconds since its last successful poll) and its last_error. |
most_errored_pairs | The five pairs with the most failed polls, each with errors_last_hour. |
drift_alerts | Pairs whose delivery latency has moved off its own baseline (a CUSUM over the latency series, which fires on a sustained slowdown of roughly twice the baseline). Each alert names the metric, the baseline_ms, and the current_ms. |
GET /v1/sync/pairs/{pair_id}/stats returns the same measurements for one pair, plus its state, its cursor, records_pulled_total, poll_intervals (mean_ms, last_ms, configured_ms), delivery_latency (mean_ms, p50_ms, p95_ms, last_ms), errors_last_hour, retries_current, last_successful_poll_at, and lag_secs.
The bearer-gated GET /v1/health carries a one-glance summary of the same data under federation: pairs_total, pairs_healthy (loops in starting, running, or paused), and pairs_failing.
The doctor
spl doctor adds one pairing-specific row, federation bind host: when the daemon has sync pairs configured but is bound to a loopback address, it warns that remote peers cannot reach this instance and names the --host flag to fix it. (The doctor's replication row is about the repository write-back lease, not about pairs.)
What healthy looks like
- Every pair you expect to be live reads
activeinspl federation listandrunninginpairs_by_state. lag_secsstays close to the configured poll interval. A lag that keeps growing means polls are failing, andlast_erroron that pair says why.errors_last_houris 0. A pair that is erroring is either infailingalready or on its way there; the sync loop backs off and retries, so a single transient error is not a problem, a steady count is.drift_alertsis empty. An alert with a risingcurrent_msand no error is usually the network between the peers, not the pair.
If a pair looks wrong, the lifecycle section above is the recovery path: spl federation refresh to re-read the peer, pause and resume to restart the loop from its saved cursor, and revoke plus a fresh pair when the divergence is not repairable.
5. Capacity planning
Per pair, in continuous mode:
| Resource | Per pair (steady) | Per pair (4× spike) |
|---|---|---|
| Memory (responder side) | ~12 MB | ~24 MB |
| Memory (initiator side) | ~8 MB | ~16 MB |
| Bandwidth | 2-4 KB/s | 8-16 KB/s |
| Open TCP connections | 1 (SSE) | 1 (SSE) |
| File descriptors | 4-6 | 6-10 |
Per pair, in polling mode (5 min default):
| Resource | Per pair |
|---|---|
| Memory | ~3 MB |
| Bandwidth | < 100 B/s average (bursty around poll cycles) |
| Open TCP connections | 0 (only during poll) |
Practical limits:
- A single instance comfortably runs 50+ pairs in polling mode on a small instance (2 GB / 1 vCPU is plenty), assuming the workload is dominated by sync rather than user-facing work.
- For continuous mode, plan ~10-20 pairs before you start contending with adapter throughput. Each open SSE adds a per-tick overhead the reconciler has to absorb.
- Bandwidth is rarely the constraint. The constraint is record-ingest throughput on the receiving side, which is bounded by the adapter pipeline, not by the pair primitive.
To find your own ceiling, add pairs under your real workload and watch GET /v1/sync/health as you go. The pair primitive is healthy as long as p95_delivery_latency_ms stays flat, lag_secs on the slowest pair stays near its poll interval, and drift_alerts stays empty. The first of those to move is the one telling you where the ceiling is.
What's next
- Federation discovery internals (DNS, manifest, signature): Operate / Relay.
- Cross-namespace consent:
spl federation grant --help. - Whole-instance backup and recovery: Instance Lifecycle.
Actor portability
Export an actor's identity + work trail from one Syncropel instance and import it into another. What migrates, what doesn't, and the consent implications.
Running an async-federation relay
Install, configure, monitor, and troubleshoot a Syncropel async-federation relay. Covers Docker and systemd deployment, Prometheus metrics, bearer-token auth for receivers, and the failure modes you'll hit in practice.