26 KiB
Reticulum Chat Sync Architecture
Purpose
This document describes the current Reticulum chat sync model, why it does not scale well, and the target architecture we want for long-term correctness and performance.
The goal is not only to reduce noisy logs. The goal is to make Reticulum chat work well when:
- a user belongs to many groups,
- each group may have many active participants,
- overlay peers appear, disappear, or change often,
- Reticulum wire payloads are small,
- attachments and resource downloads are active,
- background unread/mention summaries must stay fresh,
- the active open chat must stay responsive.
Current System
Reticulum chat currently uses a custom RCHAT protocol over the Reticulum overlay.
At a high level:
Reticulum overlay
-> RCHAT wire messages
-> ReticulumChatManager
-> SQLite event DB
-> UI summaries / active chat / resource downloads
Events are signed and stored locally. The UI reads local history from SQLite and receives live events from the manager.
Current Background Subscription Flow
When the app loads the user's joined groups, the renderer sends all local group memberships to the main process and subscribes to each group.
Group.tsx
-> setLocalGroupMemberships(all joined group IDs)
-> subscribeGroup(groupId) for each joined group
That means Reticulum chat subscribes to groups before the user opens a specific group chat. This is intentional because background summaries, unread counts, and mentions need to work from the group list.
Current Subscribe Behavior
Today, subscribeGroup() is too eager. A group subscription sends a full sync starter pack:
subscribeGroup(group)
-> sub
-> author_heads_req
-> sync_req after latest local event
-> sync_req before earliest local event
The peer replies with:
author_heads
sync_hints
sync_hints
sync_hints
...
If any hinted event is missing locally, we request the full event:
event_hint / sync_hints
-> event_req for missing event
-> event_offer / resource transfer for full event payload
Current Active Chat Flow
When the user opens a group/channel:
useReticulumGroupChat(group, channel)
-> subscribeGroup(group)
-> subscribeChannel(group, channel)
-> getHistory(group, channel)
The active chat gets local DB history immediately. But the network sync path still overlaps with the same group subscription machinery used by background groups.
Current Resource Flow
Attachments and large event payloads use the resource transfer path:
resource_find (rf)
-> resource_have from a verified provider
-> direct linked byte-range resource transfer
-> local resource store
This is conceptually separate from chat sync. RCHAT overlay messages only discover providers and route provider hints; the byte-range request is sent on the dedicated resource link after a provider is known.
Problems With The Current Model
1. Subscription Means Too Much
sub currently means:
I am interested in this group.
Also send author heads.
Also send recent history.
Also send older history.
That is too much work for a background subscription.
2. Background Sync Can Behave Like Active Chat Sync
The app may be joined to many groups. Background group subscriptions are needed for summaries and mentions, but they should not trigger full history sync for every group.
Current behavior can scale like:
joined groups * overlay peers * subscription replay * sync windows
That creates unnecessary sub, author_heads_req, sync_req, and sync_hints traffic.
3. Peer Churn Replays Too Much
When overlay peers change or the bridge recovers, subscriptions can be reannounced. Today that can replay full sync intent, not just liveness.
peer changed
-> reannounce subscriptions
-> sub + author_heads_req + sync_req + sync_req
Peer churn should not cause history sync storms.
4. Sync Happens Before We Prove Anything Is Missing
The current system asks for history windows first, then dedupes after receiving hints.
ask for history
-> receive many hints
-> discover most are already known
The better model is to exchange state first, then request exact missing data.
5. It Does Not Scale To Large Groups Cleanly
A group can have many users. The sync system must never depend on:
- full group membership lists,
- all authors in a group,
- all events in a group,
- all channels in one packet.
Sync must scale with actual missing event activity, not group size or network size.
Target Architecture
The target architecture is:
Local-first signed event-log replication
+ compact checkpoints/digests
+ cursor-based feed sync
+ exact missing-range repair
+ adaptive peer/source planner
+ strict priority scheduling
+ isolated resource transfer
In simpler terms:
Reticulum = transport
RCHAT = replication protocol
SQLite = local source of truth
UI = local projection
resources = large object transfer
RCHAT should sync signed event logs, not chat screens.
Core Concept
Every chat event is immutable and signed.
eventId
groupId
channelId
authorAddress
authorSeq
timestamp
eventType
payloadHash
signature
Each author has an append-only event sequence within a group:
Alice in group 716:
seq 1
seq 2
seq 3
Bob in group 716:
seq 1
seq 2
Nodes do not sync users. They sync events.
There are two sync tools:
Normal sync:
group/channel feed cursors
"I have this timeline up to event X"
Gap repair:
author sequence ranges
"I saw Bob seq 42 but only have Bob seq 39"
Author sequence state is important for integrity repair, but it must not be the main digest for large groups. A group can have many chatters, so the primary sync primitive must be bounded group/channel cursors.
Cursor And Continuity Semantics
Cursor semantics must be deterministic. Every peer must agree on what "after this event" and "before this event" means.
Feed Cursor
A feed cursor is a position in a group/channel timeline.
feedCursor
feedTimestamp
eventId
The canonical feed order is:
ORDER BY feed_timestamp ASC, event_id ASC
All feed queries, feed responses, hashes, and comparisons must use this exact ordering. A cursor means "the position at this feedTimestamp/eventId in the canonical feed order".
The feed cursor is cross-author. It should not include authorAddress or authorSeq, because those fields describe a single author's append-only sequence, not the mixed channel timeline.
Timestamp Trust
The signed event timestamp must not be trusted blindly for feed ordering.
The system should distinguish:
signedTimestamp:
timestamp inside the signed event
retained as event metadata
acceptedAt:
local receive/insert time
useful for abuse handling and local projections
feedTimestamp:
canonical timestamp used for feed ordering and cursors
Rules:
if signedTimestamp is within accepted past/future bounds:
feedTimestamp = signedTimestamp
if signedTimestamp is too far in the future or past:
reject or quarantine the event
if suspicious events are accepted for diagnostics:
feedTimestamp = acceptedAt
keep signedTimestamp as metadata only
The existing event timestamp validation rules should be tightened around this model. A malicious or broken client must not be able to force an event far into the past or future and distort feed cursors, unread summaries, or sync windows.
Author Cursor
An author cursor is used only for per-author integrity repair.
authorCursor
authorAddress
authorSeq
Author cursors answer a different question:
Do I have every event for this author up to this sequence?
Feed cursors and author cursors must not be treated as interchangeable.
Latest Does Not Mean Complete
A latest cursor only proves freshness:
I have seen an event at least as new as this cursor.
It does not prove continuity:
I have every event before this cursor.
The implementation must not assume that having the latest event means there are no holes behind it.
Continuity must be tracked separately:
background summaries:
latest cursor is enough for freshness
active visible chat:
visible window should be verified as continuous
author sequence gaps:
repaired with author range requests
Page Metadata
Every feed response should include page metadata:
pageStartCursor
pageEndCursor
hasMore
windowHash
windowHash is a deterministic hash of the ordered event IDs in the returned page. It lets peers detect divergence when they believe they are talking about the same cursor window but have different local contents.
Proposed Protocol Model
1. Hello
Peers announce the protocol version and runtime capabilities.
hello
version: 1
features:
digest
feed_req
range_req
event_batch
resource_v2
This is not a backward-compatibility layer. It is a sanity check so peers can reject unknown or incomplete protocol implementations instead of silently falling back to noisy behavior.
2. Group Subscription
sub becomes liveness and interest only.
group_sub
groups: [716, 812]
mode: summary
It should not automatically mean "send me history".
3. State Digest
A digest says what local event state we already have.
group_digest
groupId: 716
latestCursor:
eventId: abc123
feedTimestamp: 1782340000
channels:
general:
latestCursor:
eventId: abc123
feedTimestamp: 1782340000
oldestCursor:
eventId: def456
feedTimestamp: 1782330000
visibleWindowHash: ...
digestHash: ...
This is cursor based. A cursor is a bookmark that means "I have events up to this point" or "I have visible history back to this point".
The latest cursor is freshness state only. It is not a promise that all previous events are present locally.
For large groups, the digest remains bounded:
group_digest
groupId: 716
latestCursor:
eventId: abc123
feedTimestamp: 1782340000
channelCount: 100
channels: limited page only
moreChannels: true
The system never sends all authors for a huge group by default.
A small recentAuthorSample can exist as an optimization, but it must be optional, bounded, and never required for correctness.
recentAuthorSample:
max 10-50 authors
recently observed locally
used only to detect likely gaps faster
4. Feed Request
After comparing digests, request only missing feed pages.
feed_req
groupId: 716
channelId: general
after:
eventId: abc123
feedTimestamp: 1782340000
limit: 25
The receiver must answer according to the canonical feed order:
WHERE (feed_timestamp > cursor.feedTimestamp)
OR (feed_timestamp = cursor.feedTimestamp AND event_id > cursor.eventId)
ORDER BY feed_timestamp ASC, event_id ASC
LIMIT pageLimit
For older visible history:
feed_req
groupId: 716
channelId: general
before:
eventId: def456
feedTimestamp: 1782330000
limit: 25
Older history uses the inverse cursor comparison but still returns events in canonical ascending order.
5. Gap Repair Request
Author sequence ranges are used only when a received event or hint proves a gap for a known author.
range_req
groupId: 716
ranges:
Bob: 19-20
Carol: 8-12
This should not require enumerating every author in the group.
6. Event Batch
The peer returns compact event envelopes or full events only when they fit.
event_batch
groupId: 716
channelId: general
pageStartCursor:
eventId: event-a
feedTimestamp: 1782340100
pageEndCursor:
eventId: event-z
feedTimestamp: 1782340200
hasMore: true
windowHash: ...
events:
...
Oversized event payloads continue through the resource path.
The windowHash is computed from the ordered event IDs in events. It is not a security primitive; event signatures still provide authenticity. It is a cheap divergence detector.
7. Resource Transfer
Attachments and large event blobs remain separate:
resource_find (rf)
-> resource_have
-> linked byte-range resource transfer
Resource transfer should not be starved by background sync traffic. Overlay discovery stays bounded, and bulk/request bytes stay on the direct resource link.
Background Vs Active Sync
The target design splits sync into two tiers.
Background Group Sync
Used for:
- unread counts,
- mentions,
- latest summaries,
- notification readiness.
Behavior:
background group membership
-> group_sub summary
-> compact cursor digest pages
-> newest missing feed pages only when needed
Background sync must be cheap, low priority, and bounded.
It must not pull full history for every joined group.
Active Chat Sync
Used when the user opens a group/channel.
Behavior:
active channel opened
-> load local history immediately
-> active channel cursor digest
-> request missing visible feed pages
-> page older history only when user scrolls
Active chat gets higher priority than background group sync.
Priority Scheduler
RCHAT should have a small internal scheduler so not all work competes equally.
Suggested priorities:
P0 live active chat hints
P1 active chat missing event pulls
P2 channel metadata for open group
P3 resource transfer control
P4 background unread/mention sync
P5 periodic subscription refresh
Rules:
- Active chat should stay responsive.
- Background sync must yield during resource transfers.
- Periodic refresh must never flood the bridge.
- Duplicate control messages should be suppressed before sending.
Peer Planner
Overlay peers are sources, not state owners.
The planner should track:
peer health
groups peer appears useful for
last successful event/resource response
recent failed requests
known digest/cursor state
in-flight requests
Then decide:
what is missing?
which peer likely has it?
how urgent is it?
have we asked recently?
will it fit in one wire?
is a resource transfer active?
This prevents repeated requests to bad peers and avoids duplicating work across peers.
Wire Size Strategy
Every control message must be pageable and checked before sending.
Build page:
add next group/channel/range
check wireFitsReticulum()
if too large, flush current page
If even a single group digest is too large:
send minimal group checkpoint:
groupId
latest cursor
channelCount
digestHash
The peer can request detailed pages only if needed.
Cursor Edge Cases And Conflict Rules
The cursor rules must be explicit so different implementations do not diverge.
Same Feed Timestamp
If two events have the same feedTimestamp, the canonical tie-breaker is eventId.
ORDER BY feed_timestamp ASC, event_id ASC
This rule applies everywhere:
- local DB queries,
- digest generation,
- feed requests,
- event batches,
- window hashes,
- cursor comparisons.
Late Older Event
If a valid event arrives later with a feedTimestamp older than the current latest cursor:
insert it into the local event DB normally
do not move latest cursor backward
update or invalidate affected verified windows
trigger repair only if it affects an active visible window or summary range
The latest cursor is monotonic forward. It represents freshness, not complete history.
Overlapping Feed Page
If a feed page overlaps events we already have:
dedupe by eventId
validate any duplicate event if needed
verify windowHash if this is a known comparable window
continue from pageEndCursor
The sync cursor advances from the page metadata, not from the last newly inserted event. This prevents getting stuck on pages that are mostly duplicates.
Peer Sends Events Outside Requested Bounds
If a peer returns events outside the requested feed_req bounds:
ignore out-of-bounds events for that response
do not update peer cursor from the invalid page
record a peer protocol violation
penalize the peer if violations repeat
This prevents one bad peer from poisoning local cursor state.
Duplicate Event
If the same eventId arrives again:
if stored event hash/signature matches:
ignore as duplicate
if stored event hash/signature differs:
reject as conflict/malicious
record diagnostic evidence
The event ID is immutable. A different payload or signature for the same eventId must never replace local state.
Window Hash Mismatch
If two peers claim the same cursor window but provide different windowHash values:
do not mark the window verified
request the page from another peer if available
fall back to exact event ID/range repair if needed
penalize peers that repeatedly disagree with valid majority data
The window hash is a divergence detector. Event signatures remain the authority for individual event validity.
Initial Implementation Constants
The protocol must define concrete limits before implementation. Without shared constants, two correct-looking implementations can behave very differently under load.
These are initial defaults. They should be tuned from live metrics, but every implementation should start from the same bounded behavior.
MAX_FEED_PAGE_EVENTS = 25
MAX_DIGEST_GROUPS_PER_PAGE = 20
MAX_GROUPS_PER_SUB_PAGE = 50
MAX_DIGEST_CHANNELS_PER_GROUP = 16
MAX_RECENT_AUTHOR_SAMPLE = 20
MAX_IN_FLIGHT_PER_PEER = 4
MAX_IN_FLIGHT_ACTIVE = 8
MAX_IN_FLIGHT_BACKGROUND = 8
MAX_BACKGROUND_WORK_PER_TICK = 10
SYNC_TICK_MS = 250
DIGEST_DEDUPE_TTL_MS = 30_000
BACKGROUND_DIGEST_REFRESH_MS = 60_000
ACTIVE_DIGEST_REFRESH_MS = 10_000
TIMESTAMP_FUTURE_TOLERANCE_MS = 5 * 60_000
TIMESTAMP_PAST_TOLERANCE_MS = 30 * 24 * 60 * 60_000
PEER_VIOLATION_COOLDOWN_MS = 5 * 60_000
MAX_PEER_VIOLATIONS_BEFORE_COOLDOWN = 3
Meaning:
MAX_FEED_PAGE_EVENTSbounds event batch size.MAX_DIGEST_GROUPS_PER_PAGEandMAX_GROUPS_PER_SUB_PAGEprevent group membership replay from becoming one huge packet.MAX_DIGEST_CHANNELS_PER_GROUPprevents a large group with many channels from filling one digest.MAX_RECENT_AUTHOR_SAMPLEis only an optimization hint. Correctness must not depend on it.MAX_IN_FLIGHT_*prevents one peer or background sync from occupying the whole scheduler.MAX_BACKGROUND_WORK_PER_TICKandSYNC_TICK_MSmake background sync cooperative.- timestamp tolerances protect feed ordering from broken or malicious device clocks.
- peer violation cooldown prevents repeated malformed pages from wasting bandwidth.
All message builders must still call wireFitsReticulum(). These constants are upper bounds, not permission to exceed the real wire limit.
Data Model Additions
The immutable event table remains the source of truth.
Useful additions:
rchat_peer_group_state
peer_hash
group_id
latest_event_id
latest_feed_timestamp
digest_hash
updated_at
rchat_peer_channel_state
peer_hash
group_id
channel_id
latest_event_id
latest_feed_timestamp
oldest_event_id
oldest_feed_timestamp
visible_window_hash
updated_at
rchat_verified_windows
group_id
channel_id
start_event_id
start_feed_timestamp
end_event_id
end_feed_timestamp
window_hash
verified_at
rchat_missing_ranges
group_id
author_address
from_seq
to_seq
preferred_peer
attempts
next_attempt_at
rchat_sync_queue
priority
group_id
channel_id
peer_hash
operation
dedupe_key
next_attempt_at
Projection data remains local:
summaries
mentions
search index
channels/categories
read watermarks
These are derived from events, not synced as primary state.
Required Database Indexes
The sync design depends on fast cursor and repair queries. These indexes should exist before enabling the new protocol.
The existing immutable event table should keep event_id as its unique event identity. If the table remains named reticulum_chat_events, use indexes like:
CREATE INDEX IF NOT EXISTS idx_reticulum_chat_events_feed
ON reticulum_chat_events(group_id, channel_id, feed_timestamp, event_id);
CREATE INDEX IF NOT EXISTS idx_reticulum_chat_events_author_seq
ON reticulum_chat_events(group_id, author_address, author_seq);
The new sync state tables should have:
CREATE INDEX IF NOT EXISTS idx_rchat_sync_queue_ready
ON rchat_sync_queue(priority, next_attempt_at);
CREATE INDEX IF NOT EXISTS idx_rchat_peer_channel_state
ON rchat_peer_channel_state(peer_hash, group_id, channel_id);
CREATE INDEX IF NOT EXISTS idx_rchat_peer_group_state
ON rchat_peer_group_state(peer_hash, group_id);
CREATE INDEX IF NOT EXISTS idx_rchat_missing_ranges_ready
ON rchat_missing_ranges(group_id, author_address, next_attempt_at);
CREATE INDEX IF NOT EXISTS idx_rchat_verified_windows_lookup
ON rchat_verified_windows(group_id, channel_id, start_feed_timestamp, end_feed_timestamp);
If the event table is renamed during the rewrite, keep the same index shapes on the replacement table.
Current Vs Proposed
Current
subscribe / replay
-> send sub
-> send author_heads_req
-> send sync_req after
-> send sync_req before
-> receive many sync_hints
-> dedupe after traffic already happened
Proposed
subscribe / replay
-> send liveness sub
-> send compact cursor digest
-> compare state
-> request missing feed pages or exact repair ranges
-> transfer only missing events/resources
Implementation Plan
This design should replace the current eager sync model directly. Since this is not production protocol compatibility work, the implementation should remove the old behavior instead of carrying it indefinitely.
1. Define The New Wire Protocol
- Define implementation constants and database indexes first.
- Add
hello. - Add
group_sub. - Add
group_digest. - Add
feed_req. - Add
range_req. - Add
event_batch. - Use
resource_find/resource_havefor resource provider discovery. - Send byte-range resource requests only after opening the dedicated provider link.
- Every message builder must use
wireFitsReticulum()while paging.
2. Replace Subscription Semantics
subscribeGroup()becomes background-light.- Group subscription sends only
group_suband bounded summary digest work. - Subscription replay sends liveness/digest only.
- Remove automatic
author_heads_req,sync_req after, andsync_req beforefrom group subscription. subscribeChannel()becomes the active chat sync trigger.
3. Build Cursor Digest Sync
- Add group-level latest cursors.
- Add channel-level latest and oldest cursors.
- Build bounded digest pages by group/channel.
- Store peer digest/cursor state.
- Define and enforce canonical feed ordering everywhere.
- Use validated
feedTimestamp, not raw signed timestamps, for feed cursors and ordering. - Reject or quarantine events with timestamps outside accepted bounds.
- Track latest freshness separately from verified continuity.
- Suppress duplicate digest pages per peer/group/channel.
4. Build Feed Requests And Event Batches
- Compare peer cursors with local cursors.
- Request newer feed pages with
feed_req after. - Request older visible pages with
feed_req beforeonly for active chat or scrollback. - Return
event_batchpages that always fit the wire limit. - Include
pageStartCursor,pageEndCursor,hasMore, andwindowHashin every event batch. - Verify active visible windows independently from latest freshness.
- Use resource transfer for oversized event payloads.
5. Keep Author Sequence Repair
- Keep
authorSeqvalidation on every event. - When an inbound event proves a known author's sequence gap, enqueue
range_req. - Do not enumerate all authors for a group.
- Treat author heads as optional bounded repair/diagnostic data only.
6. Add Scheduler And Peer Planner
- Prioritize active chat and resource control.
- Track peer usefulness and failures.
- Add per-peer/per-group dedupe keys.
- Cap background work per tick.
- Keep background sync low priority.
- Make resource transfer control higher priority than background digest chatter.
7. Update Resource Transfer Integration
- Keep resource bytes on the existing resource transfer path.
- Keep event/resource manifests as signed chat events.
- Ensure resource progress/events are coalesced so transfer UI does not flood the bridge.
- Ensure background sync yields while large resource transfers are active.
8. Remove Legacy Eager Sync Paths
- Remove subscription-triggered
author_heads_req. - Remove subscription-triggered
sync_req after. - Remove subscription-triggered
sync_req before. - Remove automatic continuation behavior that can create unbounded
sync_hintsbursts. - Replace
sync_hintswith boundedevent_batchresponses.
9. Add Diagnostics And Tests
- Log digest pages sent/skipped.
- Log feed requests and event batches.
- Log duplicate suppression.
- Log scheduler queue depth by priority.
- Test many groups, large groups, peer churn, active chat, background summaries, and resource transfer under load.
Success Criteria
- Subscribing to many groups does not create a sync storm.
- Overlay peer churn causes small digest/liveness traffic only.
- Background summaries remain fresh.
- Active chat receives missing visible events quickly.
- Large resource downloads are not slowed by background sync chatter.
sync_hintsvolume drops sharply for already-synced peers.- Bridge event queue does not build multi-second backlogs during normal chat/resource use.
- Wire-size violations are impossible by construction.
- No protocol step depends on enumerating all group members or all group authors.
- Author heads are optional repair/optimization data, not the primary sync state.
Final Principle
The final system should behave like this:
Local DB is truth.
Peers exchange compact state.
Planner fetches missing feed pages and exact repair gaps.
UI renders local projections.
Large bytes use resource transfer.
That is the scalable foundation for Reticulum chat.