Skip to content

Introduce incremental offers - #2385

Open
nfrisby wants to merge 3 commits into
ch1bo/leios-protocol-message-indicesfrom
nfrisby/leios-incremental-offers
Open

nfrisby wants to merge 3 commits into
ch1bo/leios-protocol-message-indicesfrom
nfrisby/leios-incremental-offers

Conversation

@nfrisby

@nfrisby nfrisby commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

This PR is a step towards Issue input-output-hk/ouroboros-leios#1109, done now so that the first "official" version of the mini protocols will be able to express a feature we've anticipated ever since drafting the CIP but haven't yet prioritized fleshing out: incremental offers.

The key idea of incremental offers is to be able to relay a part of an EB closure before acquiring the entire EB closure. As we receive more and more of it, we're able to serve more and more of it.

There are plenty of details to work through yet, but these new MsgLeiosBlockTxsOffer semantics accommodates all of the ones I've yet anticipated.

These commits should be squash before merged into main, but they show a backtracking journey for the reviewer interested in how simpler schemes might look (eg nodes can only offer a longer and longer prefix).


These commits make the node-as-a-receiver correctly respect the new granular semantics of offers. However, the node-as-a-server still only ever serves entire ranges---it's a more involved change to teach the node to notice when it can offer certain ranges. In particular, it probably involves moving the ownership of emission of AcquiredEbTxs (and probably AcquiredEb while we're at it) into the LeiosFetch logic, out of the LeiosDb.

@nfrisby
nfrisby requested review from ch1bo and coot October 8, 2026 23:48
@nfrisby nfrisby linked an issue Oct 8, 2026 that may be closed by this pull request
@@ -1006,6 +1051,7 @@ demoLeiosFetchStaticEnv =
, maxRequestBytesSize = 500 * thousand
, maxJobBytesSize = 64 * thousandBase2

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we change maxJobBytesSize and/or maxRequestBytesSize to better match leiosClosureOfferMinLength _version = 200 * 1024 in some way?

--
-- TODO negotiate it in the handshake; until then every version gets the stub.
leiosClosureOfferMinLength :: NodeToNodeVersion -> BytesSize
leiosClosureOfferMinLength _version = 200 * 1024

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TODO confirm with @coot and @karknu what this number should be. One constraint from my end: it can't be much smaller than this, or else each peer can send "too many" offers per election.

In fact, maybe we want to make it bigger?

Or maybe it needs to scale as they send more offers? Like this is the bound on their first offer, and their later offers need to be even bigger? .... that might be too restrictive, preventing them from filling in gaps in what they've offered 🤔

@nfrisby nfrisby Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I discussed on Slack with the Network team. EG coot shared these tables

TCP cwnd Growth During Cold Start (Slow Start)
MSS=1460B, IW=10 segments (IW*MSS=14,600B)
SDU = 12,288B data + 8B overhead = 12,296B total

RTT # | cwnd (bytes) | cwnd (segments) | Cumulative sent | SDUs / Payload
------|--------------|------------------|------------------|------------------
  1   |   14,600 B   |        10        |    14,600 B      |   1 / 12,288 B
  2   |   29,200 B   |        20        |    43,800 B      |   3 / 36,864 B
  3   |   58,400 B   |        40        |   102,200 B      |   8 / 98,304 B
  4   |  116,800 B   |        80        |   219,000 B      |  17 / 208,896 B
  5   |  233,600 B   |       160        |   452,600 B      |  36 / 442,368 B

cwnd(RTT_n)        = IW*MSS * 2^(n-1)
Cumulative(RTT_n)  = IW*MSS * (2^n - 1)
SDUs fit           = floor(Cumulative / 12,296)
Payload            = SDUs fit * 12,288
┌─────┬──────────┬──────────────────┬────────────────┐
  │  M  │    c     │ fluid 1500 (750) │ DES 1500 (750) │
  ├─────┼──────────┼──────────────────┼────────────────┤
  │ 1   │ 12 MB    │ 233.7 (202.2)    │ 226.4 (202.6)  │
  ├─────┼──────────┼──────────────────┼────────────────┤
  │ 4   │ 3 MB     │ 95.2 (85.1)      │ 122.4 (110.6)  │
  ├─────┼──────────┼──────────────────┼────────────────┤
  │ 16  │ 750 kB   │ 54.1 (51.9)      │ 68.3 (64.2)    │
  ├─────┼──────────┼──────────────────┼────────────────┤
  │ 64  │ 187.5 kB │ 38.5 (41.3)      │ 51.3 (45.8)    │
  ├─────┼──────────┼──────────────────┼────────────────┤
  │ 184 │ 65 kB    │ 31.9 (37.2)      │ 49.3 (—)       │
  └─────┴──────────┴──────────────────┴────────────────┘

Which suggest that 200 * 1024 B is in the right ballpark.

Message (LeiosNotify point announcement vote) StBusy StIdle
MsgLeiosBlockTxsOffer ::
!point ->
-- the closure byte range [start, end) on offer; an end of maxBound runs

@nfrisby nfrisby Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My original vision was that this would be just an integer or a bitfield: each offset corresponds to a part of the EB closure that all nodes agree on.

That seems good as long as all nodes are using the same job size, and the job size and the objective upon partition scheme agree on which tx boundaries are delineate "parts". But I couldn't figure out how to accommodate changes in the partitioning scheme without having to give up alignment with jobs. The complicated part is: a single node's jobs would have to be compatible with multiple different partitioning schemes, since it might have some upstream/downstream peers on older NodeToNodeVersions. If the same EB couldn't be announced in more than one slot, then the slot could determine the scheme and the job scheme could change at the same time, but alas, announcements in different slots can name the same EbHash.

That's why I just gave up and used a very generic vocabulary in this message: an interval of bytes.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But the interval on offer can still not align with the job sizing, right?

I would have also expected us to start simple, but a complete scheme (server and client side figured out)

You say that it's easy to change network protocols (or at least feasible) without a hard-fork. So maybe the pragmatic way is to just slice it into pieces which happen to align with the job size we also pick?

@ch1bo
ch1bo added this pull request to stack #2364 October 9, 2026 09:03
@ch1bo
ch1bo force-pushed the ch1bo/leios-protocol-message-indices branch from 3cea3a5 to 9d9d05a Compare October 9, 2026 12:23
@ch1bo
ch1bo force-pushed the nfrisby/leios-incremental-offers branch from 3899a03 to e1a4311 Compare October 9, 2026 12:46
@nfrisby

nfrisby commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Sebastian squashed in order to rebase (totally fair). For (my) future reference, the unsquashed (exploratory) commits are on this tag: https://github.com/IntersectMBO/ouroboros-consensus/releases/tag/leios-incremental-offers-exploration

@nfrisby

nfrisby commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

I just pushed up one restriction that I realized overnight was missing: LeiosFetch should not create a "job" (ie one fetchable unit) that spans the incremental offers it expects to get from honest nodes on the same version.

After that, these are the loose-ends, as Claude summarized my explanations:

  • Incremental sending: is deferred for now.
  • One seam spacing per pool: it is accepted. Misalignment of seams among peers with different values of that parameter cost a (hopefully) minor inefficiency that fades as more nodes update.

@ch1bo ch1bo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Damn.. had these prepared and not sent off.

-- LeiosNotify client.
--
-- TODO never pruned: one entry per point ever offered, for the connection's
-- lifetime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

which the fetch logic prunes once consumed

TODO never pruned

What now?

-- start below it, which bounds how many a peer can send per point.
--
-- TODO a Leios protocol parameter that varies with the slot; static stub for
-- now.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it hard to have a ledger state here?

If no, we should do it right away to tie up this loose end.
If yes, it won't become easier when needing to implement this on main. Two options:

  • Figure it out now
  • Avoid it now

Message (LeiosNotify point announcement vote) StBusy StIdle
MsgLeiosBlockTxsOffer ::
!point ->
-- the closure byte range [start, end) on offer; an end of maxBound runs

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But the interval on offer can still not align with the job sizing, right?

I would have also expected us to start simple, but a complete scheme (server and client side figured out)

You say that it's easy to change network protocols (or at least feasible) without a hard-fork. So maybe the pragmatic way is to just slice it into pieces which happen to align with the job size we also pick?

@ch1bo
ch1bo force-pushed the nfrisby/leios-incremental-offers branch from a8acccb to bcfbfab Compare October 9, 2026 21:09
@ch1bo
ch1bo removed this pull request from stack #2364 October 9, 2026 21:10
@ch1bo
ch1bo force-pushed the ch1bo/leios-protocol-message-indices branch from 27e0cb3 to c54515e Compare October 9, 2026 21:52
…2363)

~~Both protocols had holes. LeiosNotify ran 0-7 only because MsgQuit and
MsgCanceled were appended after MsgDone; LeiosFetch used 0-3 and then 9,
with 6-8 held for batch messages that were never built.
Each now numbers from zero contiguously, ordered so the numbering says
something. Termination comes first in both. Then, in LeiosNotify, the
request and the responses it can draw, with votes last so that moving
them to a votes mini-protocol later costs no renumbering. In LeiosFetch,
each request sits beside the reply it draws.
The commented-out batch and vote placeholders go with them: never
implemented, and no longer planned.
This is a wire-format change on both mini-protocols, so every node has
to adopt it together.~~

This is purely cosmetic and just fits as we are changing the
mini-protocols for a good reason anyways.
An error occurred while trying to automatically change base from ch1bo/leios-protocol-message-indices to nfrisby/LeiosNotify-graceful-shutdown October 9, 2026 21:56
nfrisby and others added 2 commits October 9, 2026 23:59
… ranges

MsgLeiosBlockTxsOffer now names the closure bytes [start, end) it offers.
That lets a peer offer a closure in parts, as it acquires them, without
nodes having to agree how a closure is divided, and lets it offer later
parts out of order, since they can arrive out of order even when
requested in order.

This node still offers each closure once and whole, as [0, maxBound),
since it only offers what it has entirely acquired. As a client it now
requests only the jobs inside the ranges a peer has offered: each job
records the closure bytes it spans.

A peer may offer one endorser block's closure in several ranges, but
never the same bytes twice (ExnLeiosRepeatedOffer). A range must be
non-empty, start below maxEbClosureBytesSize, and span at least
leiosClosureOfferMinLength unless it runs to the end of the closure
(ExnLeiosInvalidClosureRange), which bounds how many offers a peer can
make per endorser block. Both bounds are stubs, awaiting the handshake
and the slot's protocol parameter respectively.

Squashed from the seven commits of nfrisby/leios-incremental-offers and
ported onto the omnibus's offer model: PeerOffer's closure field becomes
the range set, and admission is part of checkLeiosClosureOffer on the
per-peer LeiosNotify state, which is pruned to the immutable tip, in
place of a separate, never-pruned baseline map. The interim closure size
on AcquiredEbTxs and the 12 MB offer stub fall away, since the final
form offers [0, maxBound).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@ch1bo
ch1bo force-pushed the nfrisby/leios-incremental-offers branch from bcfbfab to da71a2f Compare October 9, 2026 21:59

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPIKE: Stream EB closures

2 participants