介绍

在 CometBFT 的预期行为 一节中, 我们介绍了最常见的行为,通常称为理想情况。 不过,同一节中给出的语法更加通用,覆盖了应用程序设计者需要考虑的更多场景。 本节将进一步说明这些可能出现的场景。我们重点关注 ABCI++ 引入的方法:PrepareProposal 和 ProcessProposal。更具体地说,我们关注下面这部分语法。
consensus-height    = *consensus-round decide commit
consensus-round     = proposer / non-proposer

proposer            = [prepare-proposal process-proposal]
non-proposer        = [process-proposal]
从这段语法可以看出,在决定某个区块之前,可能会经历多个轮次。一个轮次可能不足以完成决定,原因包括:
  • 网络异步性;以及
  • 提议者是拜占庭进程。
如果我们假设共识算法在第 rr 轮决定了区块 XX,那么在满足 r′<=rr' <= r 的各轮中,CometBFT 可能表现出以下任意一种行为:
  1. 为区块 XX 调用 PrepareProposal 和/或 ProcessProposal。
  2. 为区块 Y≠XY \neq X 调用 PrepareProposal 和/或 ProcessProposal。
  3. 不调用 PrepareProposal 和/或 ProcessProposal。
在进程担任提议者的轮次中,CometBFT 对 PrepareProposal 的调用总是会紧跟一次 ProcessProposal 调用。原因在于该进程也会把提案广播给自己,而该提案会在本地被投递, 从而触发 ProcessProposal 调用。 ProcessProposal 处理的提案,与此前在相同高度和轮次上任意一次 PrepareProposal 返回的内容相同。 如果没有重启,那么此前这样的调用只有一次;但如果提议者发生重启,则每次重启都可能额外产生一次 针对同一高度和轮次的 PrepareProposal 调用。 由于共识算法在某次运行中究竟需要多少轮才能完成决定,这一点事先是未知的,因此应用程序需要能够处理任意数量的轮次,并且每一轮都可能表现为上述三种行为之一。请注意,应用程序并不了解共识内部机制,因此也感知不到这些轮次。

可能的场景

在遵循共识算法时,轮次数量未知,因此可能出现大量可预期的场景。 将它们全部列出并不可行。不过,这里会给出其中若干场景,并提炼主要结论。 具体来说,我们将说明在区块 XX 被决定之前:
  1. 在一个正确节点上,PrepareProposal 可能会被多次调用,并且对应不同区块(场景 1)。
  2. 在一个正确节点上,ProcessProposal 可能会被多次调用,并且对应不同区块(场景 2)。
  3. 在一个正确节点上,针对区块 XX 的 PrepareProposal 和 ProcessProposal 可能根本不会被调用(场景 3)。
  4. 在一个正确节点上,PrepareProposal 和 ProcessProposal 甚至可能完全不会被调用(场景 4)。

基本信息

每个场景都从某个进程 pp 的视角进行描述。更准确地说,我们展示的是 Tendermint 共识算法 中每一轮的 stepstep 会发生什么。 尽管在实际中,共识算法是基于验证者投票权来运作的,但为了简化表述,本文使用进程数量 (例如 nn、f+1f+1、2f+12f+1)来说明。图例如下:

第 X 轮

  1. 提案阶段: 描述当 stepp=proposestep_p = propose 时发生的情况。
  2. 预投票阶段: 描述当 stepp=prevotestep_p = prevote 时发生的情况。
  3. 预提交阶段: 描述当 stepp=precommitstep_p = precommit 时发生的情况。

场景 1

pp 多次调用 ProcessProposal,且每次对应的值都不同。

第 0 轮

  1. 提案阶段: 本轮的提议者是一个拜占庭进程,并且它选择不发送提案消息。因此,pp 的 timeoutProposetimeoutPropose 超时,它会为 nilnil 发送 PrevotePrevote,并且不会调用 ProcessProposal。所有正确进程都会这样做。
  2. 预投票阶段: pp 最终会收到 2f+12f+1 条针对 nilnil 的 PrevotePrevote 消息,并启动 timeoutPrevotetimeoutPrevote。当 timeoutPrevotetimeoutPrevote 超时后,它会为 nilnil 发送 PrecommitPrecommit。
  3. 预提交阶段: pp 最终会收到 2f+12f+1 条针对 nilnil 的 PrecommitPrecommit 消息,并启动 timeoutPrecommittimeoutPrecommit。当其超时后,它会进入下一轮。

第 1 轮

  1. 提案阶段: 本轮的提议者是一个正确进程。它的 validValuevalidValue 为 nilnil,因此可以自由生成并提议一个新区块 YY。进程 pp 及时收到该提案,对区块 YY 调用 ProcessProposal,并为其广播一条 PrevotePrevote 消息。
  2. 预投票阶段: 由于网络异步性,为该区块发送 PrevotePrevote 的进程少于 2f+12f+1 个。因此,pp 在本轮不会更新 validValuevalidValue。
  3. 预提交阶段: 由于发送 PrevotePrevote 的进程少于 2f+12f+1 个,因此不会有正确进程锁定该区块并发送 PrecommitPrecommit 消息。结果就是,pp 不会对 YY 作出决定。

第 2 轮

  1. 提案阶段: 与 第 1 轮 相同,只是这次由另一个正确进程担任提议者,并提议另一个值 ZZ。进程 pp 及时收到该提案,对新区块 ZZ 调用 ProcessProposal,并为其广播一条 PrevotePrevote 消息。
  2. 预投票阶段: 与 第 1 轮 相同。
  3. 预提交阶段: 与 第 1 轮 相同。
像这样的轮次可能会持续发生,直到进程 pp 在某一轮更新了它的 validValuevalidValue,或者直到到达第 rr 轮并由进程 pp 对某个区块作出决定。在那之后,它将不会再针对该高度调用 ProcessProposal。

场景 2

pp 多次调用 PrepareProposal,且每次对应的值都不同。

第 0 轮

  1. 提案阶段: 进程 pp 是本轮的提议者。它的 validValuevalidValue 为 nilnil,因此可以自由生成并提议新区块 YY。在提议之前,它会先针对 YY 调用 PrepareProposal。之后,它会广播该提案,将其投递给自己,调用 ProcessProposal,并为其广播 PrevotePrevote。
  2. 预投票阶段: 由于网络异步性,及时收到该提案并为其发送 PrevotePrevote 的进程少于 2f+12f+1 个。因此,pp 在本轮不会更新 validValuevalidValue。
  3. 预提交阶段: 由于发送 PrevotePrevote 的进程少于 2f+12f+1 个,因此不会有正确进程锁定该区块,也不会发送非 nilnil 的 PrecommitPrecommit 消息。结果就是,pp 不会对 YY 作出决定。
在这一轮之后,可能会出现多个与 场景 1 类似的轮次。关键在于,进程 pp 不应更新它的 validValuevalidValue。因此,当进程 pp 进入它再次担任提议者的轮次时,它会再次向 mempool 请求新区块,而 mempool 可能返回另一个不同的区块 ZZ,于是就可能再次出现与 第 0 轮 相同的情况,只不过对应的是另一个区块。结果就是,进程 pp 会再次调用 PrepareProposal,但对应的是不同的值。当到达第 rr 轮时,某个进程会提议区块 XX;如果 pp 收到 2f+12f+1 条 PrecommitPrecommit 消息,它就会对该值作出决定。

场景 3

pp 会针对多个值调用 PrepareProposal 和 ProcessProposal,但最终决定的值却不是它调用过 PrepareProposal 或 ProcessProposal 的那个值。 在这个场景中,在第 rr 轮之前的所有轮次里,都可能出现 场景 1 或 场景 2 中展示的任意轮次。需要注意的是:
  • 没有任何提议者提议过区块 XX,或者即使提议过,由于异步性,进程 pp 也没有及时收到,因此它没有调用 ProcessProposal;并且
  • 如果 pp 是提议者,那么它提议的是另一个不等于 XX 的值。

第 rr 轮

  1. 提案阶段: 本轮的提议者是一个正确进程,并且它提议区块 XX。由于异步性,该提案消息在进程 pp 的 timeoutProposetimeoutPropose 超时并为 nilnil 发送 PrevotePrevote 之后才到达。因此,进程 pp 不会针对区块 XX 调用 ProcessProposal。不过,同一个提案会在其他进程的 timeoutProposetimeoutPropose 超时之前送达,于是它们会为该提案发送 PrevotePrevote。
  2. 预投票阶段: 进程 pp 收到 2f+12f+1 条针对提案 XX 的 PrevotePrevote 消息,相应地更新它的 validValuevalidValue 和 lockedValuelockedValue,并发送 PrecommitPrecommit 消息。所有正确进程都会这样做。
  3. 预提交阶段: 最终,进程 pp 收到 2f+12f+1 条 PrecommitPrecommit 消息,并对区块 XX 作出决定。

场景 4

可以将 场景 3 转换为另一种场景:pp 完全不调用 PrepareProposal 和 ProcessProposal。要满足这一点,进程 pp 必须在所有满足 0<=r′<=r0 <= r' <= r 的轮次中都不是提议者, 并且由于网络异步性或提议者是拜占庭进程,它在 timeoutProposetimeoutPropose 超时之前始终收不到提案。 结果就是,在进入第 rr 轮之前,它从未调用过 PrepareProposal 和 ProcessProposal;并且正如 场景 3 的第 rr 轮所示,它会在这一轮完成决定。同样,整个过程中不会调用这两个方法。

Introduction

In the section CometBFT’s expected behaviour, we presented the most common behaviour, usually referred to as the good case. However, the grammar specified in the same section is more general and covers more scenarios that an Application designer needs to account for. In this section, we give more information about these possible scenarios. We focus on methods introduced by ABCI++: PrepareProposal and ProcessProposal. Specifically, we concentrate on the part of the grammar presented below.
consensus-height    = *consensus-round decide commit
consensus-round     = proposer / non-proposer

proposer            = [prepare-proposal process-proposal]
non-proposer        = [process-proposal]
We can see from the grammar that we can have several rounds before deciding a block. The reasons why one round may not be enough are:
  • network asynchrony, and
  • a Byzantine process being the proposer.
If we assume that the consensus algorithm decides on block XX in round rr, in the rounds r′<=rr' <= r, CometBFT can exhibit any of the following behaviours:
  1. Call PrepareProposal and/or ProcessProposal for block XX.
  2. Call PrepareProposal and/or ProcessProposal for block Y≠XY \neq X.
  3. Does not call PrepareProposal and/or ProcessProposal.
In the rounds in which the process is the proposer, CometBFT’s PrepareProposal call is always followed by the ProcessProposal call. The reason is that the process also broadcasts the proposal to itself, which is locally delivered and triggers the ProcessProposal call. The proposal processed by ProcessProposal is the same as what was returned by any of the preceding PrepareProposal invoked for the same height and round. While in the absence of restarts there is only one such preceding invocations, if the proposer restarts there could have been one extra invocation to PrepareProposal for each restart. As the number of rounds the consensus algorithm needs to decide in a given run is a priori unknown, the application needs to account for any number of rounds, where each round can exhibit any of these three behaviours. Recall that the application is unaware of the internals of consensus and thus of the rounds.

Possible scenarios

The unknown number of rounds we can have when following the consensus algorithm yields a vast number of scenarios we can expect. Listing them all is unfeasible. However, here we give several of them and draw the main conclusions. Specifically, we will show that before block XX is decided:
  1. On a correct node, PrepareProposal may be called multiple times and for different blocks (Scenario 1).
  2. On a correct node, ProcessProposal may be called multiple times and for different blocks (Scenario 2).
  3. On a correct node, PrepareProposal and ProcessProposal for block XX may not be called (Scenario 3).
  4. On a correct node, PrepareProposal and ProcessProposal may not be called at all (Scenario 4).

Basic information

Each scenario is presented from the perspective of a process pp. More precisely, we show what happens in each round’s stepstep of the Tendermint consensus algorithm. While in practice the consensus algorithm works with respect to voting power of the validators, in this document we refer to number of processes (e.g., nn, f+1f+1, 2f+12f+1) for simplicity. The legend is below:

Round X

  1. Propose: Describes what happens while stepp=proposestep_p = propose.
  2. Prevote: Describes what happens while stepp=prevotestep_p = prevote.
  3. Precommit: Describes what happens while stepp=precommitstep_p = precommit.

Scenario 1

pp calls ProcessProposal many times with different values.

Round 0

  1. Propose: The proposer of this round is a Byzantine process, and it chooses not to send the proposal message. Therefore, pp‘s timeoutProposetimeoutPropose expires, it sends PrevotePrevote for nilnil, and it does not call ProcessProposal. All correct processes do the same.
  2. Prevote: pp eventually receives 2f+12f+1 PrevotePrevote messages for nilnil and starts timeoutPrevotetimeoutPrevote. When timeoutPrevotetimeoutPrevote expires it sends PrecommitPrecommit for nilnil.
  3. Precommit: pp eventually receives 2f+12f+1 PrecommitPrecommit messages for nilnil and starts timeoutPrecommittimeoutPrecommit. When it expires, it moves to the next round.

Round 1

  1. Propose: A correct process is the proposer in this round. Its validValuevalidValue is nilnil, and it is free to generate and propose a new block YY. Process pp receives this proposal in time, calls ProcessProposal for block YY, and broadcasts a PrevotePrevote message for it.
  2. Prevote: Due to network asynchrony less than 2f+12f+1 processes send PrevotePrevote for this block. Therefore, pp does not update validValuevalidValue in this round.
  3. Precommit: Since less than 2f+12f+1 processes send PrevotePrevote, no correct process will lock on this block and send PrecommitPrecommit message. As a consequence, pp does not decide on YY.

Round 2

  1. Propose: Same as in Round 1, just another correct process is the proposer, and it proposes another value ZZ. Process pp receives the proposal on time, calls ProcessProposal for new block ZZ, and broadcasts a PrevotePrevote message for it.
  2. Prevote: Same as in Round 1.
  3. Precommit: Same as in Round 1.
Rounds like these can continue until we have a round in which process pp updates its validValuevalidValue or until we reach round rr where process pp decides on a block. After that, it will not call ProcessProposal anymore for this height.

Scenario 2

pp calls PrepareProposal many times with different values.

Round 0

  1. Propose: Process pp is the proposer in this round. Its validValuevalidValue is nilnil, and it is free to generate and propose new block YY. Before proposing, it calls PrepareProposal for YY. After that, it broadcasts the proposal, delivers it to itself, calls ProcessProposal and broadcasts PrevotePrevote for it.
  2. Prevote: Due to network asynchrony less than 2f+12f+1 processes receive the proposal on time and send PrevotePrevote for it. Therefore, pp does not update validValuevalidValue in this round.
  3. Precommit: Since less than 2f+12f+1 processes send PrevotePrevote, no correct process will lock on this block and send non-nilnil PrecommitPrecommit message. As a consequence, pp does not decide on YY.
After this round, we can have multiple rounds like those in Scenario 1. The important thing is that process pp should not update its validValuevalidValue. Consequently, when process pp reaches the round when it is again the proposer, it will ask the mempool for the new block again, and the mempool may return a different block ZZ, and we can have the same round as Round 0 just for a different block. As a result, process pp calls PrepareProposal again but for a different value. When it reaches round rr some process will propose block XX and if pp receives 2f+12f+1 PrecommitPrecommit messages, it will decide on this value.

Scenario 3

pp calls PrepareProposal and ProcessProposal for many values, but decides on a value for which it did not call PrepareProposal or ProcessProposal. In this scenario, in all rounds before rr we can have any round presented in Scenario 1 or Scenario 2. What is important is that:
  • no proposer proposed block XX or if it did, process pp, due to asynchrony, did not receive it in time, so it did not call ProcessProposal, and
  • if pp was the proposer it proposed some other value ≠X\neq X.

Round rr

  1. Propose: A correct process is the proposer in this round, and it proposes block XX. Due to asynchrony, the proposal message arrives to process pp after its timeoutProposetimeoutPropose expires and it sends PrevotePrevote for nilnil. Consequently, process pp does not call ProcessProposal for block XX. However, the same proposal arrives at other processes before their timeoutProposetimeoutPropose expires, and they send PrevotePrevote for this proposal.
  2. Prevote: Process pp receives 2f+12f+1 PrevotePrevote messages for proposal XX, updates correspondingly its validValuevalidValue and lockedValuelockedValue and sends PrecommitPrecommit message. All correct processes do the same.
  3. Precommit: Finally, process pp receives 2f+12f+1 PrecommitPrecommit messages, and decides on block XX.

Scenario 4

Scenario 3 can be translated into a scenario where pp does not call PrepareProposal and ProcessProposal at all. For this, it is necessary that process pp is not the proposer in any of the rounds 0<=r′<=r0 <= r' <= r and that due to network asynchrony or Byzantine proposer, it does not receive the proposal before timeoutProposetimeoutPropose expires. As a result, it will enter round rr without calling PrepareProposal and ProcessProposal before it, and as shown in Round rr of Scenario 3 it will decide in this round. Again without calling any of these two calls.