正式要求

共识连接要求

本节说明 CometBFT 对应用程序的期望。其组织形式是一组正式要求,可用于测试和验证应用程序逻辑。 设 p 和 q 是两个正确进程。 设 rp(分别地,rq)是在高度 h 上、由 p(分别地,q)担任提议者的一个轮次。 设 sp,h-1 是 p 的应用程序在高度 h-1 已提交的状态。 设 vp(分别地,vq)是 p 的(分别地,q 的)CometBFT 作为高度 h、轮次 rp(分别地,rq)的提议者, 通过 RequestPrepareProposal 传递给应用程序的区块, 也称为原始提案。 设 up(分别地,uq)是 p 的(分别地,q 的)应用程序 通过 ResponsePrepareProposal 返回给 CometBFT 的、可能经过修改的区块,也称为准备后的提案。 当 p 在不同轮次中担任提议者时,进程 p 的准备后提案可能不同。
  • 要求 1 [PrepareProposal, 及时性]:如果 p 的应用程序在 PrepareProposal 中会完整执行准备后的区块,并且当进程 p 和 q 处于 rp 时网络处于同步时期, 那么 q 处的 TimeoutPropose 取值必须保证 q 的提议计时器不会超时 (否则会导致 q 在 rp 中对 nil 进行 prevote)。
在 PrepareProposal 阶段完整执行区块位于 CometBFT 的关键路径上。因此, 要求 1 确保应用程序或运维者会为 TimeoutPropose 设置一个取值,使得 在 PrepareProposal 中完整执行区块所需的时间不会干扰 CometBFT 的提议计时器。 请注意,违反要求 1 可能会导致进入后续轮次,但不会 破坏活性,因为尽管 TimeoutPropose 被用作提议超时的初始值, CometBFT 会动态调整这些超时, 使其最终足以完成 PrepareProposal。
  • 要求 2 [PrepareProposal, 交易大小]:当 p 的应用程序调用 ResponsePrepareProposal 时, 返回的交易总字节数不得超过 RequestPrepareProposal.max_tx_bytes。
繁忙的区块链可能希望能够完整看到 CometBFT 内存池中的交易, 而不只是看到其中适合装入区块的某个子集。 应用程序可以通过将 ConsensusParams.Block.MaxBytes 设置为 -1 来实现这一点。 这会指示 CometBFT:(a) 在 CometBFT 层面强制使用 MaxBytes 的最大可能值(100 MB), 以及 (b) 在调用 RequestPrepareProposal 时提供内存池中的全部交易。 在这些设置下,所有交易的聚合大小可能会超过 RequestPrepareProposal.max_tx_bytes。 因此,要求 2 确保应用程序返回的交易列表在字节数上永远不会 导致生成的区块超出其字节大小上限。
  • 要求 3 [PrepareProposal, ProcessProposal, 一致性]:对于任意两个正确进程 p 和 q, 如果 q 的 CometBFT 对 up 调用 RequestProcessProposal, 则 q 的应用程序会在 ResponseProcessProposal 中返回 Accept。
要求 3 确保由正确进程提出的区块总是能够通过正确接收进程的 ProcessProposal 检查。 另一方面,如果 PrepareProposal 或 ProcessProposal(或二者同时)中存在确定性缺陷, 严格来说,这会使所有触发该缺陷的进程都变成拜占庭进程。这在实践中是个问题, 因为验证者往往运行的是来自同一代码库的应用程序,因此潜在地所有验证者都可能 同时触发该缺陷。这会导致大多数(甚至全部)进程对 nil 进行 prevote, 从而对 CometBFT 的活性造成严重影响。由于其关键性,要求 3 是 广泛测试和自动化验证的重点目标。
  • 要求 4 [ProcessProposal, 确定性-1]:ProcessProposal 是当前状态 和即将应用的区块的一个(确定性)函数。换句话说,对于任意正确进程 p 以及任意区块 u, 如果 p 的 CometBFT 在高度 h 上对 u 调用 RequestProcessProposal, 那么 p 的应用程序是否接受或拒绝 仅仅 取决于 u 和 sp,h-1。
  • 要求 5 [ProcessProposal, 确定性-2]:对于任意两个正确进程 p 和 q,以及任意 区块 u, 如果 p 的(分别地,q 的)CometBFT 在高度 h 上对 u 调用 RequestProcessProposal, 那么 p 的应用程序接受 u 当且仅当 q 的应用程序接受 u。 请注意,该要求可由要求 4 和共识的 Agreement 性质推出。
要求 4 和要求 5 确保所有正确进程都会以相同方式响应被提议的区块,即使 提议者是拜占庭进程也是如此。然而,ProcessProposal 可能包含某种缺陷,使得 区块的接受或拒绝变为非确定性的,因此会使触发该缺陷的进程 无法满足要求 4 或要求 5(实际上使这些进程变成拜占庭进程)。 在这种场景下,CometBFT 的活性无法得到保证。 同样,如果大多数验证者运行的是同一套软件,这在实践中也是个问题,因为他们很可能 在相同位置触发该缺陷。目前还没有明确的解决方案可以缓解这种情况,因此 应用程序的设计者/实现者在 ProcessProposal 的逻辑/实现上 必须非常谨慎。一般来说,ProcessProposal SHOULD 总是接受该区块。 根据当前由 CometBFT 采用的 Tendermint 共识算法, 一个正确进程在轮次 r、高度 h 上 最多只能广播一条 precommit 消息。 因此,正如 Methods 一节所述,ResponseExtendVote 只会在共识算法 即将广播一条非 nil 的 precommit 消息时被调用,所以一个正确进程在 轮次 r、高度 h 上也只能产生一个投票扩展。 设 erp 是正确进程 p 的应用程序在 轮次 r、高度 h 上通过 ResponseExtendVote 返回的投票扩展。 设 wrp 是 p 的 CometBFT 在轮次 r、高度 h 上通过 RequestExtendVote 传递给应用程序的被提议区块。
  • 要求 6 [ExtendVote, VerifyVoteExtension, 一致性]:对于任意两个不同的正确 进程 p 和 q,如果 q 在高度 h 上从 p 收到 erp,则 q 的 应用程序会在 ResponseVerifyVoteExtension 中返回 Accept。
要求 6 以与要求 3 约束提议区块的创建和处理类似的方式,约束投票扩展的创建与处理。 要求 6 确保由正确进程创建的扩展总是能够通过接收这些扩展的正确进程所执行的 VerifyVoteExtension 检查。 然而,如果 ExtendVote 或 VerifyVoteExtension(或二者同时)中存在(确定性)缺陷, 我们将面临与要求 5 所描述相同的活性问题,因为带有无效投票 扩展的 Precommit 消息会被丢弃。
  • 要求 7 [VerifyVoteExtension, 确定性-1]:VerifyVoteExtension 是关于 当前状态、收到的投票扩展以及该扩展所指向的准备后提案的一个(确定性)函数。 换句话说,对于任意正确进程 p、任意投票扩展 e 以及任意 区块 w,如果 p 的(分别地,q 的)CometBFT 在高度 h 上对 e 和 w 调用 RequestVerifyVoteExtension, 那么 p 的应用程序是否接受或拒绝 仅仅 取决于 e、w 和 sp,h-1。
  • 要求 8 [VerifyVoteExtension, 确定性-2]:对于任意两个正确进程 p 和 q, 以及任意投票扩展 e 和任意区块 w, 如果 p 的(分别地,q 的)CometBFT 在高度 h 上对 e 和 w 调用 RequestVerifyVoteExtension, 那么 p 的应用程序接受 e 当且仅当 q 的应用程序接受 e。 请注意,该要求可由要求 7 和共识的 Agreement 性质推出。
要求 7 和要求 8 确保投票扩展的验证在所有 正确进程上都是确定性的。 要求 7 和要求 8 用于防御来自拜占庭进程的任意投票扩展数据, 方式类似于要求 4 和要求 5 对任意提议区块的防御。 要求 7 和要求 8 可能会因在 VerifyVoteExtension 中引入非确定性的缺陷而被破坏。在这种情况下,活性可能受损。 在实现 ExtendVote 和 VerifyVoteExtension 时应格外谨慎。 一般来说,VerifyVoteExtension SHOULD 总是接受该投票扩展。
  • 要求 9 [全部, 无副作用]:p 在高度 h 上对 RequestPrepareProposal、 RequestProcessProposal、RequestExtendVote 和 RequestVerifyVoteExtension 的调用 不会修改 sp,h-1。
  • 要求 10 [ExtendVote, FinalizeBlock, 非依赖性]:对于任意正确进程 p, 以及 p 在高度 h 收到的任意投票扩展 e, sp,h 的计算不依赖于 e。
对正确进程 p 在高度 h 上的 RequestFinalizeBlock 调用,以区块 vp,h 作为传入参数,会创建状态 sp,h。 此外,p 的 FinalizeBlock 还会创建一组交易结果 Tp,h。
  • 要求 11 [FinalizeBlock, 确定性-1]:对于任意正确进程 p, sp,h 仅取决于 sp,h-1 和 vp,h。
  • 要求 12 [FinalizeBlock, 确定性-2]:对于任意正确进程 p, Tp,h 的内容仅取决于 sp,h-1 和 vp,h。
请注意,要求 11 和要求 12 结合共识的 Agreement 性质,可确保 状态机复制,即应用程序状态会在所有正确进程上一致演进。 另外还要注意,PrepareProposal 和 ExtendVote 都没有与确定性相关的 要求。 确实,PrepareProposal 不要求是确定性的:
  • up 可以依赖于 vp 和 sp,h-1,但也可以依赖于其他值或操作。
  • vp = vq ⇏ up = uq。
同样,ExtendVote 也可以是非确定性的:
  • erp 可以依赖于 wrp 和 sp,h-1, 但也可以依赖于其他值或操作。
  • wrp = wrq ⇏ erp = erq

内存池连接要求

令 CheckTxCodestx,p,h 表示进程 p 的应用通过 ResponseCheckTx 对连续调用 RequestCheckTx 时返回的结果码集合,这些调用发生在应用处于高度 h 期间,且交易 tx 作为参数。 CheckTxCodestx,p,h 是一个集合,因为 p 的应用在高度 h 期间可能返回不同的结果码。 如果 CheckTxCodestx,p,h 是单元素集合,也就是说应用在高度 h 期间的 ResponseCheckTx 中始终返回相同的结果码, 我们将 CheckTxCodetx,p,h 定义为 CheckTxCodestx,p,h 的那个唯一值。 如果 CheckTxCodestx,p,h 不是单元素集合,则 CheckTxCodetx,p,h 未定义。 令谓词 OK(CheckTxCodetx,p,h) 表示 CheckTxCodetx,p,h 是否为 SUCCESS。
  • 要求 13 [CheckTx,最终不振荡]:对于任意交易 tx, 存在一个布尔值 b, 以及一个高度 hstable,使得 对任意正确进程 p, CheckTxCodetx,p,h 都是已定义的,并且 对任意高度 h ≥ hstable,都有 OK(CheckTxCodetx,p,h) = b。
要求 13 保证,如果一笔交易在 p 的内存池中停留足够长时间, 它最终会停止在 CheckTx 成功与失败之间来回振荡。 对应用行为的这一约束使内存池能够确保 一笔交易最终会离开所有全节点的内存池: 要么因为 CheckTx 调用失败而在各处被清除, 要么因为其保持有效足够长时间,从而被 gossip、被提议并被决定。 虽然要求 13 定义了一个全局的 hstable,应用开发者 可以不失一般性地将这种稳定高度视为进程 p 的局部值(hp,stable)。 相比之下,b 的值在所有进程中都必须相同。

管理应用状态及相关主题

连接状态

CometBFT 维护四个并发的 ABCI++ 连接,即 共识连接、 内存池连接、 Info/Query 连接 和 快照连接。 应用通常会为每个连接维护一份独立的状态副本,并在 Commit 调用时进行同步。

并发性

原则上,四个 ABCI++ 连接彼此并发运行。 这意味着应用需要确保对状态的访问是线程安全的。 默认的进程内 ABCI 客户端 和 默认的 Go ABCI 服务器 都使用一个全局锁来保护跨所有连接的事件处理,因此它们实际上完全不是并发的。 这意味着,无论你的应用是使用 NewLocalClient 与 CometBFT 一起编译为进程内模式, 还是使用 SocketServer 以进程外方式运行, 来自所有连接的 ABCI 消息都会按顺序、一次一条地被接收。 这个全局互斥锁的存在意味着,Go 应用开发者只要把所有读写都通过 ABCI 系统路由,就可以为应用状态获得线程安全保证。因此,直接向 RPC 接口暴露应用状态可能是不安全的;除非采取了明确的额外措施,否则所有查询都应通过 ABCI 的 Query 方法路由。

FinalizeBlock

当共识算法决定了一个区块后,CometBFT 会使用 FinalizeBlock 将该已决定区块的数据发送给应用,应用据此推进自身状态,但不得持久化该状态;持久化必须在 Commit 期间完成。 应用必须记住它最近一次成功执行 Commit 时的高度, 这样在从崩溃中恢复时,它才能告知 CometBFT 应该从哪里继续。 参见这里关于握手的信息。

Commit

应用应在 Commit 期间持久化其状态,并在返回之前完成。 在调用 Commit 之前,CometBFT 会锁定内存池并刷新内存池连接。这可以确保 在这个处理步骤期间, 内存池连接上不会收到新消息,从而提供一个安全的时机, 以便同时将四个连接的状态都更新到最新的已提交状态。 当 Commit 返回时,CometBFT 会解锁内存池。 警告:如果 ABCI 应用在处理 Commit 消息的逻辑中发送了 /broadcast_tx_sync 或 /broadcast_tx,并在继续之前等待响应, 就会发生死锁。执行 broadcast_tx 调用 需要获取 CometBFT 在 Commit 调用期间持有的内存池锁。 必须避免将与内存池相关的同步调用纳入 Commit 函数的顺序逻辑中。

候选状态

当 CometBFT 即将把一个提议区块发送到网络时,会调用 PrepareProposal。 同样地,当 CometBFT 从网络接收到一个提议区块时,会调用 ProcessProposal。 这两个方法向应用披露的提议区块数据包括:
  • 交易列表
  • 指向前一个区块的 LastCommit
  • 区块头的哈希值(PrepareProposal 中除外,因为此时尚未知晓)
  • 行为异常的验证者列表
  • 区块时间戳
  • NextValidatorsHash
  • 提议者地址
应用可以决定立即执行给定区块(即在 PrepareProposal 或 ProcessProposal 时执行)。应用可能希望这样做,主要有两个原因:
  • 避免区块中包含无效交易。 为了确保区块中不包含任何无效交易,可能 除了像对待一个已决定区块那样完整执行区块中的交易之外,没有其他办法。
  • 加快 FinalizeBlock 执行。 当通过 FinalizeBlock 接收到已决定区块时,如果同一个区块此前已经在 PrepareProposal 或 ProcessProposal 时执行过,且由此得到的状态仍保存在内存中, 应用就可以直接将该状态(更快)应用到主状态,而不是重新执行 这个已决定区块(更慢)。
在给定高度上,PrepareProposal/ProcessProposal 可能会被调用很多次。 此外,无法准确预测某个高度上被提议的多个区块中,最终哪一个会被决定, 并在该高度的 FinalizeBlock 中交付给应用。 因此,执行某个提议区块后得到的状态(称为候选状态) 应保存在内存中,作为该高度可能的最终状态。当调用 FinalizeBlock 时,应用应当 检查这个已决定区块是否对应于它的某个候选状态;如果对应,应用就会将其作为 自己的 ExecuteTxState(参见下文的共识连接)进行应用, 并在接下来的 Commit 调用期间持久化。 在不利条件下(例如网络不稳定),共识算法可能需要很多轮。 这种情况下,对于给定高度,可能会有许多提议区块被披露给应用。 由于 CometBFT 当前采用的 Tendermint 共识算法的特性,应用在某一特定高度接收到的提议区块数量 无法被限定,因此应用开发者必须谨慎处理,并使用机制 来限制内存使用。一般来说,应用应准备好在 FinalizeBlock 之前丢弃候选状态, 即使其中某个候选状态最终可能恰好对应于 已决定区块,从而不得不在 FinalizeBlock 时重新执行。

状态与 ABCI++ 连接

共识连接

共识连接应维护一个 ExecuteTxState,即区块执行的工作状态。 它应在区块执行期间由 FinalizeBlock 调用更新, 并在 Commit 期间作为“最新已提交状态”提交到磁盘。 对提议区块的执行(通过 PrepareProposal/ProcessProposal) 不得更新 ExecuteTxState,而应当作为独立的候选状态保留,直到 FinalizeBlock 确认哪些候选状态(如果有)可以用于更新 ExecuteTxState。

内存池连接

内存池连接维护 CheckTxState。CometBFT 会按顺序针对 CheckTxState 处理传入的 交易(通过来自客户端的 RPC 或来自 gossip 层的 P2P)。 如果处理未返回任何错误,该交易就会被接收到内存池中, 随后 CometBFT 开始对其进行 gossip 传播。 CheckTxState 应在每次 Commit 结束时 重置为最新的已提交状态。 在一次共识实例的执行期间,CheckTxState 可能会与 ExecuteTxState 并发更新,因为消息可能会在共识连接和内存池连接上并发发送。 如上所述,在共识实例结束时,CometBFT 会在调用 Commit 之前锁定内存池并刷新 内存池连接。这可以确保所有待处理的 CheckTx 调用都已 得到响应,并且新的调用无法开始。 在 Commit 调用返回后,CometBFT 在仍持有内存池锁的情况下,会对节点本地内存池中 在过滤掉已被包含进区块的交易之后仍然保留的所有交易再次运行 CheckTx。 RequestCheckTx 中的参数 Type 用于指示一笔传入交易是新交易(CheckTxType_New),还是一次 重检(CheckTxType_Recheck)。 最后,在重新检查完内存池中的交易后,CometBFT 会解锁 内存池连接。新交易又可以再次通过 CheckTx 被处理。 请注意,CheckTx 只是一个较弱的过滤器,用于把无效交易挡在内存池之外,并最终挡在区块链之外。 由于无法保证交易在 CheckTx 中被检查时所依据的状态,与它作为一个(潜在)已决定区块的一部分执行时所依据的状态完全相同,因此 CheckTx 不应检查影响交易有效性的所有因素,特别是不应检查那些其有效性可能依赖于交易顺序的条件。CheckTx 之所以是弱的,是因为拜占庭节点并不需要在意 CheckTx;如果它愿意,完全可以提议一个塞满无效交易的区块。ABCI++ 用于处理这类行为的机制是 ProcessProposal。
重放保护
旧交易有可能再次被发送给应用。通常这对绝大多数交易来说 都是不希望发生的,只有其中通常很小的一部分幂等交易例外。 内存池提供了一种机制来防止重复交易被处理。 不过,这种机制只是尽力而为的(当前基于索引器), 并不能提供不重复的任何保证。 因此,应用需要在 CheckTx 的逻辑中自行实现 具备强保证的、特定于应用的重放保护机制。

Info/Query 连接

Info(或 Query)连接应维护一个 QueryState。该连接有两个 用途:1)让应用回答 CometBFT 从用户接收到的查询 (参见Query一节), 2)在启动时同步 CometBFT 与应用(参见 崩溃恢复) 或在状态同步之后进行同步(参见状态同步)。 QueryState 是 ExecuteTxState 在最近一次 Commit 之后的只读副本,也就是 在完整区块处理完毕并且状态已提交到磁盘之后的状态。

快照连接

快照连接用于为其他节点提供状态同步快照, 和/或将状态同步快照恢复到正在引导启动的本地节点。 快照管理是可选的:应用可以选择不实现它。 更多信息请参见状态同步一节。

交易结果

应用程序应在 ResponseFinalizeBlock 中返回一个 ExecTxResult 列表。交易结果列表 必须与通过 RequestFinalizeBlock 传递的交易列表保持相同顺序。 本节将讨论该结构中的字段,以及 ResponseCheckTx 中语义相近的字段。 Info 和 Log 字段是 用于调试或便利性的非确定性值。CometBFT 会记录它们,但除此之外会忽略它们。

Gas

以太坊引入了 gas 的概念,将其作为节点处理交易时所消耗资源成本的抽象表示。以太坊虚拟机中的每个操作都会消耗一定数量的 gas。 Gas 的价格是由市场决定且可变的,矿工可以基于该价格接受或拒绝执行某个特定操作。 用户会为其交易提出一个最大 gas 用量;如果交易实际使用得更少,差额会返还给用户。CometBFT 采用了类似的抽象, 但它只是可选且较弱地使用该机制,允许应用程序自行定义 执行成本的含义。 在 CometBFT 中,ConsensusParams.Block.MaxGas 限制了 一个区块中所有交易可使用的 gas 总量。 默认值是 -1,表示不强制执行区块 gas 限制,或者 gas 这一概念 没有实际意义。 响应中包含 GasWanted 和 GasUsed 字段。前者表示交易发送者愿意使用的 gas 最大 数量,后者表示实际使用量。应用程序应确保 GasUsed <= GasWanted — 即交易执行 或校验应在使用超过其所请求资源之前失败。 当 MaxGas > -1 时,CometBFT 会强制执行以下规则:
  • 对于 mempool 中的每笔交易,GasWanted <= MaxGas
  • 在提议区块时,(sum of GasWanted in a block) <= MaxGas
如果 MaxGas == -1,则不会强制执行任何与 gas 有关的规则。 在 v0.34.x 及更早版本中,CometBFT 在共识阶段不会对 Gas 强制执行任何限制, 只会在 mempool 中处理。 这意味着它不保证已提交的区块满足这些规则。 当执行区块中的交易时,如果超出 gas 限制,应用程序有责任返回非零响应码。 自 v0.37.x 引入 PrepareProposal 和 ProcessProposal 以来,应用程序现在可以 确保在共识中所有被提议的区块(以及被投票的区块)— 因而所有最终决定的区块 — 都遵守上述 MaxGas 限制。 由于应用程序在执行交易时应确保 GasUsed <= GasWanted,并且 它可以使用 PrepareProposal 和 ProcessProposal 来保证所有被提议或预投票的区块中 (sum of GasWanted in a block) <= MaxGas, 因此有:
  • 对于每个区块,(sum of GasUsed in a block) <= MaxGas
GasUsed 字段会被 CometBFT 忽略。

ResponseCheckTx 的细节

如果 Code != 0,该交易将被 mempool 拒绝,因此 不会广播给其他对等节点,也不会被包含进提议区块。 Data 包含 CheckTx 交易执行的结果(如果有)。它不需要是 确定性的,因为对于同一笔交易,不同节点上的应用程序在接收该交易并通过 CheckTx 检查其有效性时,可能具有不同的 CheckTxState 值。 CometBFT 会忽略 ResponseCheckTx 中的这个值。 从 v0.34.x 开始,ResponseCheckTx 中有一个 Priority 字段,可以 用于在 mempool 中显式地为交易设定优先级,以便将其纳入区块提议。

ExecTxResult 的细节

FinalizeBlock 是区块链中的核心执行环节。CometBFT 会将已决定的区块 以及其中所有交易的列表同步交付给应用程序。 根据共识的 Agreement 性质保证,交付的区块(因此交易顺序) 在所有正确节点上都是相同的。 ExecTxResult 中的 Data 字段包含一个表示交易结果的字节数组。 它必须是确定性的(即所有节点必须返回相同的值),但其中可以包含任意 数据。同样,Code 的值也必须是确定性的。 如果 Code != 0,该交易将被标记为无效, 尽管它仍然会包含在区块中。无效交易不会被索引,因为它们被视为 类似于那些未通过 CheckTx 的交易。 Code 和 Data 都会被包含在一个结构中,该结构会被哈希到 下一个高度的区块头 LastResultsHash 中。 Events 包含执行期间产生的任何事件,CometBFT 会使用它们来建立 该交易的索引。这样就可以根据交易执行期间发生的 事件来查询交易。

更新验证者集合

应用程序可以在 InitChain 期间设置验证者集合,也可以在 FinalizeBlock 期间更新它。在这两种情况下,都会返回一个 ValidatorUpdate 类型的结构。 用于初始化应用程序的 InitChain 方法可以返回一个验证者列表。 如果该列表为空,CometBFT 将使用从创世 文件中加载的验证者。 如果 InitChain 返回的列表非空,CometBFT 将使用其内容作为验证者集合。 这样,应用程序就可以为该 区块链设置初始验证者集合。 应用程序必须确保单个验证者更新集合中不包含重复项,也就是说, 同一个公钥在一次更新中只能出现一次。如果某次更新包含 重复项,区块执行将发生不可恢复的失败。 ValidatorUpdate 结构包含一个用于标识验证者的公钥: 当前该公钥支持三种类型:
  • ed25519
  • secp256k1
  • bls12381
ValidatorUpdate 结构还包含一个 ìnt64 字段,用于表示验证者的新投票权重。 应用程序必须确保 ValidatorUpdate 结构遵守以下规则:
  • 权重必须为非负数
  • 如果权重被设置为 0,则该验证者必须已经在验证者集合中;它将从集合中移除
  • 如果权重大于 0:
    • 如果该验证者不在验证者集合中,则会以给定权重将其加入 集合
    • 如果该验证者已在验证者集合中,其权重将被调整为给定值
  • 新验证者集合的总权重不得超过 MaxTotalVotingPower,其中 MaxTotalVotingPower = MaxInt64 / 8
请注意,在处理高度为 H 的区块后返回的更新,只有到区块 H+2 时才会生效 (参见 Methods 一节)。

共识参数

ConsensusParams 是适用于区块链中所有验证者的全局参数。 它们会在区块链中强制执行某些限制,例如区块的最大大小、 区块中使用的 gas 数量,以及证据可接受的最大年龄。 它们可以在 InitChain 中设置,也可以在 FinalizeBlock 中更新。 这些参数由应用程序以确定性方式设置和/或更新,因此 所有全节点在给定高度上都具有相同的值。

参数列表

以下是当前的共识参数(截至 v0.38.x):
  1. ABCIParams.VoteExtensionsEnableHeight
  2. BlockParams.MaxBytes
  3. BlockParams.MaxGas
  4. EvidenceParams.MaxAgeDuration
  5. EvidenceParams.MaxAgeNumBlocks
  6. EvidenceParams.MaxBytes
  7. ValidatorParams.PubKeyTypes
  8. VersionParams.App

ABCIParams.VoteExtensionsEnableHeight

该参数要么为 0,要么为一个正高度,表示投票扩展 从该高度开始变为强制要求。如果该值为 0(默认值),则不要求 提供投票扩展。否则,在所有大于所配置高度 H 的高度上, 都必须存在投票扩展(即使其内容为空)。 当达到配置的高度 H 时,PrepareProposal 还不会 包含投票扩展,但会调用 ExtendVote 和 VerifyVoteExtension。 随后,当达到高度 H+1 时,PrepareProposal 将 包含来自高度 H 的投票扩展。对于所有高于 H 的高度
  • 不能禁用投票扩展,
  • 它们是强制性的:所有发送的 precommit 消息都必须附带扩展。 不过,应用程序可以提供长度为 0 的 扩展。
必须始终设置为未来的某个高度、0,或者之前已设置的同一高度。 一旦链高度达到所设置的值,就不能再将其更改为其他值。
BlockParams.MaxBytes
完整 Protobuf 编码区块的最大大小。 该限制由共识算法强制执行。 这意味着交易的最大大小为 MaxBytes,减去区块中预计的 头部、验证者集合以及任何已包含证据所占的大小。 应用程序应意识到,诚实的验证者 可能 会生成并 广播大小最多达到所配置 MaxBytes 的区块。 因此,节点采用的共识 超时参数 应进行配置,以考虑最坏情况下将一个大小为 MaxBytes 的完整区块传递给所有验证者所需的 延迟。 如果应用程序希望完全控制区块大小, 它可以通过在应用程序层强制执行一个字节限制来实现。 这个应用程序内部限制会被 PrepareProposal 用来约束其返回的 交易总大小,也会被 ProcessProposal 用来拒绝任何接收到的、 交易总大小超过该限制的区块。 在这种情况下,应用程序可以将 MaxBytes 设置为 -1。 如果应用程序将该值设为 -1,共识将会:
  • 认为实际要强制执行的值是 100 MB
  • 在调用 PrepareProposal 时提供 mempool 中的 所有 交易
必须满足 MaxBytes == -1 或 0 < MaxBytes <= 100 MB。
请注意,BlockParams.MaxBytes 共识参数的默认值 会将大小最多为 21 MB 的区块视为有效。 如果应用程序的使用场景不需要这么大的区块, 或者尚未评估传播此大小区块带来的影响(尤其是带宽消耗和区块延迟), 强烈建议下调这个默认值。
BlockParams.MaxGas
提议区块中允许的 GasWanted 总和上限。 这 不会 由共识算法强制执行。 是否执行该限制由应用程序负责(即如果包含交易后超过该限制, 它们应返回非零代码)。CometBFT 使用它来限制 提议区块中包含的交易。 必须满足 MaxGas >= -1。 如果 MaxGas == -1,则不强制执行任何限制。
EvidenceParams.MaxAgeDuration
这是证据按时间单位计算的最大年龄。 该限制由共识算法强制执行。 如果某个区块包含比这更早的证据(并且该证据创建于 MaxAgeNumBlocks 之前),该区块将被拒绝(验证者不会为其 投票)。 必须满足 MaxAgeDuration > 0。
EvidenceParams.MaxAgeNumBlocks
这是证据按区块数计算的最大年龄。 该限制由共识算法强制执行。 如果某个区块包含比这更早的证据(并且该证据创建于 MaxAgeDuration 之前),该区块将被拒绝(验证者不会为其 投票)。 必须满足 MaxAgeNumBlocks > 0。
EvidenceParams.MaxBytes
这是单个区块中可提交的证据总大小上限,单位为字节。它应当明显低于区块字节数上限。 它的值不得超过区块大小减去其开销后的大小(约为 BlockParams.MaxBytes)。 必须满足 MaxBytes > 0。
ValidatorParams.PubKeyTypes
该参数限制验证者可使用的密钥类型。该参数使用 ABCI 的公钥命名,而不是 Amino 名称。
VersionParams.App
这是 ABCI 应用程序的版本。

更新共识参数

应用程序可以在 InitChain 期间设置 ConsensusParams,并在 FinalizeBlock 期间更新它们。 如果 ConsensusParams 为空,则会被忽略。每个非空字段 都会被完整应用。例如,如果更新 Block.MaxBytes,应用程序也必须同时设置其他 Block 字段(如 Block.MaxGas),即使这些字段未发生变化也是如此,否则它们会被更新为默认值。
InitChain
ResponseInitChain 包含一个 ConsensusParams 参数。 如果 ConsensusParams 为 nil,CometBFT 将使用创世文件中加载的参数。 如果 ConsensusParams 不为 nil,CometBFT 将使用它。 通过这种方式,应用程序可以确定区块链的初始共识参数。
FinalizeBlock, PrepareProposal/ProcessProposal
ResponseFinalizeBlock 接受一个 ConsensusParams 参数。 如果 ConsensusParams 为 nil,CometBFT 不会执行任何操作。 如果 ConsensusParams 不为 nil,CometBFT 将使用它。 通过这种方式,应用程序可以随着时间推移更新共识参数。 在区块 H 中返回的更新会立即对区块 H+1 生效。

Query

Query 是一个高度灵活的通用方法,可用于支持针对应用状态的多种查询。 CometBFT 使用 Query 基于 ID 和 IP 过滤新对等节点, 并通过 RPC 将 Query 暴露给用户。 请注意,对 Query 的调用不会在节点之间复制,而是查询 本地节点的状态,因此它们可能返回过期读取。对于需要共识保证的读取,请使用交易。 Query 最重要的用途是在某个高度返回应用状态的 Merkle 证明, 这些证明可用于构建高效的应用专用轻客户端。 需要注意的是,CometBFT 在正常运行时从技术上讲并不要求 Query 消息具备任何特定能力,也就是说,如果 ABCI 应用开发者不希望实现 Query 功能,可以不实现。

Query 证明

CometBFT 的区块头包含多个哈希值,每个哈希值都为某类区块链证明 提供锚点。ValidatorsHash 支持快速验证验证者集合, DataHash 支持快速验证区块中包含的交易。 AppHash 的独特之处在于它是应用特定的,并且允许提供关于应用状态的 应用专用 Merkle 证明。 有些应用将所有相关状态都保存在交易本身中 (例如 Bitcoin 及其 UTXO),而另一些应用则维护一个独立状态, 该状态是根据交易确定性计算出来的,但并不直接包含在 交易本身中(例如 Ethereum 的合约和账户)。 对于这类应用,AppHash 为验证轻客户端证明提供了高效得多的方法。 ABCI 应用可以按如下方式利用更高效的轻客户端状态证明:
  • 在 ResponseFinalizeBlock.Data 中返回确定性应用状态的 Merkle 根。 该 Merkle 根会作为 AppHash 包含在下一个区块中。
  • 在 ResponseQuery.Proof 中返回关于该应用状态的高效 Merkle 证明, 这些证明可以使用对应区块的 AppHash 进行验证。
例如,这使得应用的轻客户端能够验证应用状态中的不存在性证明, 而使用区块哈希来做这件事的效率要低得多。 某些应用(例如 Ethereum、Cosmos-SDK)具有多层级的 Merkle 树, 其中一棵树的叶子节点是其他树的根哈希。为了支持这一点,以及 Merkle 证明在总体上的多样性,ResponseQuery.Proof 具有如下最小结构:
message ProofOps {
  repeated ProofOp ops = 1
}

message ProofOp {
  string type = 1;
  bytes key   = 2;
  bytes data  = 3;
}
每个 ProofOp 都包含单棵 Merkle 树中单个键的一份证明,其类型由指定的 type 决定。 这使得 ABCI 只需通过变化 type,就能够支持多种不同类型的 Merkle 树、编码 格式以及证明(例如存在性证明和不存在性证明)。 data 包含实际编码后的证明,并按照 type 对应的方式进行编码。 在验证完整证明时,某个 ProofOp 的根哈希会作为列表中下一个 ProofOp 要验证的值。列表中最后一个 ProofOp 的根哈希应当与待验证的 AppHash 匹配。

对等节点过滤

当 CometBFT 连接到一个对等节点时,它会使用以下路径向 ABCI 应用发送两次查询, 且不附带额外数据:
  • /p2p/filter/addr/<IP:PORT>,其中 <IP:PORT> 表示连接的 IP 地址和 端口
  • p2p/filter/id/<ID>,其中 <ID> 是对等节点的节点 ID(即 对等节点 PubKey 的 pubkey.Address())
如果这些查询中的任意一个返回非零的 ABCI code,CometBFT 将拒绝 连接该对等节点。

路径

查询是针对路径发起的,并且可以选择性地包含额外数据。 预期会存在若干高层级路径,用于区分不同关注点, 例如 /p2p、/store 和 /app。目前, CometBFT 仅使用 /p2p 来过滤对等节点。对于更高级的用法,请参见 Cosmos-SDK 中 Query 的实现。

崩溃恢复

预期 CometBFT 与应用程序会一同崩溃,并且不应存在这样一种场景: 应用程序持久化的状态高度高于 CometBFT 已持久化的最新高度。 在实践中,持久化某个高度的状态由三个步骤组成,其中最后一步 是调用应用程序的 Commit 方法,这是唯一一个预期由应用程序 持久化/提交其状态的地方。 在启动时(恢复时),CometBFT 会通过 Info Connection 调用 Info 方法,以获取应用程序最新的 已提交状态。应用程序返回的信息必须与其成功完成 Commit 的最后一个区块保持一致。 在某个高度的状态被视为已持久化之前,需要执行以下三个步骤:
  • CometBFT 将区块存储到 blockstore 中
  • CometBFT 存储应用程序通过 FinalizeBlockResponse 返回的状态
  • 应用程序在 Commit 中提交其状态。
下图展示了这些事件发生的顺序,以及 CometBFT 与应用程序调用并执行的 对应 ABCI 函数:
APP:                                              Execute block                         Persist application state
                                                 /     return ResultFinalizeBlock            /
                                                /                                           /
Event: ------------- block_stored ------------ / ------------ state_stored --------------- / ----- app_persisted_state
                          |                   /                   |                       /        |
CometBFT: Decide --- Persist block -- Call FinalizeBlock - Persist results ---------- Call Commit --
            on        in the                                (txResults, validator
           Block      block store                              updates...)

由于这三个步骤不是原子性的,因此我们会根据崩溃发生前已经执行了哪些步骤 观察到不同情况 (我们假设至少已经执行了 block_stored,否则就没有任何状态被持久化, 并且该高度的操作将被完整重做):
  • block_stored:我们重放 FinalizeBlock 以及其后的步骤。
  • block_stored 和 state_stored:由于应用程序没有在 Commit 中持久化其状态,我们需要重新执行 FinalizeBlock 以获取结果,并将其与 CometBFT 在 state_stored 中保存的状态进行比较。 预期情况下这些状态应当匹配,否则 CometBFT 会 panic。
  • block_stored、state_stored、app_persisted_state:我们进入下一个高度。
基于这些事件的顺序,如果序列中的任一步骤乱序发生,CometBFT 将会 panic, 也就是说,如果:
  • 应用程序持久化的某个区块高度高于 state_stored 阶段保存的区块高度。
  • block_stored 步骤持久化的区块高度小于 state_stored 的高度。
  • 并且 state_stored 与 block_stored 所持久化区块高度之间的差值大于 1 这对应于一种场景:我们在区块存储中保存了两个区块,但从未持久化第一个 区块的状态,这种情况绝不应发生。
还有一种特殊情况:如果崩溃发生在第一个区块提交之前,也就是调用 InitChain 之后。在这种情况下,应用程序的状态仍应处于高度 0,因此 InitChain 会再次被调用。

状态同步

加入网络的新节点可以直接从创世高度加入共识,并重放全部 历史区块直到追上当前进度。然而,对于大型链来说,这可能需要相当长的 时间,通常以天或周为单位。 状态同步是一种为新节点完成引导的替代机制,节点会获取某个高度上的状态机快照 并进行恢复。根据应用程序的不同,这可能比重放区块快上数个数量级。 请注意,状态同步目前不会回填历史区块,因此节点将只保留 截断后的区块历史。建议用户从区块可用性和可审计性的角度,考虑这对整个网络的更广泛影响。 未来可能会加入此功能。 有关具体的 ABCI 调用和类型,请参见 methods 一节。

创建快照

想要支持状态同步的应用必须按固定间隔创建状态快照。具体如何实现完全由应用自行决定。一个快照由一些元数据和一组任意格式的二进制分块组成:
  • Height (uint64):创建快照时的高度。快照必须在给定高度已提交之后生成,并且不得包含任何更高高度的数据。
  • Format (uint32):任意的快照格式标识符。可用于对快照格式进行版本管理,例如将序列化方式从 Protobuf 切换为 MessagePack。应用在恢复时可以据此决定接受还是拒绝某个快照。
  • Chunks (uint32):快照中的分块数量。每个分块包含任意二进制数据,并且应小于 16 MB;10 MB 是一个不错的起始值。
  • Hash ([]byte):快照的任意哈希值。下载分块时,它用于检查各节点上的快照是否相同。
  • Metadata ([]byte):任意快照元数据,例如用于校验的分块哈希或其他必要信息。
要让一个快照在不同节点之间被视为相同,上述所有字段都必须完全一致。通过网络发送时,快照元数据消息的大小限制为 4 MB。 当新节点运行状态同步并发现快照时,CometBFT 会通过 ABCI 的 ListSnapshots 方法查询已有应用,以发现可用快照,并通过 LoadSnapshotChunk 加载二进制快照分块。应用可以自由决定如何实现以及使用哪些格式,但必须提供以下保证:
  • 一致性: 快照必须在单一且隔离的高度生成,不受并发写入影响。这可以通过使用支持带快照隔离的 ACID 事务的数据存储来实现。
  • 异步性: 创建快照可能耗时较长,因此不能阻塞链的推进,例如可以在独立线程中执行。
  • 确定性: 在相同高度、相同格式下生成的快照,在所有节点之间必须完全一致(字节级别一致),包括所有元数据。这可确保分块具有良好的可用性,并且跨节点能够正确拼接。
一种非常基础的方法是使用支持 MVCC 事务的数据存储(例如 RocksDB),在区块提交后立即启动一个事务,并创建一个新线程,将事务句柄传递给它。然后该线程可以导出所有数据项,使用例如 Protobuf 进行序列化,对字节流求哈希,将其拆分为多个分块,并将这些分块连同一些元数据一起存储到文件系统中;与此同时,区块链仍可并行应用新区块。 更高级的方法还可以包括针对单个分块与链应用哈希的增量验证、并行或批量导出、压缩等。 一段时间后应删除旧快照 - 通常只需要保留最近两个快照(以防节点在恢复最后一个快照时,该快照被删除)。

引导节点启动

可以通过将配置项 statesync.enabled = true 设置为真,对一个空节点执行状态同步。该节点还需要链的创世文件以获取基本链信息,以及用于轻客户端验证已恢复快照的配置:一组 CometBFT RPC 服务器,以及来自可信来源的受信任头部哈希及其对应高度,这些都通过 statesync 配置段提供。 启动后,节点将连接到 P2P 网络并开始发现快照。发现到的快照会通过 OfferSnapshot ABCI 方法提供给本地应用。一旦某个快照被接受,CometBFT 就会获取并应用该快照分块。所有分块成功应用后,CometBFT 会使用轻客户端将应用的 AppHash 与链上数据进行校验,然后将节点切换到正常的共识运行模式。

快照发现

当空节点加入 P2P 网络时,它会要求所有对等节点通过 ListSnapshots ABCI 调用上报快照(每个节点最多 10 个)。经过一段时间后,节点会选择最合适的快照(通常按高度、格式和提供该快照的对等节点数量优先排序),并通过 OfferSnapshot 将其提供给应用。应用可以选择多种响应方式,包括接受或拒绝它、拒绝所提供的格式、拒绝发送它的对等节点等。CometBFT 会持续发现并提供快照,直到某个快照被接受,或者应用中止流程。

快照恢复

一旦通过 OfferSnapshot 接受了某个快照,CometBFT 就会开始从任何拥有相同快照的对等节点下载分块(即元数据字段完全一致的快照)。分块会先缓存在临时目录中,然后按顺序通过 ApplySnapshotChunk 提供给应用,直到所有分块都被接受。 如何恢复快照分块完全由应用自行决定。 在恢复过程中,应用可以通过 ApplySnapshotChunk 返回关于如何继续的指令。通常是接受当前分块并等待下一个分块,但也可以要求重新获取分块(当前分块或任意数量的前序分块)、封禁 P2P 对等节点、拒绝或重试快照,以及其他多种响应方式,详见 ABCI 参考文档。 如果 CometBFT 在一段时间后仍无法获取某个分块,它会拒绝该快照,并通过 OfferSnapshot 尝试其他快照 - 应用可以自行决定是否支持重新开始恢复,或者直接以错误中止。

快照验证

当所有分块都已被接受后,CometBFT 会发起一次 Info ABCI 调用以获取 LastBlockAppHash。然后将其与通过轻客户端获取并验证的链上受信任应用哈希进行比较。CometBFT 还会检查 LastBlockHeight 是否与快照高度一致。 此验证可确保应用在加入网络前是有效的。不过,快照恢复可能需要较长时间才能完成,因此应用可能希望在恢复过程中采用额外的验证机制,以便尽早发现故障。例如,可以对每个分块相对于应用哈希进行增量验证(使用随附的 Merkle 证明),使用校验和来防止磁盘或网络导致的数据损坏等。不过,需要注意的是,唯一可信的信息只有应用哈希,其他所有快照元数据都可能被攻击者伪造。 应用还可能需要考虑状态同步中的拒绝服务攻击向量,即攻击者提供无效或有害的快照,阻止节点加入网络。应用可以通过要求 CometBFT 封禁对等节点来应对。作为最后手段,节点运营者可以使用 P2P 配置选项,将能够提供有效快照的一组受信任对等节点加入白名单。

切换到共识

当所有快照都恢复完成后,CometBFT 会从创世文件和轻客户端 RPC 服务器中收集引导节点启动所需的附加信息(例如链 ID、共识参数、验证者集合和区块头)。它还会调用 Info 来验证以下内容:
  • 已交付给应用的快照中的应用哈希,与下一高度区块中存储的 apphash 一致
  • 应用在 ResponseInfo 中返回的版本,与当前高度区块头中的版本一致
一旦状态机恢复完成且 CometBFT 收集到了这些附加信息,它就会切换到共识。自 ABCI 2.0 起,CometBFT 会确保满足切换所需的必要条件 RFC-100。 从应用的角度看,这些操作是透明的,除非应用刚刚升级到 ABCI 2.0。 在这种情况下,应用需要正确配置,并了解在何时提供投票扩展方面的一些约束。更多细节可参见下文对应章节。 一旦节点切换到共识模式,它的运行方式就与其他任何节点一样,只是其区块历史会在已恢复快照的高度处截断。

切换到 ABCI 2.0 所需的应用配置

引入投票扩展需要修改应用的配置。 首先,切换到支持投票扩展的 CometBFT 版本需要一次协调升级。 有关升级路径的详细说明,请参阅 RFC-100 中对应的章节。 新增了一个 共识参数:VoteExtensionsEnableHeight。 该参数表示共识继续推进时要求启用投票扩展的高度,默认值为 0(即不启用投票扩展)。 一条链可以通过以下任一方式启用投票扩展:
  • 在创世时设置 VoteExtensionsEnableHeight,例如令其等于 InitialHeight
  • 或通过应用逻辑修改 ConsensusParam,以配置 VoteExtensionsEnableHeight
一旦在高度 hu 完成了到 ABCI 2.0 的(协调)升级,VoteExtensionsEnableHeight 的值 MAY 被设置为某个高度 he,且该高度 MUST 高于链的当前高度。因此,he 的最早取值是 hu + 1。 当节点达到配置的高度后, 对于所有高度 h ≥ he,共识算法都会将任何未携带已签名投票扩展数据的预提交消息判定为无效。 如果应用需要,允许使用长度为 0 的投票扩展,但它 MUST 已签名, 并且必须出现在预提交消息中。 同样地,对于所有高度 h < he,任何确实携带投票扩展的预提交消息 也会被拒绝并视为格式错误。 高度 he 比较特殊,因为对 PrepareProposal 的调用 MUST NOT 包含投票扩展数据,但该高度上的所有预提交投票 MUST 携带投票扩展, 即使该扩展为 nil。 高度 he + 1 是第一个要求 PrepareProposal MUST 携带投票扩展数据, 且该高度上的所有预提交投票 MUST 携带投票扩展的高度。 由此可推,CometBFT 将决定基于链的当前高度,需要存储哪些数据,以及成功执行操作时需要哪些数据。

Formal Requirements

Consensus Connection Requirements

This section specifies what CometBFT expects from the Application. It is structured as a set of formal requirements that can be used for testing and verification of the Application’s logic. Let p and q be two correct processes. Let rp (resp. rq) be a round of height h where p (resp. q) is the proposer. Let sp,h-1 be p’s Application’s state committed for height h-1. Let vp (resp. vq) be the block that p’s (resp. q’s) CometBFT passes on to the Application via RequestPrepareProposal as proposer of round rp (resp rq), height h, also known as the raw proposal. Let up (resp. uq) the possibly modified block p’s (resp. q’s) Application returns via ResponsePrepareProposal to CometBFT, also known as the prepared proposal. Process p’s prepared proposal can differ in two different rounds where p is the proposer.
  • Requirement 1 [PrepareProposal, timeliness]: If p’s Application fully executes prepared blocks in PrepareProposal and the network is in a synchronous period while processes p and q are in rp, then the value of TimeoutPropose at q must be such that q’s propose timer does not time out (which would result in q prevoting nil in rp).
Full execution of blocks at PrepareProposal time stands on CometBFT’s critical path. Thus, Requirement 1 ensures the Application or operator will set a value for TimeoutPropose such that the time it takes to fully execute blocks in PrepareProposal does not interfere with CometBFT’s propose timer. Note that the violation of Requirement 1 may lead to further rounds, but will not compromise liveness because even though TimeoutPropose is used as the initial value for proposal timeouts, CometBFT will be dynamically adjust these timeouts such that they will eventually be enough for completing PrepareProposal.
  • Requirement 2 [PrepareProposal, tx-size]: When p’s Application calls ResponsePrepareProposal, the total size in bytes of the transactions returned does not exceed RequestPrepareProposal.max_tx_bytes.
Busy blockchains might seek to gain full visibility into transactions in CometBFT’s mempool, rather than having visibility only on a subset of those transactions that fit in a block. The application can do so by setting ConsensusParams.Block.MaxBytes to -1. This instructs CometBFT (a) to enforce the maximum possible value for MaxBytes (100 MB) at CometBFT level, and (b) to provide all transactions in the mempool when calling RequestPrepareProposal. Under these settings, the aggregated size of all transactions may exceed RequestPrepareProposal.max_tx_bytes. Hence, Requirement 2 ensures that the size in bytes of the transaction list returned by the application will never cause the resulting block to go beyond its byte size limit.
  • Requirement 3 [PrepareProposal, ProcessProposal, coherence]: For any two correct processes p and q, if q’s CometBFT calls RequestProcessProposal on up, q’s Application returns Accept in ResponseProcessProposal.
Requirement 3 makes sure that blocks proposed by correct processes always pass the correct receiving process’s ProcessProposal check. On the other hand, if there is a deterministic bug in PrepareProposal or ProcessProposal (or in both), strictly speaking, this makes all processes that hit the bug byzantine. This is a problem in practice, as very often validators are running the Application from the same codebase, so potentially all would likely hit the bug at the same time. This would result in most (or all) processes prevoting nil, with the serious consequences on CometBFT’s liveness that this entails. Due to its criticality, Requirement 3 is a target for extensive testing and automated verification.
  • Requirement 4 [ProcessProposal, determinism-1]: ProcessProposal is a (deterministic) function of the current state and the block that is about to be applied. In other words, for any correct process p, and any arbitrary block u, if p’s CometBFT calls RequestProcessProposal on u at height h, then p’s Application’s acceptance or rejection exclusively depends on u and sp,h-1.
  • Requirement 5 [ProcessProposal, determinism-2]: For any two correct processes p and q, and any arbitrary block u, if p’s (resp. q’s) CometBFT calls RequestProcessProposal on u at height h, then p’s Application accepts u if and only if q’s Application accepts u. Note that this requirement follows from Requirement 4 and the Agreement property of consensus.
Requirements 4 and 5 ensure that all correct processes will react in the same way to a proposed block, even if the proposer is Byzantine. However, ProcessProposal may contain a bug that renders the acceptance or rejection of the block non-deterministic, and therefore prevents processes hitting the bug from fulfilling Requirements 4 or 5 (effectively making those processes Byzantine). In such a scenario, CometBFT’s liveness cannot be guaranteed. Again, this is a problem in practice if most validators are running the same software, as they are likely to hit the bug at the same point. There is currently no clear solution to help with this situation, so the Application designers/implementors must proceed very carefully with the logic/implementation of ProcessProposal. As a general rule ProcessProposal SHOULD always accept the block. According to the Tendermint consensus algorithm, currently adopted in CometBFT, a correct process can broadcast at most one precommit message in round r, height h. Since, as stated in the Methods section, ResponseExtendVote is only called when the consensus algorithm is about to broadcast a non-nil precommit message, a correct process can only produce one vote extension in round r, height h. Let erp be the vote extension that the Application of a correct process p returns via ResponseExtendVote in round r, height h. Let wrp be the proposed block that p’s CometBFT passes to the Application via RequestExtendVote in round r, height h.
  • Requirement 6 [ExtendVote, VerifyVoteExtension, coherence]: For any two different correct processes p and q, if q receives erp from p in height h, q’s Application returns Accept in ResponseVerifyVoteExtension.
Requirement 6 constrains the creation and handling of vote extensions in a similar way as Requirement 3 constrains the creation and handling of proposed blocks. Requirement 6 ensures that extensions created by correct processes always pass the VerifyVoteExtension checks performed by correct processes receiving those extensions. However, if there is a (deterministic) bug in ExtendVote or VerifyVoteExtension (or in both), we will face the same liveness issues as described for Requirement 5, as Precommit messages with invalid vote extensions will be discarded.
  • Requirement 7 [VerifyVoteExtension, determinism-1]: VerifyVoteExtension is a (deterministic) function of the current state, the vote extension received, and the prepared proposal that the extension refers to. In other words, for any correct process p, and any arbitrary vote extension e, and any arbitrary block w, if p’s (resp. q’s) CometBFT calls RequestVerifyVoteExtension on e and w at height h, then p’s Application’s acceptance or rejection exclusively depends on e, w and sp,h-1.
  • Requirement 8 [VerifyVoteExtension, determinism-2]: For any two correct processes p and q, and any arbitrary vote extension e, and any arbitrary block w, if p’s (resp. q’s) CometBFT calls RequestVerifyVoteExtension on e and w at height h, then p’s Application accepts e if and only if q’s Application accepts e. Note that this requirement follows from Requirement 7 and the Agreement property of consensus.
Requirements 7 and 8 ensure that the validation of vote extensions will be deterministic at all correct processes. Requirements 7 and 8 protect against arbitrary vote extension data from Byzantine processes, in a similar way as Requirements 4 and 5 protect against arbitrary proposed blocks. Requirements 7 and 8 can be violated by a bug inducing non-determinism in VerifyVoteExtension. In this case liveness can be compromised. Extra care should be put in the implementation of ExtendVote and VerifyVoteExtension. As a general rule, VerifyVoteExtension SHOULD always accept the vote extension.
  • Requirement 9 [all, no-side-effects]: p’s calls to RequestPrepareProposal, RequestProcessProposal, RequestExtendVote, and RequestVerifyVoteExtension at height h do not modify sp,h-1.
  • Requirement 10 [ExtendVote, FinalizeBlock, non-dependency]: for any correct process p, and any vote extension e that p received at height h, the computation of sp,h does not depend on e.
The call to correct process p’s RequestFinalizeBlock at height h, with block vp,h passed as parameter, creates state sp,h. Additionally, p’s FinalizeBlock creates a set of transaction results Tp,h.
  • Requirement 11 [FinalizeBlock, determinism-1]: For any correct process p, sp,h exclusively depends on sp,h-1 and vp,h.
  • Requirement 12 [FinalizeBlock, determinism-2]: For any correct process p, the contents of Tp,h exclusively depend on sp,h-1 and vp,h.
Note that Requirements 11 and 12, combined with the Agreement property of consensus ensure state machine replication, i.e., the Application state evolves consistently at all correct processes. Also, notice that neither PrepareProposal nor ExtendVote have determinism-related requirements associated. Indeed, PrepareProposal is not required to be deterministic:
  • up may depend on vp and sp,h-1, but may also depend on other values or operations.
  • vp = vq ⇏ up = uq.
Likewise, ExtendVote can also be non-deterministic:
  • erp may depend on wrp and sp,h-1, but may also depend on other values or operations.
  • wrp = wrq ⇏ erp = erq

Mempool Connection Requirements

Let CheckTxCodestx,p,h denote the set of result codes returned by p’s Application, via ResponseCheckTx, to successive calls to RequestCheckTx occurring while the Application is at height h and having transaction tx as parameter. CheckTxCodestx,p,h is a set since p’s Application may return different result codes during height h. If CheckTxCodestx,p,h is a singleton set, i.e. the Application always returned the same result code in ResponseCheckTx while at height h, we define CheckTxCodetx,p,h as the singleton value of CheckTxCodestx,p,h. If CheckTxCodestx,p,h is not a singleton set, CheckTxCodetx,p,h is undefined. Let predicate OK(CheckTxCodetx,p,h) denote whether CheckTxCodetx,p,h is SUCCESS.
  • Requirement 13 [CheckTx, eventual non-oscillation]: For any transaction tx, there exists a boolean value b, and a height hstable such that, for any correct process p, CheckTxCodetx,p,h is defined, and OK(CheckTxCodetx,p,h) = b for any height h ≥ hstable.
Requirement 13 ensures that a transaction will eventually stop oscillating between CheckTx success and failure if it stays in p’s mempool for long enough. This condition on the Application’s behavior allows the mempool to ensure that a transaction will leave the mempool of all full nodes, either because it is expunged everywhere due to failing CheckTx calls, or because it stays valid long enough to be gossipped, proposed and decided. Although Requirement 13 defines a global hstable, application developers can consider such stabilization height as local to process p (hp,stable), without loss for generality. In contrast, the value of b MUST be the same across all processes.

Connection State

CometBFT maintains four concurrent ABCI++ connections, namely Consensus Connection, Mempool Connection, Info/Query Connection, and Snapshot Connection. It is common for an application to maintain a distinct copy of the state for each connection, which are synchronized upon Commit calls.

Concurrency

In principle, each of the four ABCI++ connections operates concurrently with one another. This means applications need to ensure access to state is thread safe. Both the default in-process ABCI client and the default Go ABCI server use a global lock to guard the handling of events across all connections, so they are not concurrent at all. This means whether your app is compiled in-process with CometBFT using the NewLocalClient, or run out-of-process using the SocketServer, ABCI messages from all connections are received in sequence, one at a time. The existence of this global mutex means Go application developers can get thread safety for application state by routing all reads and writes through the ABCI system. Thus it may be unsafe to expose application state directly to an RPC interface, and unless explicit measures are taken, all queries should be routed through the ABCI Query method.

FinalizeBlock

When the consensus algorithm decides on a block, CometBFT uses FinalizeBlock to send the decided block’s data to the Application, which uses it to transition its state, but MUST NOT persist it; persisting MUST be done during Commit. The Application must remember the latest height from which it has run a successful Commit so that it can tell CometBFT where to pick up from when it recovers from a crash. See information on the Handshake here.

Commit

The Application should persist its state during Commit, before returning from it. Before invoking Commit, CometBFT locks the mempool and flushes the mempool connection. This ensures that no new messages will be received on the mempool connection during this processing step, providing an opportunity to safely update all four connection states to the latest committed state at the same time. When Commit returns, CometBFT unlocks the mempool. WARNING: if the ABCI app logic processing the Commit message sends a /broadcast_tx_sync or /broadcast_tx and waits for the response before proceeding, it will deadlock. Executing broadcast_tx calls involves acquiring the mempool lock that CometBFT holds during the Commit call. Synchronous mempool-related calls must be avoided as part of the sequential logic of the Commit function.

Candidate States

CometBFT calls PrepareProposal when it is about to send a proposed block to the network. Likewise, CometBFT calls ProcessProposal upon reception of a proposed block from the network. The proposed block’s data that is disclosed to the Application by these two methods is the following:
  • the transaction list
  • the LastCommit referring to the previous block
  • the block header’s hash (except in PrepareProposal, where it is not known yet)
  • list of validators that misbehaved
  • the block’s timestamp
  • NextValidatorsHash
  • Proposer address
The Application may decide to immediately execute the given block (i.e., upon PrepareProposal or ProcessProposal). There are two main reasons why the Application may want to do this:
  • Avoiding invalid transactions in blocks. In order to be sure that the block does not contain any invalid transaction, there may be no way other than fully executing the transactions in the block as though it was the decided block.
  • Quick FinalizeBlock execution. Upon reception of the decided block via FinalizeBlock, if that same block was executed upon PrepareProposal or ProcessProposal and the resulting state was kept in memory, the Application can simply apply that state (faster) to the main state, rather than reexecuting the decided block (slower).
PrepareProposal/ProcessProposal can be called many times for a given height. Moreover, it is not possible to accurately predict which of the blocks proposed in a height will be decided, being delivered to the Application in that height’s FinalizeBlock. Therefore, the state resulting from executing a proposed block, denoted a candidate state, should be kept in memory as a possible final state for that height. When FinalizeBlock is called, the Application should check if the decided block corresponds to one of its candidate states; if so, it will apply it as its ExecuteTxState (see Consensus Connection below), which will be persisted during the upcoming Commit call. Under adverse conditions (e.g., network instability), the consensus algorithm might take many rounds. In this case, potentially many proposed blocks will be disclosed to the Application for a given height. By the nature of Tendermint consensus algorithm, currently adopted in CometBFT, the number of proposed blocks received by the Application for a particular height cannot be bound, so Application developers must act with care and use mechanisms to bound memory usage. As a general rule, the Application should be ready to discard candidate states before FinalizeBlock, even if one of them might end up corresponding to the decided block and thus have to be reexecuted upon FinalizeBlock.

States and ABCI++ Connections

Consensus Connection

The Consensus Connection should maintain an ExecuteTxState — the working state for block execution. It should be updated by the call to FinalizeBlock during block execution and committed to disk as the “latest committed state” during Commit. Execution of a proposed block (via PrepareProposal/ProcessProposal) must not update the ExecuteTxState, but rather be kept as a separate candidate state until FinalizeBlock confirms which of the candidate states (if any) can be used to update ExecuteTxState.

Mempool Connection

The mempool Connection maintains CheckTxState. CometBFT sequentially processes an incoming transaction (via RPC from client or P2P from the gossip layer) against CheckTxState. If the processing does not return any error, the transaction is accepted into the mempool and CometBFT starts gossipping it. CheckTxState should be reset to the latest committed state at the end of every Commit. During the execution of a consensus instance, the CheckTxState may be updated concurrently with the ExecuteTxState, as messages may be sent concurrently on the Consensus and Mempool connections. At the end of the consensus instance, as described above, CometBFT locks the mempool and flushes the mempool connection before calling Commit. This ensures that all pending CheckTx calls are responded to and no new ones can begin. After the Commit call returns, while still holding the mempool lock, CheckTx is run again on all transactions that remain in the node’s local mempool after filtering those included in the block. Parameter Type in RequestCheckTx indicates whether an incoming transaction is new (CheckTxType_New), or a recheck (CheckTxType_Recheck). Finally, after re-checking transactions in the mempool, CometBFT will unlock the mempool connection. New transactions are once again able to be processed through CheckTx. Note that CheckTx is just a weak filter to keep invalid transactions out of the mempool and, ultimately, ouf of the blockchain. Since the transaction cannot be guaranteed to be checked against the exact same state as it will be executed as part of a (potential) decided block, CheckTx shouldn’t check everything that affects the transaction’s validity, in particular those checks whose validity may depend on transaction ordering. CheckTx is weak because a Byzantine node need not care about CheckTx; it can propose a block full of invalid transactions if it wants. The mechanism ABCI++ has in place for dealing with such behavior is ProcessProposal.
Replay Protection
It is possible for old transactions to be sent again to the Application. This is typically undesirable for all transactions, except for a generally small subset of them which are idempotent. The mempool has a mechanism to prevent duplicated transactions from being processed. This mechanism is nevertheless best-effort (currently based on the indexer) and does not provide any guarantee of non duplication. It is thus up to the Application to implement an application-specific replay protection mechanism with strong guarantees as part of the logic in CheckTx.

Info/Query Connection

The Info (or Query) Connection should maintain a QueryState. This connection has two purposes: 1) having the application answer the queries CometBFT receives from users (see section Query), and 2) synchronizing CometBFT and the Application at start up time (see Crash Recovery) or after state sync (see State Sync). QueryState is a read-only copy of ExecuteTxState as it was after the last Commit, i.e. after the full block has been processed and the state committed to disk.

Snapshot Connection

The Snapshot Connection is used to serve state sync snapshots for other nodes and/or restore state sync snapshots to a local node being bootstrapped. Snapshot management is optional: an Application may choose not to implement it. For more information, see Section State Sync.

Transaction Results

The Application is expected to return a list of ExecTxResult in ResponseFinalizeBlock. The list of transaction results MUST respect the same order as the list of transactions delivered via RequestFinalizeBlock. This section discusses the fields inside this structure, along with the fields in ResponseCheckTx, whose semantics are similar. The Info and Log fields are non-deterministic values for debugging/convenience purposes. CometBFT logs them but they are otherwise ignored.

Gas

Ethereum introduced the notion of gas as an abstract representation of the cost of the resources consumed by nodes when processing a transaction. Every operation in the Ethereum Virtual Machine uses some amount of gas. Gas has a market-variable price based on which miners can accept or reject to execute a particular operation. Users propose a maximum amount of gas for their transaction; if the transaction uses less, they get the difference credited back. CometBFT adopts a similar abstraction, though uses it only optionally and weakly, allowing applications to define their own sense of the cost of execution. In CometBFT, the ConsensusParams.Block.MaxGas limits the amount of total gas that can be used by all transactions in a block. The default value is -1, which means the block gas limit is not enforced, or that the concept of gas is meaningless. Responses contain a GasWanted and GasUsed field. The former is the maximum amount of gas the sender of a transaction is willing to use, and the latter is how much it actually used. Applications should enforce that GasUsed <= GasWanted — i.e. transaction execution or validation should fail before it can use more resources than it requested. When MaxGas > -1, CometBFT enforces the following rules:
  • GasWanted <= MaxGas for every transaction in the mempool
  • (sum of GasWanted in a block) <= MaxGas when proposing a block
If MaxGas == -1, no rules about gas are enforced. In v0.34.x and earlier versions, CometBFT does not enforce anything about Gas in consensus, only in the mempool. This means it does not guarantee that committed blocks satisfy these rules. It is the application’s responsibility to return non-zero response codes when gas limits are exceeded when executing the transactions of a block. Since the introduction of PrepareProposal and ProcessProposal in v.0.37.x, it is now possible for the Application to enforce that all blocks proposed (and voted for) in consensus — and thus all blocks decided — respect the MaxGas limits described above. Since the Application should enforce that GasUsed <= GasWanted when executing a transaction, and it can use PrepareProposal and ProcessProposal to enforce that (sum of GasWanted in a block) <= MaxGas in all proposed or prevoted blocks, we have:
  • (sum of GasUsed in a block) <= MaxGas for every block
The GasUsed field is ignored by CometBFT.

Specifics of ResponseCheckTx

If Code != 0, it will be rejected from the mempool and hence not broadcasted to other peers and not included in a proposal block. Data contains the result of the CheckTx transaction execution, if any. It does not need to be deterministic since, given a transaction, nodes’ Applications might have a different CheckTxState values when they receive it and check their validity via CheckTx. CometBFT ignores this value in ResponseCheckTx. From v0.34.x on, there is a Priority field in ResponseCheckTx that can be used to explicitly prioritize transactions in the mempool for inclusion in a block proposal.

Specifics of ExecTxResult

FinalizeBlock is the workhorse of the blockchain. CometBFT delivers the decided block, including the list of all its transactions synchronously to the Application. The block delivered (and thus the transaction order) is the same at all correct nodes as guaranteed by the Agreement property of consensus. The Data field in ExecTxResult contains an array of bytes with the transaction result. It must be deterministic (i.e., the same value must be returned at all nodes), but it can contain arbitrary data. Likewise, the value of Code must be deterministic. If Code != 0, the transaction will be marked invalid, though it is still included in the block. Invalid transactions are not indexed, as they are considered analogous to those that failed CheckTx. Both the Code and Data are included in a structure that is hashed into the LastResultsHash of the block header in the next height. Events include any events for the execution, which CometBFT will use to index the transaction by. This allows transactions to be queried according to what events took place during their execution.

Updating the Validator Set

The application may set the validator set during InitChain, and may update it during FinalizeBlock. In both cases, a structure of type ValidatorUpdate is returned. The InitChain method, used to initialize the Application, can return a list of validators. If the list is empty, CometBFT will use the validators loaded from the genesis file. If the list returned by InitChain is not empty, CometBFT will use its contents as the validator set. This way the application can set the initial validator set for the blockchain. Applications must ensure that a single set of validator updates does not contain duplicates, i.e. a given public key can only appear once within a given update. If an update includes duplicates, the block execution will fail irrecoverably. Structure ValidatorUpdate contains a public key, which is used to identify the validator: The public key currently supports three types:
  • ed25519
  • secp256k1
  • bls12381
Structure ValidatorUpdate also contains an ìnt64 field denoting the validator’s new power. Applications must ensure that ValidatorUpdate structures abide by the following rules:
  • power must be non-negative
  • if power is set to 0, the validator must be in the validator set; it will be removed from the set
  • if power is greater than 0:
    • if the validator is not in the validator set, it will be added to the set with the given power
    • if the validator is in the validator set, its power will be adjusted to the given power
  • the total power of the new validator set must not exceed MaxTotalVotingPower, where MaxTotalVotingPower = MaxInt64 / 8
Note the updates returned after processing the block at height H will only take effect at block H+2 (see Section Methods).

Consensus Parameters

ConsensusParams are global parameters that apply to all validators in a blockchain. They enforce certain limits in the blockchain, like the maximum size of blocks, amount of gas used in a block, and the maximum acceptable age of evidence. They can be set in InitChain, and updated in FinalizeBlock. These parameters are deterministically set and/or updated by the Application, so all full nodes have the same value at a given height.

List of Parameters

These are the current consensus parameters (as of v0.38.x):
  1. ABCIParams.VoteExtensionsEnableHeight
  2. BlockParams.MaxBytes
  3. BlockParams.MaxGas
  4. EvidenceParams.MaxAgeDuration
  5. EvidenceParams.MaxAgeNumBlocks
  6. EvidenceParams.MaxBytes
  7. ValidatorParams.PubKeyTypes
  8. VersionParams.App

ABCIParams.VoteExtensionsEnableHeight

This parameter is either 0 or a positive height at which vote extensions become mandatory. If the value is zero (which is the default), vote extensions are not expected. Otherwise, at all heights greater than the configured height H vote extensions must be present (even if empty). When the configured height H is reached, PrepareProposal will not include vote extensions yet, but ExtendVote and VerifyVoteExtension will be called. Then, when reaching height H+1, PrepareProposal will include the vote extensions from height H. For all heights after H
  • vote extensions cannot be disabled,
  • they are mandatory: all precommit messages sent MUST have an extension attached. Nevertheless, the application MAY provide 0-length extensions.
Must always be set to a future height, 0, or the same height that was previously set. Once the chain’s height reaches the value set, it cannot be changed to a different value.
BlockParams.MaxBytes
The maximum size of a complete Protobuf encoded block. This is enforced by the consensus algorithm. This implies a maximum transaction size that is MaxBytes, less the expected size of the header, the validator set, and any included evidence in the block. The Application should be aware that honest validators may produce and broadcast blocks with up to the configured MaxBytes size. As a result, the consensus timeout parameters adopted by nodes should be configured so as to account for the worst-case latency for the delivery of a full block with MaxBytes size to all validators. If the Application wants full control over the size of blocks, it can do so by enforcing a byte limit set up at the Application level. This Application-internal limit is used by PrepareProposal to bound the total size of transactions it returns, and by ProcessProposal to reject any received block whose total transaction size is bigger than the enforced limit. In such case, the Application MAY set MaxBytes to -1. If the Application sets value -1, consensus will:
  • consider that the actual value to enforce is 100 MB
  • will provide all transactions in the mempool in calls to PrepareProposal
Must have MaxBytes == -1 OR 0 < MaxBytes <= 100 MB.
Bear in mind that the default value for the BlockParams.MaxBytes consensus parameter accepts as valid blocks with size up to 21 MB. If the Application’s use case does not need blocks of that size, or if the impact (specially on bandwidth consumption and block latency) of propagating blocks of that size was not evaluated, it is strongly recommended to wind down this default value.
BlockParams.MaxGas
The maximum of the sum of GasWanted that will be allowed in a proposed block. This is not enforced by the consensus algorithm. It is left to the Application to enforce (ie. if transactions are included past the limit, they should return non-zero codes). It is used by CometBFT to limit the transactions included in a proposed block. Must have MaxGas >= -1. If MaxGas == -1, no limit is enforced.
EvidenceParams.MaxAgeDuration
This is the maximum age of evidence in time units. This is enforced by the consensus algorithm. If a block includes evidence older than this (AND the evidence was created more than MaxAgeNumBlocks ago), the block will be rejected (validators won’t vote for it). Must have MaxAgeDuration > 0.
EvidenceParams.MaxAgeNumBlocks
This is the maximum age of evidence in blocks. This is enforced by the consensus algorithm. If a block includes evidence older than this (AND the evidence was created more than MaxAgeDuration ago), the block will be rejected (validators won’t vote for it). Must have MaxAgeNumBlocks > 0.
EvidenceParams.MaxBytes
This is the maximum size of total evidence in bytes that can be committed to a single block. It should fall comfortably under the max block bytes. Its value must not exceed the size of a block minus its overhead ( ~ BlockParams.MaxBytes). Must have MaxBytes > 0.
ValidatorParams.PubKeyTypes
The parameter restricts the type of keys validators can use. The parameter uses ABCI pubkey naming, not Amino names.
VersionParams.App
This is the version of the ABCI application.

Updating Consensus Parameters

The application may set the ConsensusParams during InitChain, and update them during FinalizeBlock. If the ConsensusParams is empty, it will be ignored. Each field that is not empty will be applied in full. For instance, if updating the Block.MaxBytes, applications must also set the other Block fields (like Block.MaxGas), even if they are unchanged, as they will otherwise cause the value to be updated to the default.
InitChain
ResponseInitChain includes a ConsensusParams parameter. If ConsensusParams is nil, CometBFT will use the params loaded in the genesis file. If ConsensusParams is not nil, CometBFT will use it. This way the application can determine the initial consensus parameters for the blockchain.
FinalizeBlock, PrepareProposal/ProcessProposal
ResponseFinalizeBlock accepts a ConsensusParams parameter. If ConsensusParams is nil, CometBFT will do nothing. If ConsensusParams is not nil, CometBFT will use it. This way the application can update the consensus parameters over time. The updates returned in block H will take effect right away for block H+1.

Query

Query is a generic method with lots of flexibility to enable diverse sets of queries on application state. CometBFT makes use of Query to filter new peers based on ID and IP, and exposes Query to the user over RPC. Note that calls to Query are not replicated across nodes, but rather query the local node’s state - hence they may return stale reads. For reads that require consensus, use a transaction. The most important use of Query is to return Merkle proofs of the application state at some height that can be used for efficient application-specific light-clients. Note CometBFT has technically no requirements from the Query message for normal operation - that is, the ABCI app developer need not implement Query functionality if they do not wish to.

Query Proofs

The CometBFT block header includes a number of hashes, each providing an anchor for some type of proof about the blockchain. The ValidatorsHash enables quick verification of the validator set, the DataHash gives quick verification of the transactions included in the block. The AppHash is unique in that it is application specific, and allows for application-specific Merkle proofs about the state of the application. While some applications keep all relevant state in the transactions themselves (like Bitcoin and its UTXOs), others maintain a separated state that is computed deterministically from transactions, but is not contained directly in the transactions themselves (like Ethereum contracts and accounts). For such applications, the AppHash provides a much more efficient way to verify light-client proofs. ABCI applications can take advantage of more efficient light-client proofs for their state as follows:
  • return the Merkle root of the deterministic application state in ResponseFinalizeBlock.Data. This Merkle root will be included as the AppHash in the next block.
  • return efficient Merkle proofs about that application state in ResponseQuery.Proof that can be verified using the AppHash of the corresponding block.
For instance, this allows an application’s light-client to verify proofs of absence in the application state, something which is much less efficient to do using the block hash. Some applications (eg. Ethereum, Cosmos-SDK) have multiple “levels” of Merkle trees, where the leaves of one tree are the root hashes of others. To support this, and the general variability in Merkle proofs, the ResponseQuery.Proof has some minimal structure:
message ProofOps {
  repeated ProofOp ops = 1
}

message ProofOp {
  string type = 1;
  bytes key   = 2;
  bytes data  = 3;
}
Each ProofOp contains a proof for a single key in a single Merkle tree, of the specified type. This allows ABCI to support many different kinds of Merkle trees, encoding formats, and proofs (eg. of presence and absence) just by varying the type. The data contains the actual encoded proof, encoded according to the type. When verifying the full proof, the root hash for one ProofOp is the value being verified for the next ProofOp in the list. The root hash of the final ProofOp in the list should match the AppHash being verified against.

Peer Filtering

When CometBFT connects to a peer, it sends two queries to the ABCI application using the following paths, with no additional data:
  • /p2p/filter/addr/<IP:PORT>, where <IP:PORT> denote the IP address and the port of the connection
  • p2p/filter/id/<ID>, where <ID> is the peer node ID (ie. the pubkey.Address() for the peer’s PubKey)
If either of these queries return a non-zero ABCI code, CometBFT will refuse to connect to the peer.

Paths

Queries are directed at paths, and may optionally include additional data. The expectation is for there to be some number of high level paths differentiating concerns, like /p2p, /store, and /app. Currently, CometBFT only uses /p2p, for filtering peers. For more advanced use, see the implementation of Query in the Cosmos-SDK.

Crash Recovery

CometBFT and the application are expected to crash together and there should not exist a scenario where the application has persisted state of a height greater than the latest height persisted by CometBFT. In practice, persisting the state of a height consists of three steps, the last of which is the call to the application’s Commit method, the only place where the application is expected to persist/commit its state. On startup (upon recovery), CometBFT calls the Info method on the Info Connection to get the latest committed state of the app. The app MUST return information consistent with the last block for which it successfully completed Commit. The three steps performed before the state of a height is considered persisted are:
  • The block is stored by CometBFT in the blockstore
  • CometBFT has stored the state returned by the application through FinalizeBlockResponse
  • The application has committed its state within Commit.
The following diagram depicts the order in which these events happen, and the corresponding ABCI functions that are called and executed by CometBFT and the application:
APP:                                              Execute block                         Persist application state
                                                 /     return ResultFinalizeBlock            /
                                                /                                           /
Event: ------------- block_stored ------------ / ------------ state_stored --------------- / ----- app_persisted_state
                          |                   /                   |                       /        |
CometBFT: Decide --- Persist block -- Call FinalizeBlock - Persist results ---------- Call Commit --
            on        in the                                (txResults, validator
           Block      block store                              updates...)

As these three steps are not atomic, we observe different cases based on which steps have been executed before the crash occurred (we assume that at least block_stored has been executed, otherwise, there is no state persisted, and the operations for this height are repeated entirely):
  • block_stored: we replay FinalizeBlock and the steps afterwards.
  • block_stored and state_stored: As the app did not persist its state within Commit, we need to re-execute FinalizeBlock to retrieve the results and compare them to the state stored by CometBFT within state_stored. The expected case is that the states will match, otherwise CometBFT panics.
  • block_stored, state_stored, app_persisted_state: we move on to the next height.
Based on the sequence of these events, CometBFT will panic if any of the steps in the sequence happen out of order, that is if:
  • The application has persisted a block at a height higher than the blocked saved during state_stored.
  • The block_stored step persisted a block at a height smaller than the state_stored
  • And the difference between the heights of the blocks persisted by state_stored and block_stored is more than 1 (this corresponds to a scenario where we stored two blocks in the block store but never persisted the state of the first block, which should never happen).
A special case is when a crash happens before the first block is committed - that is, after calling InitChain. In that case, the application’s state should still be at height 0 and thus InitChain will be called again.

State Sync

A new node joining the network can simply join consensus at the genesis height and replay all historical blocks until it is caught up. However, for large chains this can take a significant amount of time, often on the order of days or weeks. State sync is an alternative mechanism for bootstrapping a new node, where it fetches a snapshot of the state machine at a given height and restores it. Depending on the application, this can be several orders of magnitude faster than replaying blocks. Note that state sync does not currently backfill historical blocks, so the node will have a truncated block history - users are advised to consider the broader network implications of this in terms of block availability and auditability. This functionality may be added in the future. For details on the specific ABCI calls and types, see the methods section.

Taking Snapshots

Applications that want to support state syncing must take state snapshots at regular intervals. How this is accomplished is entirely up to the application. A snapshot consists of some metadata and a set of binary chunks in an arbitrary format:
  • Height (uint64): The height at which the snapshot is taken. It must be taken after the given height has been committed, and must not contain data from any later heights.
  • Format (uint32): An arbitrary snapshot format identifier. This can be used to version snapshot formats, e.g. to switch from Protobuf to MessagePack for serialization. The application can use this when restoring to choose whether to accept or reject a snapshot.
  • Chunks (uint32): The number of chunks in the snapshot. Each chunk contains arbitrary binary data, and should be less than 16 MB; 10 MB is a good starting point.
  • Hash ([]byte): An arbitrary hash of the snapshot. This is used to check whether a snapshot is the same across nodes when downloading chunks.
  • Metadata ([]byte): Arbitrary snapshot metadata, e.g. chunk hashes for verification or any other necessary info.
For a snapshot to be considered the same across nodes, all of these fields must be identical. When sent across the network, snapshot metadata messages are limited to 4 MB. When a new node is running state sync and discovering snapshots, CometBFT will query an existing application via the ABCI ListSnapshots method to discover available snapshots, and load binary snapshot chunks via LoadSnapshotChunk. The application is free to choose how to implement this and which formats to use, but must provide the following guarantees:
  • Consistent: A snapshot must be taken at a single isolated height, unaffected by concurrent writes. This can be accomplished by using a data store that supports ACID transactions with snapshot isolation.
  • Asynchronous: Taking a snapshot can be time-consuming, so it must not halt chain progress, for example by running in a separate thread.
  • Deterministic: A snapshot taken at the same height in the same format must be identical (at the byte level) across nodes, including all metadata. This ensures good availability of chunks, and that they fit together across nodes.
A very basic approach might be to use a datastore with MVCC transactions (such as RocksDB), start a transaction immediately after block commit, and spawn a new thread which is passed the transaction handle. This thread can then export all data items, serialize them using e.g. Protobuf, hash the byte stream, split it into chunks, and store the chunks in the file system along with some metadata - all while the blockchain is applying new blocks in parallel. A more advanced approach might include incremental verification of individual chunks against the chain app hash, parallel or batched exports, compression, and so on. Old snapshots should be removed after some time - generally only the last two snapshots are needed (to prevent the last one from being removed while a node is restoring it).

Bootstrapping a Node

An empty node can be state synced by setting the configuration option statesync.enabled = true. The node also needs the chain genesis file for basic chain info, and configuration for light client verification of the restored snapshot: a set of CometBFT RPC servers, and a trusted header hash and corresponding height from a trusted source, via the statesync configuration section. Once started, the node will connect to the P2P network and begin discovering snapshots. These will be offered to the local application via the OfferSnapshot ABCI method. Once a snapshot is accepted CometBFT will fetch and apply the snapshot chunks. After all chunks have been successfully applied, CometBFT verifies the app’s AppHash against the chain using the light client, then switches the node to normal consensus operation.

Snapshot Discovery

When the empty node joins the P2P network, it asks all peers to report snapshots via the ListSnapshots ABCI call (limited to 10 per node). After some time, the node picks the most suitable snapshot (generally prioritized by height, format, and number of peers), and offers it to the application via OfferSnapshot. The application can choose a number of responses, including accepting or rejecting it, rejecting the offered format, rejecting the peer who sent it, and so on. CometBFT will keep discovering and offering snapshots until one is accepted or the application aborts.

Snapshot Restoration

Once a snapshot has been accepted via OfferSnapshot, CometBFT begins downloading chunks from any peers that have the same snapshot (i.e. that have identical metadata fields). Chunks are spooled in a temporary directory, and then given to the application in sequential order via ApplySnapshotChunk until all chunks have been accepted. The method for restoring snapshot chunks is entirely up to the application. During restoration, the application can respond to ApplySnapshotChunk with instructions for how to continue. This will typically be to accept the chunk and await the next one, but it can also ask for chunks to be refetched (either the current one or any number of previous ones), P2P peers to be banned, snapshots to be rejected or retried, and a number of other responses - see the ABCI reference for details. If CometBFT fails to fetch a chunk after some time, it will reject the snapshot and try a different one via OfferSnapshot - the application can choose whether it wants to support restarting restoration, or simply abort with an error.

Snapshot Verification

Once all chunks have been accepted, CometBFT issues an Info ABCI call to retrieve the LastBlockAppHash. This is compared with the trusted app hash from the chain, retrieved and verified using the light client. CometBFT also checks that LastBlockHeight corresponds to the height of the snapshot. This verification ensures that an application is valid before joining the network. However, the snapshot restoration may take a long time to complete, so applications may want to employ additional verification during the restore to detect failures early. This might e.g. include incremental verification of each chunk against the app hash (using bundled Merkle proofs), checksums to protect against data corruption by the disk or network, and so on. However, it is important to note that the only trusted information available is the app hash, and all other snapshot metadata can be spoofed by adversaries. Apps may also want to consider state sync denial-of-service vectors, where adversaries provide invalid or harmful snapshots to prevent nodes from joining the network. The application can counteract this by asking CometBFT to ban peers. As a last resort, node operators can use P2P configuration options to whitelist a set of trusted peers that can provide valid snapshots.

Transition to Consensus

Once the snapshots have all been restored, CometBFT gathers additional information necessary for bootstrapping the node (e.g. chain ID, consensus parameters, validator sets, and block headers) from the genesis file and light client RPC servers. It also calls Info to verify the following:
  • that the app hash from the snapshot it has delivered to the Application matches the apphash stored in the next height’s block
  • that the version that the Application returns in ResponseInfo matches the version in the current height’s block header
Once the state machine has been restored and CometBFT has gathered this additional information, it transitions to consensus. As of ABCI 2.0, CometBFT ensures the necessary conditions to switch are met RFC-100. From the application’s point of view, these operations are transparent, unless the application has just upgraded to ABCI 2.0. In that case, the application needs to be properly configured and aware of certain constraints in terms of when to provide vote extensions. More details can be found in the section below. Once a node switches to consensus, it operates like any other node, apart from having a truncated block history at the height of the restored snapshot.

Application configuration required to switch to ABCI 2.0

Introducing vote extensions requires changes to the configuration of the application. First of all, switching to a version of CometBFT with vote extensions, requires a coordinated upgrade. For a detailed description on the upgrade path, please refer to the corresponding section in RFC-100. There is a newly introduced consensus parameter: VoteExtensionsEnableHeight. This parameter represents the height at which vote extensions are required for consensus to proceed, with 0 being the default value (no vote extensions). A chain can enable vote extensions either:
  • at genesis by setting VoteExtensionsEnableHeight to be equal, e.g., to the InitialHeight
  • or via the application logic by changing the ConsensusParam to configure the VoteExtensionsEnableHeight.
Once the (coordinated) upgrade to ABCI 2.0 has taken place, at height hu, the value of VoteExtensionsEnableHeight MAY be set to some height, he, which MUST be higher than the current height of the chain. Thus the earliest value for he is hu + 1. Once a node reaches the configured height, for all heights h ≥ he, the consensus algorithm will reject as invalid any precommit messages that do not have signed vote extension data. If the application requires it, a 0-length vote extension is allowed, but it MUST be signed and present in the precommit message. Likewise, for all heights h < he, any precommit messages that do have vote extensions will also be rejected as malformed. Height he is somewhat special, as calls to PrepareProposal MUST NOT have vote extension data, but all precommit votes in that height MUST carry a vote extension, even if the extension is nil. Height he + 1 is the first height for which PrepareProposal MUST have vote extension data and all precommit votes in that height MUST have a vote extension. Corollary, CometBFT will decide which data to store, and require for successful operations, based on the current height of the chain.