分叉问责

问题陈述

Tendermint 共识算法在所有高度上保证以下规范:
  • 一致性 — 任意两个正确的全节点不会做出不同的决定。
  • 有效性 — 被决定的区块满足预定义谓词 valid()。
  • 终止性 — 所有正确的全节点最终都会做出决定。
前提是当前验证者集合中的故障验证者所拥有的投票权少于 1/3。如果这一假设 不成立,上述每一条规范都可能被破坏。 一致性属性表示:对于某一给定高度,任何两个在该高度上对区块做出决定的正确验证者,决定的都是同一个区块。该区块确实由区块链生成这一点,可以从一个可信(创世)区块开始验证,并检查其后的所有区块都已被正确签名。 然而,故障节点可能伪造区块,并试图让用户(轻客户端)相信这些区块是被正确生成的。此外,当 1/3 或以上的投票权属于故障验证者时,Tendermint 的一致性也可能被破坏:两个正确的验证者会对不同的区块做出决定。后一种情况正是“分叉”一词的由来:由于 Tendermint 共识还会就下一个验证者集合达成一致,正确验证者可能已经对彼此不相交的下一个验证者集合做出了决定,于是链会分裂为两个或更多分区(这些分区之间可能共享部分故障验证者),并且每个分支都会独立于其他分支继续生成区块。 我们称“分叉”是指:在区块链的同一高度上,存在两个针对不同区块的提交。问题在于,如何确保在这些情况下我们能够识别出故障验证者(而不会错误指控正确验证者),并以此激励验证者按照协议规范行事。 概念上的限制。 为了证明某个节点存在不当行为,我们必须说明其行为相对于某个给定算法的正确行为发生了偏离。因此,用于检测执行某个算法 A 的节点是否存在不当行为的算法,必须相对于算法 A 来定义。在我们的场景中,A 是 Tendermint 共识(加上基础设施中的其他协议;例如 Cosmos 全节点和轻客户端)。如果未来共识算法发生变化、更新或优化,我们就必须检查问责算法是否也需要相应调整。因此,本文中的所有讨论天然都是针对 Tendermint 共识和轻客户端规范的。 问: 对于一致性,我们是否应该区分验证者的一致性和全节点的一致性?如果所有正确的验证者都对同一个区块达成一致,但某个正确的全节点却对另一个区块做出了决定,这种情况似乎比两个正确验证者对不同区块做出决定要轻一些。尽管如此,如果一个被污染的全节点后来成为验证者,后续可能会带来问题。此外,如果一个被污染的全节点处在不同分支上,gossip 传播会如何受影响,目前也并不清楚。 备注。 当 1/3 或以上的投票权属于故障验证者时,有效性和终止性也可能被破坏。如果故障进程 simply 不发送推进所需的消息,终止性就可能被破坏。由于系统具有异步性,这种行为无法被惩罚,因为故障验证者总可以声称自己从未收到那些会迫使其发送消息的消息。

故障验证者的不当行为

分叉是故障验证者偏离协议的结果。原则上,即使分叉并未真正发生,这种偏离中的若干种也可以被检测出来:
  1. 双重提议:某个故障提议者在 Tendermint 共识中,对同一高度和同一轮提出两个不同的值(区块)。
  2. 双重签名:Tendermint 共识要求正确验证者在每一轮中至多只对一个值发送 prevote 和 precommit。如果某个故障验证者针对同一高度/轮次的不同值发送多个 prevote 和/或 precommit 消息,这就是不当行为。
  3. 狂乱验证者:Tendermint 共识要求正确验证者只对满足 valid(v) 的值 v 发送 prevote 和 precommit。如果故障验证者在 valid(v)=false 的情况下仍然对 v 发送 prevote 和 precommit,这就是不当行为。
备注。 单独来看,第 3 点针对的是有效性攻击(而不是一致性攻击)。不过,这些 prevote 和 precommit 也可以被用于伪造区块。
  1. 失忆:Tendermint 共识具有锁定机制。如果某个验证者锁定了某个值 v,那么它之后只能对 v 或 nil 发送 prevote/precommit。如果在仍然持有值 v 的锁时,对另一个不同的值 v’(且不是 nil)发送 prevote/precommit 消息,这就是不当行为。
  2. 虚假消息:在 Tendermint 共识中,大多数消息发送指令都受到阈值条件的保护,例如,必须先接收到 2f + 1 条 prevote 消息,才能发送 precommit。故障验证者可能在尚未收到这些 prevote 消息的情况下就发送 precommit。
无论是否发生分叉,惩罚这类行为都可能很重要,因为这有助于从根本上防止分叉。这应当能抑制攻击者作恶:如果故障投票权少于 1/3,这类不当行为虽然可以被检测到,但不会导致安全性违规。因此,除非攻击者拥有 1/3 或更多(某些情况下甚至超过 2/3)的投票权,否则他们就有动机不去作恶。如果攻击者控制了过多投票权,我们就必须处理分叉问题,正如本文所讨论的那样。

两种类型的分叉

  • Fork-Full。两个正确的验证者在同一高度上对不同的区块做出决定。由于还需要对下一个验证者集合做出决定,正确验证者可能会被分隔到分叉链的两个不同分支中参与。
在这种情况下,我们面对的是两个不同的区块(两者同样有权存在或同样无权存在),因此一个核心系统不变式(每个高度只能由正确验证者决定一个区块)被破坏了。由于此时全节点会被污染,这种污染也可能传播到轻客户端。然而,即使这个系统不变式没有被破坏,轻客户端仍然可能遭遇分叉:
  • Fork-Light。所有正确的验证者都对高度 h 的同一个区块做出决定,但故障进程(无论是否为验证者)会伪造该高度上的另一个区块,以欺骗用户(使用轻客户端的用户)。

攻击场景

链上攻击

矛盾签名(单轮)

存在多种可能导致分叉的场景。第一种是在同一轮中发生双重签名。
  • F1. 矛盾签名:故障验证者在给定高度 h 的同一轮 r 中,对不同的值签署多条投票消息(prevote 和/或 precommit)。

反复切换

Tendermint 共识实现了锁定机制:如果某个正确验证者 p 收到值 v 的提议,并在第 r 轮收到了针对值 id(v) 的 2f + 1 条 prevote,它就会锁定 v 并记住 r。在这种情况下,p 还会发送一条针对 id(v) 的 precommit 消息,这条消息之后可以作为 p 锁定了 v 的证明。 在后续轮次中,p 只会对它之前已经锁定过的值发送 prevote 消息。然而,如果在未来某一轮 r’ > r 中,该进程收到新的提议以及针对另一个不同值 v’ 的 2f + 1 条 prevote,那么锁定值就有可能被改变。在这种情况下,p 可以针对 id(v’) 发送 prevote/precommit。这一算法特性可以被以两种方式利用:
  • F2. 故障式反复切换(失忆):故障验证者在第 r 轮对某个值 id(v) 发送 precommit(即值 v 在第 r 轮被锁定),随后又在更高轮 r’ > r 中,在未先正确解锁值 v 的情况下,对另一个不同的值 id(v’) 发送 prevote。在这种情况下,故障进程“忘记了”自己已经锁定值 v,并在后续轮次中对其他值发送 prevote。 某些正确验证者可能已经在 r 轮对 v 做出决定,而另一些正确验证者则在 r’ 轮对 v’ 做出决定。此时主链上可能出现分支(Fork-Full)。
  • F3. 正确式反复切换(回到过去):存在一些由(正确)验证者签署的、针对值 id(v) 且属于第 r 轮的 precommit 消息。尽管如此,v 并未被决定,所有进程都进入下一轮。随后,正确验证者在某个更高轮 r’ > r 中(正确地)锁定并决定了另一个不同的值 v’。之后正确验证者继续向前推进;主链上不会出现分支。 然而,故障验证者可以利用第 r 轮中正确的 precommit 消息,再配合事后生成的、针对第 r 轮的故障 precommit 消息,伪造一个针对某个并未在主链上被决定的值的区块(Fork-Light)。

链下攻击

F1-F3 可能污染全节点(甚至验证者)的状态。因此,被污染的(但在其他方面仍然正确的)全节点可能会向轻客户端传播错误区块。 同样地,即使完全不干扰主链,也可能出现以下情况:
  • F4. 幽灵验证者:故障验证者在某些高度上投票(签署 prevote 和 precommit 消息),而在这些高度上它们并不属于(主链上的)验证者集合。
  • F5. 狂乱验证者:故障验证者签署投票消息,以支持某个(任意的)应用状态,而该状态不同于由有效状态转换产生的应用状态。

受害者类型

我们考虑三类潜在攻击受害者:
  • FN:全节点
  • LCS:按顺序验证头部的轻客户端
  • LCB:基于二分法验证头部的轻客户端
F1 和 F2 可被故障验证者用来真正地在区块链上创建多个分支。这意味着,正确运行的全节点会在同一高度上对不同区块做出决定。在全节点于本地检测到分叉之前(例如通过接收到来自其他节点的证据,或因某种本地检查失败),它都可能向轻客户端传播被破坏的区块。 备注。 如果全节点跟随的分支与验证者跟随的分支不同,gossip 协议的活性可能会受到影响。我们最终应当更仔细地研究这一点。不过,由于它不影响安全性,所以这不是首要问题。 F3 与 F1 类似,不同之处在于,不会有两个正确验证者对不同区块做出决定。尽管如此,全节点仍然可能受到影响。 此外,即使不在主链上制造分叉,只要有超过三分之一的故障验证者签署伪造的区块头,轻客户端也可能被污染。 F4 无法欺骗正确的全节点,因为它们知道当前的验证者集合。同样,LCS 也知道验证者是谁。因此,F4 针对的是那些不一定知道完整头部前缀的 LCB(Fork-Light),因为它们信任一个至少由一个正确验证者签名的头部(trusting period 方法)。 下表概述了不同攻击可能如何影响不同类型的节点。F1-F3 是链上攻击,因此它们可以破坏全节点的状态。随后,如果某个轻客户端(LCS 或 LCB)联系全节点以获取头部(或区块),这种被破坏的状态就可能传播到轻客户端。 F4 和 F5 是链下攻击,也就是说,这些攻击不能用于破坏全节点的状态(因为全节点对链状态有足够了解,不会被欺骗)。
攻击FNLCSLCB
F1直接FNFN
F2直接FNFN
F3直接FNFN
F4直接
F5直接
问: 轻客户端比全节点更脆弱,因为前者只验证头部而不执行交易。执行交易的全节点究竟获得了什么样的确定性? 由于全节点会验证所有交易,它只有在区块链本身违反其不变式(每个高度只能有一个区块)时,才可能被某种攻击污染,也就是说,只会在导致链分支的分叉情况下受到污染。

详细攻击场景

基于双签的攻击

在基于双签的攻击中,故障验证者会在某个高度的同一轮中对多个投票(prevote 和/或 precommit)进行签名。该攻击既可以针对全节点,也可以针对轻客户端执行。执行该攻击需要至少 1/3 的投票权。

场景 1:主链上的双签

验证者:
  • CA - 一组正确验证者,投票权少于 1/3
  • CB - 一组正确验证者,投票权少于 1/3
  • CA 和 CB 互不相交
  • F - 一组故障验证者,拥有 1/3 或更多投票权
注意,这种设置违反了 Cosmos 的故障模型。 执行过程:
  • 一个故障提议者向 CA 提议区块 A
  • 一个故障提议者向 CB 提议区块 B
  • 集合 CA 和 CB 中的验证者分别对 A 和 B 进行 prevote。
  • 集合 F 中的故障验证者同时对 A 和 B 进行 prevote。
  • 这些故障 prevote 消息
    • 针对 A 的消息,比 B 的消息更早到达 CA
    • 针对 B 的消息,比 A 的消息更早到达 CB
  • 因此,集合 CA 和 CB 中的正确验证者将分别观察到 超过 2/3 的针对 A 和 B 的 prevote,并分别对 A 和 B 进行 precommit。
  • 集合 F 中的故障验证者同时对值 A 和 B 进行 precommit。
  • 因此,A 和 B 都会得到超过 2/3 的 commit。
后果:
  • 在这种情况下,创建不当行为证据很简单,因为同一个故障进程会在同一轮中对不同值签署多条消息。
  • 我们必须确保这些不同的消息能够到达某个正确进程(全节点、监控器?),由其提交证据。
  • 这是针对全节点层面的攻击(Fork-Full)。
  • 它也会扩展到轻客户端,
  • 对这两者都需要检测和恢复机制。

场景 2:针对轻客户端的双签(LCS)

验证者:
  • 一组故障验证者 F,拥有超过 2/3 的投票权。
执行过程:
  • 在主链上,F 表现正常
  • F 协同签署一个与主链上不同的区块 B。
  • 轻客户端获得 B,并信任它,因为它带有超过 2/3 投票权的签名。
后果: 一旦双签被用于攻击轻客户端,就为各种不同类型的攻击打开了空间,因为应用状态可以向任意方向分叉。例如,它可以修改验证者集合,使其中只包含那些没有任何质押绑定的验证者。注意,一旦轻客户端被某个分叉欺骗,攻击者就可以任意改变应用状态和验证者集合。 为了检测这种(基于双签的)攻击,轻客户端需要与某个正确验证者交叉检查其状态(或者通过带外通道从主链获取状态哈希)。 备注。 轻客户端能够创建不当行为证据,但这可能需要从正确的全节点拉取大量数据。也许我们需要设计一种不同的架构:受到攻击的轻客户端会把其当前 unbonding period 内的所有数据推送给某个正确节点,由该节点检查这些数据并提交相应的证据。还有一些架构假设存在一个特殊角色(有时称为 fisherman),其目标是尽可能从网络中收集有用数据,进行分析并创建证据交易。该功能不在本文档的讨论范围内。 备注。 LCS 和 LCB 的区别可能只在于说服轻客户端接受任意状态所需的投票权数量。在 LCB 中,当安全阈值最低时,攻击者只需拥有 1/3 或更多投票权就可以任意修改应用状态;而在 LCS 中,则需要超过 2/3 的投票权。

反复横跳:基于失忆的攻击

在失忆攻击中,故障验证者会在某一轮 r 中锁定某个值 v,然后在更高轮次中,在没有正确解锁值 v 的情况下为另一个值 v’ 投票。该攻击既可用于全节点,也可用于轻客户端。

场景 3:故障至多为 2/3

验证者:
  • 一组故障验证者 F,拥有至少 1/3 但至多 2/3 的投票权
  • 一组正确验证者 C
执行过程:
  • 故障验证者通过收集超过 2/3 的 投票权(其中包含正确和故障验证者),在第 r 轮对区块 A 达成 commit(但不在主链上公开)。
  • 所有验证者(正确和故障)都进入某个 r’ > r 的轮次。
  • C 中某些正确验证者在第 r’ 轮之前没有锁定任何值。
  • F 中的故障验证者偏离 Tendermint 共识,忽略它们曾在 r 中锁定 A 的事实,并在 r’ 中提议另一个区块 B。
  • 由于 C 中那些未锁定任何值的验证者认为 B 可以接受,它们接受 B 的提议并对区块 B 达成 commit。
备注。 在这种情况下,超过 1/3 的故障验证者不需要实施双签(F1),因为在该执行过程中它们每轮只投票一次。 如果攻击者使用这种攻击并拥有 1/3 或更多但少于 2/3 的投票权来攻击轻客户端,那么它无法任意更改应用状态。相反,攻击者只能将状态限制在某个正确验证者认为可接受的范围内:在上述执行中,正确验证者仍然认为该值可接受,但轻客户端信任的区块会偏离主链上的区块。

场景 4:故障超过 2/3

如果攻击者拥有超过 2/3 的投票权,就可以任意更改应用状态。 验证者:
  • 一组故障验证者 F1,拥有 1/3 或更多投票权
  • 一组故障验证者 F2,拥有少于 1/3 的投票权
执行过程
  • 与场景 3 类似(但不需要正确验证者的消息)
  • F1 中的故障验证者在第 r 轮锁定值 A
  • 它们在后续轮次中为不同的值签名
  • F2 在第 r 轮不锁定 A
后果:
  • F1 中的验证者可以通过分叉问责机制被检测出来。
  • F2 中的验证者无法通过该机制被检测出来。 只有在它们签署了与应用冲突的内容时,才能据此追究它们。否则,它们并没有做任何错误的事。
问题: 我们是否需要为验证者签署任意状态的情况定义一种特殊攻击?看起来,检测这类攻击需要一种不同的机制,并且需要一系列导致该状态的区块作为证据。这可能会非常难以实现。

回到过去

在这类攻击中,故障验证者利用了自己在过去某些轮次中没有签署消息这一事实。由于 Tendermint 运行在异步网络中,我们很难区分这种攻击和延迟消息。这类攻击既可用于全节点,也可用于轻客户端。

场景 5

验证者:
  • C1 - 一组正确验证者,拥有超过 1/3 的投票权
  • C2 - 一组正确验证者,拥有 1/3 的投票权
  • C1 和 C2 互不相交
  • F - 一组故障验证者,拥有少于 1/3 的投票权
  • 另一个额外的故障进程 q
  • F 和 q 违反了 Cosmos 的故障模型。
执行过程:
  • 在高度 h 的某一轮 r 中,C1 对值 A 进行 precommit,
  • C2 对 nil 进行 precommit,
  • F 不发送任何消息
  • q 对 nil 进行 precommit。
  • 在某个 r’ > r 的轮次中,F、q 和 C2 对另一个不同于 A 的值 B 达成 commit。
  • F 和 fp “回到过去”,并在第 r 轮为值 A 签署 precommit 消息。
  • 再加上 C1 的 precommit 消息,这已经足以对值 A 达成 commit。
后果:
  • 只有一个之前对 nil 做过 precommit 的故障验证者实施了双签,而另外那 1/3 的故障验证者实际上执行了一种攻击,其消息序列与失忆攻击中的一部分完全相同。检测这类攻击最终归结为针对双签和失忆的机制。
问题: 我们是否应该将其保留为一种单独的攻击类型?看起来,双签、失忆和幻影验证者似乎是我们唯一需要支持的攻击类型,而这也能在其他情况下提供安全性。这并不令人意外,因为双签和失忆是由协议本身导出的攻击,而幻影攻击实际上并不是针对 Tendermint 的攻击,而更多是针对 Cosmos Proof of Stake 模块的攻击。

幻影验证者

在幻影验证者攻击中,那些不属于当前验证者集合、但仍处于绑定状态的进程(因为攻击发生在其 unbonding period 内)可以通过签署投票消息参与攻击。该攻击既可以针对全节点,也可以针对轻客户端执行。

场景 6

验证者:
  • F — 一组故障验证者,在高度 h + k 的主链上不属于验证者集合
执行过程:
  • 存在一个分叉,并且高度 h + k 有两个不同的 header,它们对应不同的验证者集合:
    • 主链上的 VS2
    • 由 F(以及其他方)签署的伪造 header VS2’
  • 轻客户端信任高度 h 的某个 header(以及相应的验证者集合 VS1)。
  • 作为二分 header 验证的一部分,它会使用新的验证者集合 VS2’ 来验证高度 h + k 的 header。
后果:
  • 为了检测这一点,节点需要同时看到伪造的 header 和链上的规范 header。
  • 如果满足这一点,那么检测这类攻击很容易,因为它只需要验证某些进程是否在自己并不属于验证者集合的高度上签署了消息。
备注。 我们可能会遇到以幻影验证者为后续步骤的攻击,其前提是先发生了基于双签或失忆的攻击,导致分叉状态中包含了那些不在主链验证者集合中的验证者。在这种情况下,尽管它们不属于主链上的验证者集合,但仍会继续签署消息并为分叉链(错误分支)做出贡献。这种攻击也可以在全节点被 eclipse 的一段时间内用于攻击全节点。 备注。 幻影验证者证据已从实现中移除,因为虽然它可能是一种合理的证据形式,但被认为并不相关。任何涉及幻影验证者的轻客户端攻击,都必然是由 1/3+ 的 lunatic 验证者发起的,这些验证者可以伪造一个包含该幻影验证者的新验证者集合。只有在 这种情况下,轻客户端才会接受幻影验证者的投票。我们只需要关注惩罚这个 1/3+ 的 lunatic 集团,因为它才是攻击的根本原因。

Lunatic 验证者

Lunatic 验证者会同意为任意应用状态签署 commit 消息。它被用于攻击轻客户端。 注意,检测这种行为需要应用层知识。检测这种行为很可能可以通过 参考发生该高度之前的那个区块来完成。 问题: 我们是否可以说,在这种情况下,验证者在投票之前拒绝检查提议值是否有效?

Fork accountability

Problem Statement

Tendermint consensus algorithm guarantees the following specifications for all heights:
  • agreement — no two correct full nodes decide differently.
  • validity — the decided block satisfies the predefined predicate valid().
  • termination — all correct full nodes eventually decide,
If the faulty validators have less than 1/3 of voting power in the current validator set. In the case where this assumption does not hold, each of the specification may be violated. The agreement property says that for a given height, any two correct validators that decide on a block for that height decide on the same block. That the block was indeed generated by the blockchain, can be verified starting from a trusted (genesis) block, and checking that all subsequent blocks are properly signed. However, faulty nodes may forge blocks and try to convince users (light clients) that the blocks had been correctly generated. In addition, Tendermint agreement might be violated in the case where 1/3 or more of the voting power belongs to faulty validators: Two correct validators decide on different blocks. The latter case motivates the term “fork”: as Tendermint consensus also agrees on the next validator set, correct validators may have decided on disjoint next validator sets, and the chain branches into two or more partitions (possibly having faulty validators in common) and each branch continues to generate blocks independently of the other. We say that a fork is a case in which there are two commits for different blocks at the same height of the blockchain. The problem is to ensure that in those cases we are able to detect faulty validators (and not mistakenly accuse correct validators), and incentivize therefore validators to behave according to the protocol specification. Conceptual Limit. In order to prove misbehavior of a node, we have to show that the behavior deviates from correct behavior with respect to a given algorithm. Thus, an algorithm that detects misbehavior of nodes executing some algorithm A must be defined with respect to algorithm A. In our case, A is Tendermint consensus (+ other protocols in the infrastructure; e.g., Cosmos full nodes and the Light Client). If the consensus algorithm is changed/updated/optimized in the future, we have to check whether changes to the accountability algorithm are also required. All the discussions in this document are thus inherently specific to Tendermint consensus and the Light Client specification. Q: Should we distinguish agreement for validators and full nodes for agreement? The case where all correct validators agree on a block, but a correct full node decides on a different block seems to be slightly less severe that the case where two correct validators decide on different blocks. Still, if a contaminated full node becomes validator that may be problematic later on. Also it is not clear how gossiping is impaired if a contaminated full node is on a different branch. Remark. In the case 1/3 or more of the voting power belongs to faulty validators, also validity and termination can be broken. Termination can be broken if faulty processes just do not send the messages that are needed to make progress. Due to asynchrony, this is not punishable, because faulty validators can always claim they never received the messages that would have forced them to send messages.

The Misbehavior of Faulty Validators

Forks are the result of faulty validators deviating from the protocol. In principle several such deviations can be detected without a fork actually occurring:
  1. double proposal: A faulty proposer proposes two different values (blocks) for the same height and the same round in Tendermint consensus.
  2. double signing: Tendermint consensus forces correct validators to prevote and precommit for at most one value per round. In case a faulty validator sends multiple prevote and/or precommit messages for different values for the same height/round, this is a misbehavior.
  3. lunatic validator: Tendermint consensus forces correct validators to prevote and precommit only for values v that satisfy valid(v). If faulty validators prevote and precommit for v although valid(v)=false this is misbehavior.
Remark. In isolation, Point 3 is an attack on validity (rather than agreement). However, the prevotes and precommits can then also be used to forge blocks.
  1. amnesia: Tendermint consensus has a locking mechanism. If a validator has some value v locked, then it can only prevote/precommit for v or nil. Sending prevote/precomit message for a different value v’ (that is not nil) while holding lock on value v is misbehavior.
  2. spurious messages: In Tendermint consensus most of the message send instructions are guarded by threshold guards, e.g., one needs to receive 2f + 1 prevote messages to send precommit. Faulty validators may send precommit without having received the prevote messages.
Independently of a fork happening, punishing this behavior might be important to prevent forks altogether. This should keep attackers from misbehaving: if less than 1/3 of the voting power is faulty, this misbehavior is detectable but will not lead to a safety violation. Thus, unless they have 1/3 or more (or in some cases more than 2/3) of the voting power attackers have the incentive to not misbehave. If attackers control too much voting power, we have to deal with forks, as discussed in this document.

Two types of forks

  • Fork-Full. Two correct validators decide on different blocks for the same height. Since also the next validator sets are decided upon, the correct validators may be partitioned to participate in two distinct branches of the forked chain.
As in this case we have two different blocks (both having the same right/no right to exist), a central system invariant (one block per height decided by correct validators) is violated. As full nodes are contaminated in this case, the contamination can spread also to light clients. However, even without breaking this system invariant, light clients can be subject to a fork:
  • Fork-Light. All correct validators decide on the same block for height h, but faulty processes (validators or not), forge a different block for that height, in order to fool users (who use the light client).

Attack scenarios

On-chain attacks

Equivocation (one round)

There are several scenarios in which forks might happen. The first is double signing within a round.
  • F1. Equivocation: faulty validators sign multiple vote messages (prevote and/or precommit) for different values during the same round r at a given height h.

Flip-flopping

Tendermint consensus implements a locking mechanism: If a correct validator p receives proposal for value v and 2f + 1 prevotes for a value id(v) in round r, it locks v and remembers r. In this case, p also sends a precommit message for id(v), which later may serve as proof that p locked v. In subsequent rounds, p only sends prevote messages for a value it had previously locked. However, it is possible to change the locked value if in a future round r’ > r, if the process receives proposal and 2f + 1 prevotes for a different value v’. In this case, p could send a prevote/precommit for id(v’). This algorithmic feature can be exploited in two ways:
  • F2. Faulty Flip-flopping (Amnesia): faulty validators precommit some value id(v) in round r (value v is locked in round r) and then prevote for different value id(v’) in higher round r’ > r without previously correctly unlocking value v. In this case faulty processes “forget” that they have locked value v and prevote some other value in the following rounds. Some correct validators might have decided on v in r, and other correct validators decide on v’ in r’. Here we can have branching on the main chain (Fork-Full).
  • F3. Correct Flip-flopping (Back to the past): There are some precommit messages signed by (correct) validators for value id(v) in round r. Still, v is not decided upon, and all processes move on to the next round. Then correct validators (correctly) lock and decide a different value v’ in some round r’ > r. And the correct validators continue; there is no branching on the main chain. However, faulty validators may use the correct precommit messages from round r together with a posteriori generated faulty precommit messages for round r to forge a block for a value that was not decided on the main chain (Fork-Light).

Off-chain attacks

F1-F3 may contaminate the state of full nodes (and even validators). Contaminated (but otherwise correct) full nodes may thus communicate faulty blocks to light clients. Similarly, without actually interfering with the main chain, we can have the following:
  • F4. Phantom validators: faulty validators vote (sign prevote and precommit messages) in heights in which they are not part of the validator sets (at the main chain).
  • F5. Lunatic validator: faulty validator that sign vote messages to support (arbitrary) application state that is different from the application state that resulted from valid state transitions.

Types of victims

We consider three types of potential attack victims:
  • FN: full node
  • LCS: light client with sequential header verification
  • LCB: light client with bisection based header verification
F1 and F2 can be used by faulty validators to actually create multiple branches on the blockchain. That means that correctly operating full nodes decide on different blocks for the same height. Until a fork is detected locally by a full node (by receiving evidence from others or by some other local check that fails), the full node can spread corrupted blocks to light clients. Remark. If full nodes take a branch different from the one taken by the validators, it may be that the liveness of the gossip protocol may be affected. We should eventually look at this more closely. However, as it does not influence safety it is not a primary concern. F3 is similar to F1, except that no two correct validators decide on different blocks. It may still be the case that full nodes become affected. In addition, without creating a fork on the main chain, light clients can be contaminated by more than a third of validators that are faulty and sign a forged header F4 cannot fool correct full nodes as they know the current validator set. Similarly, LCS know who the validators are. Hence, F4 is an attack against LCB that do not necessarily know the complete prefix of headers (Fork-Light), as they trust a header that is signed by at least one correct validator (trusting period method). The following table gives an overview of how the different attacks may affect different nodes. F1-F3 are on-chain attacks so they can corrupt the state of full nodes. Then if a light client (LCS or LCB) contacts a full node to obtain headers (or blocks), the corrupted state may propagate to the light client. F4 and F5 are off-chain, that is, these attacks cannot be used to corrupt the state of full nodes (which have sufficient knowledge on the state of the chain to not be fooled).
AttackFNLCSLCB
F1directFNFN
F2directFNFN
F3directFNFN
F4direct
F5direct
Q: Light clients are more vulnerable than full nodes, because the former do only verify headers but do not execute transactions. What kind of certainty is gained by a full node that executes a transaction? As a full node verifies all transactions, it can only be contaminated by an attack if the blockchain itself violates its invariant (one block per height), that is, in case of a fork that leads to branching.

Detailed Attack Scenarios

Equivocation based attacks

In case of equivocation based attacks, faulty validators sign multiple votes (prevote and/or precommit) in the same round of some height. This attack can be executed on both full nodes and light clients. It requires 1/3 or more of voting power to be executed.

Scenario 1: Equivocation on the main chain

Validators:
  • CA - a set of correct validators with less than 1/3 of the voting power
  • CB - a set of correct validators with less than 1/3 of the voting power
  • CA and CB are disjoint
  • F - a set of faulty validators with 1/3 or more voting power
Observe that this setting violates the Cosmos failure model. Execution:
  • A faulty proposer proposes block A to CA
  • A faulty proposer proposes block B to CB
  • Validators from the set CA and CB prevote for A and B, respectively.
  • Faulty validators from the set F prevote both for A and B.
  • The faulty prevote messages
    • for A arrive at CA long before the B messages
    • for B arrive at CB long before the A messages
  • Therefore correct validators from set CA and CB will observe more than 2/3 of prevotes for A and B and precommit for A and B, respectively.
  • Faulty validators from the set F precommit both values A and B.
  • Thus, we have more than 2/3 commits for both A and B.
Consequences:
  • Creating evidence of misbehavior is simple in this case as we have multiple messages signed by the same faulty processes for different values in the same round.
  • We have to ensure that these different messages reach a correct process (full node, monitor?), which can submit evidence.
  • This is an attack on the full node level (Fork-Full).
  • It extends also to the light clients,
  • For both we need a detection and recovery mechanism.

Scenario 2: Equivocation to a light client (LCS)

Validators:
  • a set F of faulty validators with more than 2/3 of the voting power.
Execution:
  • for the main chain F behaves nicely
  • F coordinates to sign a block B that is different from the one on the main chain.
  • the light clients obtains B and trusts at as it is signed by more than 2/3 of the voting power.
Consequences: Once equivocation is used to attack light client it opens space for different kind of attacks as application state can be diverged in any direction. For example, it can modify validator set such that it contains only validators that do not have any stake bonded. Note that after a light client is fooled by a fork, that means that an attacker can change application state and validator set arbitrarily. In order to detect such (equivocation-based attack), the light client would need to cross check its state with some correct validator (or to obtain a hash of the state from the main chain using out of band channels). Remark. The light client would be able to create evidence of misbehavior, but this would require to pull potentially a lot of data from correct full nodes. Maybe we need to figure out different architecture where a light client that is attacked will push all its data for the current unbonding period to a correct node that will inspect this data and submit corresponding evidence. There are also architectures that assumes a special role (sometimes called fisherman) whose goal is to collect as much as possible useful data from the network, to do analysis and create evidence transactions. That functionality is outside the scope of this document. Remark. The difference between LCS and LCB might only be in the amount of voting power needed to convince light client about arbitrary state. In case of LCB where security threshold is at minimum, an attacker can arbitrarily modify application state with 1/3 or more of voting power, while in case of LCS it requires more than 2/3 of the voting power.

Flip-flopping: Amnesia based attacks

In case of amnesia, faulty validators lock some value v in some round r, and then vote for different value v’ in higher rounds without correctly unlocking value v. This attack can be used both on full nodes and light clients.

Scenario 3: At most 2/3 of faults

Validators:
  • a set F of faulty validators with 1/3 or more but at most 2/3 of the voting power
  • a set C of correct validators
Execution:
  • Faulty validators commit (without exposing it on the main chain) a block A in round r by collecting more than 2/3 of the voting power (containing correct and faulty validators).
  • All validators (correct and faulty) reach a round r’ > r.
  • Some correct validators in C do not lock any value before round r’.
  • The faulty validators in F deviate from Tendermint consensus by ignoring that they locked A in r, and propose a different block B in r’.
  • As the validators in C that have not locked any value find B acceptable, they accept the proposal for B and commit a block B.
Remark. In this case, the more than 1/3 of faulty validators do not need to commit an equivocation (F1) as they only vote once per round in the execution. If a light client is attacked using this attack with 1/3 or more of voting power (and less than 2/3), the attacker cannot change the application state arbitrarily. Rather, the attacker is limited to a state a correct validator finds acceptable: In the execution above, correct validators still find the value acceptable, however, the block the light client trusts deviates from the one on the main chain.

Scenario 4: More than 2/3 of faults

In case there is an attack with more than 2/3 of the voting power, an attacker can arbitrarily change application state. Validators:
  • a set F1 of faulty validators with 1/3 or more of the voting power
  • a set F2 of faulty validators with less than 1/3 of the voting power
Execution
  • Similar to Scenario 3 (however, messages by correct validators are not needed)
  • The faulty validators in F1 lock value A in round r
  • They sign a different value in follow-up rounds
  • F2 does not lock A in round r
Consequences:
  • The validators in F1 will be detectable by the fork accountability mechanisms.
  • The validators in F2 cannot be detected using this mechanism. Only in case they signed something which conflicts with the application this can be used against them. Otherwise, they do not do anything incorrect.
Q: do we need to define a special kind of attack for the case where a validator sign arbitrarily state? It seems that detecting such attack requires a different mechanism that would require as an evidence a sequence of blocks that led to that state. This might be very tricky to implement.

Back to the past

In this kind of attack, faulty validators take advantage of the fact that they did not sign messages in some of the past rounds. Due to the asynchronous network in which Tendermint operates, we cannot easily differentiate between such an attack and delayed message. This kind of attack can be used at both full nodes and light clients.

Scenario 5

Validators:
  • C1 - a set of correct validators with over 1/3 of the voting power
  • C2 - a set of correct validators with 1/3 of the voting power
  • C1 and C2 are disjoint
  • F - a set of faulty validators with less than 1/3 voting power
  • one additional faulty process q
  • F and q violate the Cosmos failure model.
Execution:
  • in a round r of height h we have C1 precommitting a value A,
  • C2 precommits nil,
  • F does not send any message
  • q precommits nil.
  • In some round r’ > r, F and q and C2 commit some other value B different from A.
  • F and fp “go back to the past” and sign precommit message for value A in round r.
  • Together with precomit messages of C1 this is sufficient for a commit for value A.
Consequences:
  • Only a single faulty validator that previously precommited nil did equivocation, while the other 1/3 of faulty validators actually executed an attack that has exactly the same sequence of messages as part of amnesia attack. Detecting this kind of attack boil down to mechanisms for equivocation and amnesia.
Q: should we keep this as a separate kind of attack? It seems that equivocation, amnesia and phantom validators are the only kind of attack we need to support and this gives us security also in other cases. This would not be surprising as equivocation and amnesia are attacks that followed from the protocol and phantom attack is not really an attack to Tendermint but more to the Cosmos Proof of Stake module.

Phantom validators

In case of phantom validators, processes that are not part of the current validator set but are still bonded (as attack happen during their unbonding period) can be part of the attack by signing vote messages. This attack can be executed against both full nodes and light clients.

Scenario 6

Validators:
  • F — a set of faulty validators that are not part of the validator set on the main chain at height h + k
Execution:
  • There is a fork, and there exist two different headers for height h + k, with different validator sets:
    • VS2 on the main chain
    • forged header VS2’, signed by F (and others)
  • a light client has a trust in a header for height h (and the corresponding validator set VS1).
  • As part of bisection header verification, it verifies the header at height h + k with new validator set VS2’.
Consequences:
  • To detect this, a node needs to see both, the forged header and the canonical header from the chain.
  • If this is the case, detecting these kind of attacks is easy as it just requires verifying if processes are signing messages in heights in which they are not part of the validator set.
Remark. We can have phantom-validator-based attacks as a follow up of equivocation or amnesia based attack where forked state contains validators that are not part of the validator set at the main chain. In this case, they keep signing messages contributed to a forked chain (the wrong branch) although they are not part of the validator set on the main chain. This attack can also be used to attack full node during a period of time it is eclipsed. Remark. Phantom validator evidence has been removed from implementation as it was deemed, although possibly a plausible form of evidence, not relevant. Any attack on the light client involving a phantom validator will have needed to be initiated by 1/3+ lunatic validators that can forge a new validator set that includes the phantom validator. Only in that case will the light client accept the phantom validators vote. We need only worry about punishing the 1/3+ lunatic cabal, that is the root cause of the attack.

Lunatic validator

Lunatic validator agrees to sign commit messages for arbitrary application state. It is used to attack light clients. Note that detecting this behavior require application knowledge. Detecting this behavior can probably be done by referring to the block before the one in which height happen. Q: can we say that in this case a validator declines to check if a proposed value is valid before voting for it?