核心验证
问题陈述
我们假设轻客户端已知一个它所信任的(基础)头inithead(通过社会共识,或因为轻客户端此前已决定信任该头)。目标是基于 inithead 中的数据,检查另一个头 newhead 是否可以被信任。
该协议的正确性基于这样一个假设:inithead 是由 Tendermint 共识的某个实例生成的。
故障模型
为便于给出下面的定义,我们假设存在一个函数validators,它会针对给定的哈希返回对应的验证者集合。
轻客户端协议是相对于如下故障模型定义的:
给定一个已知界限 TRUSTED_PERIOD,以及一个区块 b,其头 h 在时间 Time 生成
(即 h.Time = Time),那么在 validators(b.Header.NextValidatorsHash) 中,持有超过 2/3 投票权的那部分验证者集合,在时间 b.Header.Time + TRUSTED_PERIOD 之前都是正确的。
假设:“正确”是相对于实时定义的(某种牛顿式的全局时间概念,即墙上时钟时间),而 Header.Time 对应的是 BFT 时间。在本文中,我们假设正确进程的时钟是同步的(例如通过 NTP),因此本地时钟与 BFT 时间之间存在有界时钟漂移(CLOCK_DRIFT)。更准确地说,对于每个正确的轻客户端进程,以及每个 header.Time(即由 Tendermint 共识正确生成的头对应的 BFT 时间),都满足如下不等式:Header.Time < now + CLOCK_DRIFT,其中 now 对应轻客户端进程上的系统时钟。
此外,我们假设 TRUSTED_PERIOD 比 CLOCK_DRIFT 大(好几个)数量级(TRUSTED_PERIOD >> CLOCK_DRIFT),因为 CLOCK_DRIFT(使用 NTP 时)通常是毫秒量级,而 TRUSTED_PERIOD 通常是数周量级。
我们预期,本文定义的轻客户端进程会被用于这样一种上下文:存在一个更长的时间段,在该时间段内,行为不当的验证者能够被检测并惩罚(由于现代权益证明系统中的“bonding”机制,我们通常将其称为 UNBONDING_PERIOD)。此外,我们假设
TRUSTED_PERIOD < UNBONDING_PERIOD,并且二者通常处于同一数量级,例如
TRUSTED_PERIOD = UNBONDING_PERIOD / 2。
本文中的规范考虑的是在上述故障模型下的轻客户端实现。像 fork accountability 和 evidence submission 这样的机制,是在 UNBONDING_PERIOD 的上下文中定义的,
它们会激励验证者遵守本文定义的协议规范。如果他们不遵守,
且存在 1/3(或更多)故障验证者,则安全性可能会被破坏。此时我们的做法是
在事后检测这些情况,并采取适当的修复措施(自动化和社会层面的)。
相关内容将在 分叉问责 文档中讨论。
上文中的“trusted”一词表示,协议的正确性依赖于
这一假设。运行轻客户端的用户有责任确保:信任一个被破坏或伪造的 inithead 所带来的风险可以忽略不计。
备注:未来这个故障模型可能会变为一个同时考虑高度的混合版本。
高层方案
初始化时,轻客户端会获得一个它所信任的头inithead(通过
社会共识)。当轻客户端看到一个新的已签名头 snh 时,它必须决定是否信任这个新头。
信任可以通过以下三种方法的(可能的)组合来获得。
-
不中断的头序列。 给定一个可信头
h和一个不可信头h1, 如果轻客户端信任h与h1之间的所有头,那么它就信任头h1。 -
可信期。 给定一个可信头
h、一个不可信头h1 > h,以及故障模型成立的TRUSTED_PERIOD,我们可以检查是否至少有一个从h.Time到现在持续保持正确的验证者签署了h1。如果是这样,我们就可以信任h1。 -
二分。 如果按照 2.(可信期)进行的检查失败,轻客户端可以尝试获取一个高度位于
h和h1之间的头hp,以检查是否可以利用h为hp建立信任,以及利用hp为snh建立信任。如果可以,我们就能信任h1; 如果不行,则继续递归,直到我们找到一组头,能够在h与h1之间建立(传递性的)信任关系,或者由于两个相邻头彼此无法通过验证而失败。
定义
数据结构
下文只给出本规范所需的数据结构细节。函数
为了本轻客户端规范的目的,我们假设 Cosmos 全节点通过 RPC 暴露以下函数:trustThreshold 作为参数。为简单起见,
我们假设 trustThreshold 是一个介于 1/3 和 2/3 之间的浮点数,并且不会在伪代码中检查它。
VerifySingle。 函数 VerifySingle 尝试基于给定的可信状态,验证给定的不可信头以及对应的验证者集合。
它会确保可信状态仍处于其信任期内,
并且不可信头相对于传入时间 now 的时间在假设的 clockDrift 界限之内。
注意,此函数不会对全节点发起外部(RPC)调用;整个逻辑
都基于本地(给定)状态。该函数预期由 IBC 处理器使用。
VerifySingle 返回时没有错误(即不可信头
被成功验证),那么我们可以保证,信任已在
trustedState.SignedHeader.Header 的信任期内
从 trustedState 转移到 newTrustedState。
TODO:解释 VerifySingle 返回错误时会发生什么。
verifySingle。 函数 verifySingle 针对给定的可信状态验证单个不可信头。
它包含所有校验和签名验证。
由于它不会检查头是否过期(时间约束),
因此可能被错误使用,所以不会对外公开。
VerifyHeaderAtHeight 封装了高层逻辑,
也就是应用调用轻客户端模块来下载并验证某个高度的头。
VerifyHeaderAtHeight 返回时没有错误(即不可信头
被成功验证),那么我们可以保证,信任已在
trustedState.SignedHeader.Header 的信任期内
从 trustedState 转移到 newTrustedState。
如果 VerifyHeaderAtHeight 返回错误,那么要么 (i) 我们正在通信的全节点有故障,
要么 (ii) 可信头已过期(即超出其信任期)。对于情况 (i),全节点存在故障,因此
轻客户端应断开连接,并使用新的对等节点重新初始化。对于情况 (ii),由于可信头已过期,
我们需要使用一个新的可信头(其仍处于信任期内)重新初始化轻客户端,
但不一定需要断开与当前全节点的连接(因为在这种情况下我们并未观察到全节点有不当行为)。
VerifyBisection。 函数 VerifyBisection 实现了递归逻辑,用于检查
是否可以通过有限集合的(已下载并已验证的)头,
在 trustedState 与给定高度处的不可信头之间建立信任关系。
untrustedHeader.Height < trustedHeader.Height 的情况
在这样一种使用场景中:有人告知轻客户端,与其相关的应用数据可以在高度为 k 的区块中读取,而轻客户端信任一个更新的区块头,此时我们可以使用哈希来“沿链向下”验证区块头。也就是说,我们沿着高度递减迭代,并在每一步检查哈希。
备注。 对于轻客户端同时信任两个区块头 i 和 j,且 i < k < j 的情况,我们应当讨论或实验前向方法还是后向方法更有效。
- 轻客户端与一个全节点通信
- 轻客户端会在本地存储所有通过基础验证且仍处于轻客户端信任期内的区块头。在下方伪代码中,我们将其记为 Store.Add(header)。如果某个区块头验证失败,那么当前通信的全节点就是有故障的,我们应当断开与其的连接,并使用新的对等节点重新初始化。
- 如果
CanTrust返回 error,则说明轻客户端看到了一个伪造区块头,或者受信任区块头已过期(即超出了其信任期)。- 如果是伪造区块头,则全节点有故障,因此轻客户端应断开连接并使用新的对等节点重新初始化。如果受信任区块头已过期,我们需要用一个新的受信任区块头(且其仍在信任期内)重新初始化轻客户端,但在这种情况下不一定需要断开当前全节点连接,因为我们并未观察到全节点存在作恶行为。
轻客户端协议的正确性
定义
TRUSTED_PERIOD:信任期- 对于实时时间
t,如果验证者v在时间t之前一直遵循协议,则谓词correct(v,t)为真(关于恢复问题我们稍后讨论)。 - 验证者字段。我们将验证者写作元组
(v,p),其中v是标识符(即验证者地址;我们假设在每个验证者集合中标识符唯一)p是其投票权重
- 对于每个区块头
h,若轻客户端信任h,则记为trust(h) = true。
故障模型
如果一个带有区块头h 的区块 b 在时间 Time 被生成(即 h.Time = Time),则在 validators(h.NextValidatorsHash) 中,持有超过 2/3 投票权重的一组验证者在时间 h.Time + TRUSTED_PERIOD 之前都是正确的。
形式化地,
-
轻客户端完备性:如果某个区块头
h是由 Tendermint 共识实例正确生成的(并且其年龄小于信任期),那么轻客户端最终应将trust(h)设为true。 -
轻客户端准确性:如果某个区块头
h不是由 Tendermint 共识实例生成的,那么轻客户端绝不能将trust(h)设为 true。
trust(h) 应当在 h.Time + TRUSTED_PERIOD 之前被设为 true。否则该区块头由于过旧而不能被信任。
备注:如果某个区块头 h 被标记为 trust(h),但在某个记作 now 的时刻它已经过旧(h.Time + TRUSTED_PERIOD < now),那么轻客户端应在 now 时刻再次将 trust(h) 设为 false。
假设:初始时,轻客户端拥有一个它所信任的区块头 inithead,也就是说,inithead 是由 Tendermint 共识正确生成的。
为了论证正确性,我们可以证明如下不变式。
验证条件:轻客户端不变式。
对于每个轻客户端 l 和每个区块头 h:
如果 l 已将 trust(h) = true,
那么在时间 h.Time + TRUSTED_PERIOD 之前保持正确的验证者,在 validators(h.NextValidatorsHash) 中持有超过三分之二的投票权重。
形式化地,
细节
观察 1。 如果h.Time + TRUSTED_PERIOD > now,我们就信任验证者集合 validators(h.NextValidatorsHash)。
当我们说信任 validators(h.NextValidatorsHash) 时,不意味着我们信任 validators(h.NextValidatorsHash) 中每一个单独的验证者都是正确的;我们只信任这样一个事实:其中故障验证者少于 1/3(更准确地说,故障验证者持有的总投票权重少于 1/3)。
VerifySingle 的正确性论证
轻客户端准确性:
- 反证法。假设
untrustedHeader并非正确生成,但由于verifySingle无错误返回,轻客户端却将其设为受信任。 trustedState是受信任的且足够新。- 根据故障模型,故障验证者持有的投票权重少于
1/3,因此至少有一个正确验证者v签署了untrustedHeader。 - 由于
v到当前为止都是正确的,它至少在签署untrustedHeader之前一直遵循 Tendermint 共识协议,因此untrustedHeader是被正确生成的。 我们得到了所需的矛盾。
- 如果
trustedState中有足够多的验证者在高度untrustedHeader.Height时仍然是验证者,并且签署了untrustedHeader,则检查会成功。 - 如果
untrustedHeader.Height = trustedHeader.Height + 1,且这两个区块头都是正确生成的,则测试通过。
untrustedSignedHeader.Header.Height = trustedHeader.Height + 1,那么
signers(untrustedSignedHeader.Commit) \subseteq validators(trustedHeader.NextValidatorsHash)。
备注:如果用户认为仅依赖一个正确验证者还不够,那么可以使用变量 trustThreshold。
但是,在验证者集合频繁变更的情况下,trustThreshold 选得越高,verifySingle 对非相邻区块头返回错误的可能性就越大。
VerifyBisection的正确性论证(概要)*
- 反证法。假设从全节点获得的高度为
untrustedHeight的区块头并非正确生成,而由于VerifyBisection无错误返回,轻客户端却将其设为受信任。 - 只有当递归中的所有
verifySingle调用都无错误返回(返回nil)时,VerifyBisection才会无错误返回。 - 因此,我们得到了一串都满足
verifySingle的区块头。 - 再次得到矛盾。
Commit(pivot) 时,轻客户端始终都能拿到一个被正确生成的区块头,这一点才能得到保证。
停滞
在 VerifyBisection 中,有故障的全节点可以构造一个很长的区块头序列,使轻客户端逐个查询这些区块头,而它们在表面上都没有问题,直到轻客户端最终发现问题为止,从而导致轻客户端停滞。对此有几种处理方式:
- 每次调用
Commit都可以发给不同的全节点 - 轻客户端不必逐个查询区块头,而是告知全节点它信任哪个区块头,以及它需要哪个高度的区块头。全节点返回目标区块头,以及由中间区块头组成的证明,轻客户端可用其进行验证。粗略地说,此时
VerifyBisection会在全节点侧执行。 - 我们可以为
VerifyBisection允许花费的时间设置超时。
Core Verification
Problem statement
We assume that the light client knows a (base) headerinithead it trusts (by social consensus or because
the light client has decided to trust the header before). The goal is to check whether another header
newhead can be trusted based on the data in inithead.
The correctness of the protocol is based on the assumption that inithead was generated by an instance of
Tendermint consensus.
Failure Model
For the purpose of the following definitions we assume that there exists a functionvalidators that returns the corresponding validator set for the given hash.
The light client protocol is defined with respect to the following failure model:
Given a known bound TRUSTED_PERIOD, and a block b with header h generated at time Time
(i.e. h.Time = Time), a set of validators that hold more than 2/3 of the voting power
in validators(b.Header.NextValidatorsHash) is correct until time b.Header.Time + TRUSTED_PERIOD.
Assumption: “correct” is defined w.r.t. realtime (some Newtonian global notion of time, i.e., wall time),
while Header.Time corresponds to the BFT time. In this note, we assume that clocks of correct processes
are synchronized (for example using NTP), and therefore there is bounded clock drift (CLOCK_DRIFT) between local clocks and
BFT time. More precisely, for every correct light client process and every header.Time (i.e. BFT Time, for a header correctly
generated by the Tendermint consensus), the following inequality holds: Header.Time < now + CLOCK_DRIFT,
where now corresponds to the system clock at the light client process.
Furthermore, we assume that TRUSTED_PERIOD is (several) order of magnitude bigger than CLOCK_DRIFT (TRUSTED_PERIOD >> CLOCK_DRIFT),
as CLOCK_DRIFT (using NTP) is in the order of milliseconds and TRUSTED_PERIOD is in the order of weeks.
We expect a light client process defined in this document to be used in the context in which there is some
larger period during which misbehaving validators can be detected and punished (we normally refer to it as UNBONDING_PERIOD
due to the “bonding” mechanism in modern proof of stake systems). Furthermore, we assume that
TRUSTED_PERIOD < UNBONDING_PERIOD and that they are normally of the same order of magnitude, for example
TRUSTED_PERIOD = UNBONDING_PERIOD / 2.
The specification in this document considers an implementation of the light client under the Failure Model defined above.
Mechanisms like fork accountability and evidence submission are defined in the context of UNBONDING_PERIOD and
they incentivize validators to follow the protocol specification defined in this document. If they don’t,
and we have 1/3 (or more) faulty validators, safety may be violated. Our approach then is
to detect these cases (after the fact), and take suitable repair actions (automatic and social).
This is discussed in document on Fork accountability.
The term “trusted” above indicates that the correctness of the protocol depends on
this assumption. It is in the responsibility of the user that runs the light client to make sure that the risk
of trusting a corrupted/forged inithead is negligible.
Remark: This failure model might change to a hybrid version that takes heights into account in the future.
High Level Solution
Upon initialization, the light client is given a headerinithead it trusts (by
social consensus). When a light clients sees a new signed header snh, it has to decide whether to trust the new
header. Trust can be obtained by (possibly) the combination of three methods.
-
Uninterrupted sequence of headers. Given a trusted header
hand an untrusted headerh1, the light client trusts a headerh1if it trusts all headers in betweenhandh1. -
Trusted period. Given a trusted header
h, an untrusted headerh1 > handTRUSTED_PERIODduring which the failure model holds, we can check whether at least one validator, that has been continuously correct fromh.Timeuntil now, has signedh1. If this is the case, we can trusth1. -
Bisection. If a check according to 2. (trusted period) fails, the light client can try to
obtain a header
hpwhose height lies betweenhandh1in order to check whetherhcan be used to get trust forhp, andhpcan be used to get trust forsnh. If this is the case we can trusth1; if not, we continue recursively until either we found set of headers that can build (transitively) trust relation betweenhandh1, or we failed as two consecutive headers don’t verify against each other.
Definitions
Data structures
In the following, only the details of the data structures needed for this specification are given.Functions
For the purpose of this light client specification, we assume that the Cosmos Full Node exposes the following functions over RPC:trustThreshold as a parameter. For simplicity
we assume that trustThreshold is a float between 1/3 and 2/3 and we will not be checking it
in the pseudo-code.
VerifySingle. The function VerifySingle attempts to validate given untrusted header and the corresponding validator sets
based on a given trusted state. It ensures that the trusted state is still within its trusted period,
and that the untrusted header is within assumed clockDrift bound of the passed time now.
Note that this function is not making external (RPC) calls to the full node; the whole logic is
based on the local (given) state. This function is supposed to be used by the IBC handlers.
VerifySingle returns without an error (untrusted header
is successfully verified) then we have a guarantee that the transition of the trust
from trustedState to newTrustedState happened during the trusted period of
trustedState.SignedHeader.Header.
TODO: Explain what happens in case VerifySingle returns with an error.
verifySingle. The function verifySingle verifies a single untrusted header
against a given trusted state. It includes all validations and signature verification.
It is not publicly exposed since it does not check for header expiry (time constraints)
and hence it’s possible to use it incorrectly.
VerifyHeaderAtHeight captures high level
logic, i.e., application call to the light client module to download and verify header
for some height.
VerifyHeaderAtHeight returns without an error (untrusted header
is successfully verified) then we have a guarantee that the transition of the trust
from trustedState to newTrustedState happened during the trusted period of
trustedState.SignedHeader.Header.
In case VerifyHeaderAtHeight returns with an error, then either (i) the full node we are talking to is faulty
or (ii) the trusted header has expired (it is outside its trusted period). In case (i) the full node is faulty so
light client should disconnect and reinitialize with new peer. In the case (ii) as the trusted header has expired,
we need to reinitialize light client with a new trusted header (that is within its trusted period),
but we don’t necessarily need to disconnect from the full node we are talking to (as we haven’t observed full node misbehavior in this case).
VerifyBisection. The function VerifyBisection implements
recursive logic for checking if it is possible building trust
relationship between trustedState and untrusted header at the given height over
finite set of (downloaded and verified) headers.
The case untrustedHeader.Height < trustedHeader.Height
In the use case where someone tells the light client that application data that is relevant for it
can be read in the block of height k and the light client trusts a more recent header, we can use the
hashes to verify headers “down the chain.” That is, we iterate down the heights and check the hashes in each step.
Remark. For the case were the light client trusts two headers i and j with i < k < j, we should
discuss/experiment whether the forward or the backward method is more effective.
- the light client communicates with one full node
- the light client locally stores all the headers that has passed basic verification and that are within light client trust period. In the pseudo code below we write Store.Add(header) for this. If a header failed to verify, then the full node we are talking to is faulty and we should disconnect from it and reinitialize with new peer.
- If
CanTrustreturns error, then the light client has seen a forged header or the trusted header has expired (it is outside its trusted period).- In case of forged header, the full node is faulty so light client should disconnect and reinitialize with new peer. If the trusted header has expired, we need to reinitialize light client with new trusted header (that is within its trusted period), but we don’t necessarily need to disconnect from the full node we are talking to (as we haven’t observed full node misbehavior in this case).
Correctness of the Light Client Protocols
Definitions
TRUSTED_PERIOD: trusted period- for realtime
t, the predicatecorrect(v,t)is true if the validatorvfollows the protocol until timet(we will see about recovery later). - Validator fields. We will write a validator as a tuple
(v,p)such thatvis the identifier (i.e., validator address; we assume identifiers are unique in each validator set)pis its voting power
- For each header
h, we writetrust(h) = trueif the light client trustsh.
Failure Model
If a blockb with a header h is generated at time Time (i.e. h.Time = Time), then a set of validators that
hold more than 2/3 of the voting power in validators(h.NextValidatorsHash) is correct until time
h.Time + TRUSTED_PERIOD.
Formally,
-
Light Client Completeness: If a header
hwas correctly generated by an instance of Tendermint consensus (and its age is less than the trusted period), then the light client should eventually settrust(h)totrue. -
Light Client Accuracy: If a header
hwas not generated by an instance of Tendermint consensus, then the light client should never settrust(h)to true.
trust(h) should be set to true before h.Time + TRUSTED_PERIOD. If not, the header
cannot be trusted because it is too old.
Remark: If a header h is marked with trust(h), but it is too old at some point in time we denote with now (h.Time + TRUSTED_PERIOD < now),
then the light client should set trust(h) to false again at time now.
Assumption: Initially, the light client has a header inithead that it trusts, that is, inithead was correctly generated by the Tendermint consensus.
To reason about the correctness, we may prove the following invariant.
Verification Condition: light Client Invariant.
For each light client l and each header h:
if l has set trust(h) = true,
then validators that are correct until time h.Time + TRUSTED_PERIOD have more than two thirds of the voting power in validators(h.NextValidatorsHash).
Formally,
Details
Observation 1. Ifh.Time + TRUSTED_PERIOD > now, we trust the validator set validators(h.NextValidatorsHash).
When we say we trust validators(h.NextValidatorsHash) we do not trust that each individual validator in validators(h.NextValidatorsHash)
is correct, but we only trust the fact that less than 1/3 of them are faulty (more precisely, the faulty ones have less than 1/3 of the total voting power).
VerifySingle correctness arguments
Light Client Accuracy:
- Assume by contradiction that
untrustedHeaderwas not generated correctly and the light client sets trust to true becauseverifySinglereturns without error. trustedStateis trusted and sufficiently new- by the Failure Model, less than
1/3of the voting power held by faulty validators => at least one correct validatorvhas signeduntrustedHeader. - as
vis correct up to now, it followed the Tendermint consensus protocol at least up to signinguntrustedHeader=>untrustedHeaderwas correctly generated. We arrive at the required contradiction.
- The check is successful if sufficiently many validators of
trustedStateare still validators in the heightuntrustedHeader.Heightand signeduntrustedHeader. - If
untrustedHeader.Height = trustedHeader.Height + 1, and both headers were generated correctly, the test passes.
untrustedSignedHeader.Header.Height = trustedHeader.Height + 1 then
signers(untrustedSignedHeader.Commit) \subseteq validators(trustedHeader.NextValidatorsHash).
Remark: The variable trustThreshold can be used if the user believes that relying on one correct validator is not sufficient.
However, in case of (frequent) changes in the validator set, the higher the trustThreshold is chosen, the more unlikely it becomes that
verifySingle returns with an error for non-adjacent headers.
VerifyBisectioncorrectness arguments (sketch)*
- Assume by contradiction that the header at
untrustedHeightobtained from the full node was not generated correctly and the light client sets trust to true becauseVerifyBisectionreturns without an error. VerifyBisectionreturns without error only if all calls toverifySinglein the recursion return without error (returnnil).- Thus we have a sequence of headers that all satisfied the
verifySingle - again a contradiction
Commit(pivot) the light client is always provided with a correctly generated header.
Stalling
With VerifyBisection, a faulty full node could stall a light client by creating a long sequence of headers that are queried one-by-one by the light client and look OK,
before the light client eventually detects a problem. There are several ways to address this:
- Each call to
Commitcould be issued to a different full node - Instead of querying header by header, the light client tells a full node which header it trusts, and the height of the header it needs. The full node responds with
the header along with a proof consisting of intermediate headers that the light client can use to verify. Roughly,
VerifyBisectionwould then be executed at the full node. - We may set a timeout how long
VerifyBisectionmay take.