变更记录

  • 2020-06-23:初始版本
  • 2020-08-06:根据审查意见修订,并调整为引用版本
  • 2021-01-15:修订以支持使用替代客户端进行解冻
  • 2021-05-20:修订以简化共识状态复制,移除初始高度
  • 2022-04-08:修订以弃用 AllowUpdateAfterExpiry 和 AllowUpdateAfterMisbehaviour
  • 2022-07-15:修订以允许更新 TrustingPeriod
  • 2023-09-05:修订以从 gov v1beta1 迁移到 gov v1

状态

已接受

背景

摘要

IBC 在启动时将是一个全新的协议,尚无成熟的用户群体。在协议层,无法区分客户端过期或作恶究竟是由真实故障(拜占庭行为)引起,还是由用户失误引起(例如未能及时更新客户端,或意外双重签名)。在基础 IBC 协议以及 ICS 20 同质化代币转账实现中,如果某个客户端无法再被更新,则该通道中的资金将被永久锁定,无法继续转移。在确保安全的前提下,更理想的做法是为用户提供一种恢复机制,以便在这些异常情况下使用。

异常情况

需要关注的状态是:与连接和通道关联的某个客户端无法再被更新。出现这种情况的原因可能有多种:
  1. 该客户端所跟踪的链已经停止运行,不再产生区块或区块头,因此无法再更新该客户端
  2. 该客户端所跟踪的链仍在继续运行,但在解绑期内没有任何中继者提交新的区块头,导致客户端已过期
    1. 这可能是由真实作恶(故意的拜占庭行为)引起,也可能只是验证者失误,但客户端无法区分这两种情况
  3. 该客户端所跟踪的链发生了作恶事件,客户端因此被冻结,从而无法再被更新

安全模型

验证者集合中的三分之二(治理和模块参与所需的法定人数)本就可以对任意数据进行签名,因此允许治理在经过一段延迟后手动使用新区块头强制更新客户端,并不会从根本上改变安全模型。

决策

我们选择不处理那些已经真正停止运行的链,因为这本质上必然属于拜占庭行为,而且在这种情况下,代币恢复本来也大概率无法实现(传输中的数据包无法超时,但这一点的相对影响较小)。
  1. 要求 Tendermint 轻客户端(ICS 07)在创建时带上下列额外标志
    1. allow_update_after_expiry(布尔值,默认 true)。注意,该标志现已弃用;它仅保留用于表达意图,但不会再强制执行针对该值的检查。
  2. 要求 Tendermint 轻客户端(ICS 07)暴露下列额外的内部查询函数
    1. Expired() boolean,用于返回客户端自上次更新以来是否已经超过信任期(此时将无法验证任何区块头)
  3. 要求 Tendermint 轻客户端(ICS 07)和 solo machine 客户端(ICS 06)在创建时带上下列额外标志
    1. allow_update_after_misbehaviour(布尔值,默认 true)。注意,该标志现已弃用;它仅保留用于表达意图,但不会再强制执行针对该值的检查。
  4. 要求 Tendermint 轻客户端(ICS 07)暴露下列额外的状态变更函数
    1. Unfreeze(),用于在发生作恶后对轻客户端解冻,并清除之前设置的任何冻结高度
  5. 新增一个带有 MsgRecoverClient 的治理提案。
    1. 创建一个新的 Msg,其中包含两个客户端标识符(string)以及一个签名者。
    2. 第一个客户端标识符是拟被更新的客户端。该客户端必须处于冻结或过期状态。
    3. 第二个客户端是替代客户端。它携带了可被更新客户端所需的全部状态。除最新高度、冻结高度和 chain-id 外,它必须与可被更新客户端具有完全相同的客户端参数和链参数。在投票期间,应持续对其进行更新。
    4. 如果该治理提案通过,则待恢复客户端将被更新为替代客户端的最新状态。
    5. 签名者必须是为 ibc 模块设置的权限主体。
    先前,AllowUpdateAfterExpiry 和 AllowUpdateAfterMisbehaviour 用于表明过期或冻结客户端的恢复选项;如果这些参数被设为 false,则不允许治理提案覆盖该客户端。但现在这已经被弃用,因为无论这些参数的值为何,代码迁移都可以覆盖客户端和共识状态。既然治理愿意投票去覆盖某个客户端或共识状态,那么治理大概率也愿意通过代码迁移来完成同样的事情。 此外,在最初的设计中,客户端升级提案不允许更新 TrustingPeriod。然而,鉴于生产环境中已经出现多种场景,例如某个规范通道在初始配置时发生错误,此时应允许更新客户端的 TrustingPeriod,因此治理应被允许更新这一客户端参数。 在早于 ibc-go v8 的版本中,MsgRecoverClient 是一种名为 ClientUpdateProposal 的治理提案类型。在从 governance v1beta1 迁移到 governance v1 的过程中,它已被移除,并由 MsgRecoverClient 取代。 注意,不应轻率地进行此类更新,因为从作恶发生到作恶证据被提交之间可能存在时间空档。举例来说,如果 UnbondingPeriod 为两周,而 TrustingPeriod 也被设置为两周,那么验证者可以等到 UnbondingPeriod 即将结束时提交虚假信息,然后解绑退出,从而不会因作恶而被惩罚。因此,我们建议将 07-tendermint 客户端的信任期设置为 UnbondingPeriod 的 2/3。
请注意,因作恶而被冻结的客户端必须等待证据过期,以避免再次被冻结。 本 ADR 不涉及计划内升级;这部分内容已按规范单独处理。

结果

正面影响

  • 建立了客户端在过期情况下的恢复机制
  • 建立了客户端在作恶情况下的恢复机制
  • 构造一个 ClientUpdate 提案的难度与创建一个新客户端相当

负面影响

  • 客户端创建过程增加了额外复杂性,用户必须理解这些内容
  • 复制替代客户端状态会增加复杂性
  • 治理参与者必须对替代客户端进行投票

中性影响

没有中性影响。

参考


Changelog

  • 23-06-2020: Initial version
  • 06-08-2020: Revisions per review & to reference version
  • 15-01-2021: Revision to support substitute clients for unfreezing
  • 20-05-2021: Revision to simplify consensus state copying, remove initial height
  • 08-04-2022: Revision to deprecate AllowUpdateAfterExpiry and AllowUpdateAfterMisbehaviour
  • 15-07-2022: Revision to allow updating of TrustingPeriod
  • 05-09-2023: Revision to migrate from gov v1beta1 to gov v1

Status

Accepted

Context

Summary

At launch, IBC will be a novel protocol, without an experienced user-base. At the protocol layer, it is not possible to distinguish between client expiry or misbehaviour due to genuine faults (Byzantine behaviour) and client expiry or misbehaviour due to user mistakes (failing to update a client, or accidentally double-signing). In the base IBC protocol and ICS 20 fungible token transfer implementation, if a client can no longer be updated, funds in that channel will be permanently locked and can no longer be transferred. To the degree that it is safe to do so, it would be preferable to provide users with a recovery mechanism which can be utilised in these exceptional cases.

Exceptional cases

The state of concern is where a client associated with connection(s) and channel(s) can no longer be updated. This can happen for several reasons:
  1. The chain which the client is following has halted and is no longer producing blocks/headers, so no updates can be made to the client
  2. The chain which the client is following has continued to operate, but no relayer has submitted a new header within the unbonding period, and the client has expired
    1. This could be due to real misbehaviour (intentional Byzantine behaviour) or merely a mistake by validators, but the client cannot distinguish these two cases
  3. The chain which the client is following has experienced a misbehaviour event, and the client has been frozen & thus can no longer be updated

Security model

Two-thirds of the validator set (the quorum for governance, module participation) can already sign arbitrary data, so allowing governance to manually force-update a client with a new header after a delay period does not substantially alter the security model.

Decision

We elect not to deal with chains which have actually halted, which is necessarily Byzantine behaviour and in which case token recovery is not likely possible anyways (in-flight packets cannot be timed-out, but the relative impact of that is minor).
  1. Require Tendermint light clients (ICS 07) to be created with the following additional flags
    1. allow_update_after_expiry (boolean, default true). Note that this flag has been deprecated, it remains to signal intent but checks against this value will not be enforced.
  2. Require Tendermint light clients (ICS 07) to expose the following additional internal query functions
    1. Expired() boolean, which returns whether or not the client has passed the trusting period since the last update (in which case no headers can be validated)
  3. Require Tendermint light clients (ICS 07) & solo machine clients (ICS 06) to be created with the following additional flags
    1. allow_update_after_misbehaviour (boolean, default true). Note that this flag has been deprecated, it remains to signal intent but checks against this value will not be enforced.
  4. Require Tendermint light clients (ICS 07) to expose the following additional state mutation functions
    1. Unfreeze(), which unfreezes a light client after misbehaviour and clears any frozen height previously set
  5. Add a new governance proposal with MsgRecoverClient.
    1. Create a new Msg with two client identifiers (string) and a signer.
    2. The first client identifier is the proposed client to be updated. This client must be either frozen or expired.
    3. The second client is a substitute client. It carries all the state for the client which may be updated. It must have identical client and chain parameters to the client which may be updated (except for latest height, frozen height, and chain-id). It should be continually updated during the voting period.
    4. If this governance proposal passes, the client on trial will be updated to the latest state of the substitute.
    5. The signer must be the authority set for the ibc module.
    Previously, AllowUpdateAfterExpiry and AllowUpdateAfterMisbehaviour were used to signal the recovery options for an expired or frozen client, and governance proposals were not allowed to overwrite the client if these parameters were set to false. However, this has now been deprecated because a code migration can overwrite the client and consensus states regardless of the value of these parameters. If governance would vote to overwrite a client or consensus state, it is likely that governance would also be willing to perform a code migration to do the same. In addition, TrustingPeriod was initially not allowed to be updated by a client upgrade proposal. However, due to the number of situations experienced in production where the TrustingPeriod of a client should be allowed to be updated because of ie: initial misconfiguration for a canonical channel, governance should be allowed to update this client parameter. In versions older than ibc-go v8, MsgRecoverClient was a governance proposal type ClientUpdateProposal. It has been removed and replaced by MsgRecoverClient in the migration from governance v1beta1 to governance v1. Note that this should NOT be lightly updated, as there may be a gap in time between when misbehaviour has occurred and when the evidence of misbehaviour is submitted. For example, if the UnbondingPeriod is 2 weeks and the TrustingPeriod has also been set to two weeks, a validator could wait until right before UnbondingPeriod finishes, submit false information, then unbond and exit without being slashed for misbehaviour. Therefore, we recommend that the trusting period for the 07-tendermint client be set to 2/3 of the UnbondingPeriod.
Note that clients frozen due to misbehaviour must wait for the evidence to expire to avoid becoming refrozen. This ADR does not address planned upgrades, which are handled separately as per the specification.

Consequences

Positive

  • Establishes a mechanism for client recovery in the case of expiry
  • Establishes a mechanism for client recovery in the case of misbehaviour
  • Constructing an ClientUpdate Proposal is as difficult as creating a new client

Negative

  • Additional complexity in client creation which must be understood by the user
  • Coping state of the substitute adds complexity
  • Governance participants must vote on a substitute client

Neutral

No neutral consequences.

References