变更记录

  • 17.02.2021:初始草案

状态

已接受

摘要

本 ADR 引入了一种机制,用于在链软件升级期间执行原地状态存储迁移。

背景

当一次链升级在模块内部引入破坏状态兼容性的变更时,当前流程包括:将整个状态导出为 JSON 文件(通过 simd export 命令)、对该 JSON 文件运行迁移脚本(simd genesis migrate 命令)、清空存储(simd unsafe-reset-all 命令),然后使用迁移后的 JSON 文件作为新的 genesis 启动一条新链(可选地指定自定义初始区块高度)。此类流程的一个示例可见于 Cosmos Hub 3->4 迁移指南。 这一流程之所以繁琐,原因有多方面:
  • 该流程耗时较长。运行 export 命令可能需要数小时,之后在新链上使用迁移后的 JSON 运行 InitChain 还可能额外耗费数小时。
  • 导出的 JSON 文件可能非常庞大(约为 ~100MB-1GB),难以查看、编辑和传输,这又会带来额外工作去解决这些问题(例如 streaming genesis)。

决策

我们提出一种基于直接原地修改 KV 存储的迁移流程,不再依赖上述 JSON 导出-处理-导入流程。

模块 ConsensusVersion

我们在 AppModule 接口上引入一个新方法:
type AppModule interface {
    // --snip--
    ConsensusVersion() uint64
}
该方法返回一个 uint64,用作模块的状态破坏性版本号。模块每次引入破坏共识的变更时,都必须递增该值。为避免默认值带来的潜在错误,模块的初始版本必须设为 1。在 Cosmos SDK 中,版本 1 对应 v0.41 系列中的模块。

模块专属迁移函数

对于模块引入的每一次破坏共识的变更,都必须在 Configurator 中通过新增的 RegisterMigration 方法注册一个从 ConsensusVersion N 到版本 N+1 的迁移脚本。所有模块都会在其 AppModule 的 RegisterServices 方法中接收到 configurator 的引用,迁移函数应当在这里注册。这些迁移函数应按递增顺序注册。
func (am AppModule) RegisterServices(cfg module.Configurator) {
    // --snip--
    cfg.RegisterMigration(types.ModuleName, 1, func(ctx sdk.Context) error {
        // Perform in-place store migrations from ConsensusVersion 1 to 2.
    })

    cfg.RegisterMigration(types.ModuleName, 2, func(ctx sdk.Context) error {
        // Perform in-place store migrations from ConsensusVersion 2 to 3.
    })
    // etc.
}
例如,如果某模块新的 ConsensusVersion 为 N,那么必须在 configurator 中注册 N-1 个迁移函数。 在 Cosmos SDK 中,迁移函数由各模块的 keeper 处理,因为 keeper 持有用于执行原地存储迁移的 sdk.StoreKey。为避免 keeper 过于臃肿,每个模块都使用一个 Migrator 包装器来处理迁移函数:
// Migrator is a struct for handling in-place store migrations.
type Migrator struct {
    BaseKeeper
}
迁移函数应位于各模块的 migrations/ 目录中,并由 Migrator 的方法调用。我们建议方法命名格式为 Migrate{M}to{N}。
// Migrate1to2 migrates from version 1 to 2.
func (m Migrator) Migrate1to2(ctx sdk.Context) error {
    return v2bank.MigrateStore(ctx, m.keeper.storeKey) // v043bank is package `x/bank/migrations/v2`.
}
各模块的迁移函数都与该模块存储结构的演进密切相关,因此本文不对其具体实现进行描述。在引入 ADR-028 长度前缀地址后,x/bank 存储键迁移的一个示例可见于这段 store.go 代码。

在 x/upgrade 中跟踪模块版本

我们将在 x/upgrade 的存储中引入一个新的前缀存储。该存储将跟踪每个模块的当前版本,可以建模为从模块名到模块 ConsensusVersion 的 map[string]uint64,并将在运行迁移时使用(详见下一节)。使用的键前缀为 0x1,键值格式如下:
0x2 | {bytes(module_name)} => BigEndian(module_consensus_version)
该存储的初始状态通过 app.go 中的 InitChainer 方法设置。 UpgradeHandler 的签名需要更新,以接收一个 VersionMap,并返回升级后的 VersionMap 以及一个错误值:
- type UpgradeHandler func(ctx sdk.Context, plan Plan)
+ type UpgradeHandler func(ctx sdk.Context, plan Plan, versionMap VersionMap) (VersionMap, error)
应用升级时,我们会从 x/upgrade 存储中查询 VersionMap 并将其传入 handler。handler 运行实际的迁移函数(见下一节),如果成功,则返回一个更新后的 VersionMap,并写回状态。
func (k UpgradeKeeper) ApplyUpgrade(ctx sdk.Context, plan types.Plan) {
    // --snip--
-   handler(ctx, plan)
+   updatedVM, err := handler(ctx, plan, k.GetModuleVersionMap(ctx)) // k.GetModuleVersionMap() fetches the VersionMap stored in state.
+   if err != nil {
+       return err
+   }
+
+   // Set the updated consensus versions to state
+   k.SetModuleVersionMap(ctx, updatedVM)
}
还会新增一个 gRPC 查询端点,用于查询存储在 x/upgrade 状态中的 VersionMap,以便应用开发者在升级处理器运行前再次确认 VersionMap。

运行迁移

一旦所有迁移处理器都在 configurator 中注册完成(这会在启动时发生),就可以通过调用 module.Manager 上的 RunMigrations 方法来执行迁移。该函数会遍历所有模块,并对每个模块执行以下步骤:
  • 从其 VersionMap 参数中获取模块旧的 ConsensusVersion(记为 M)。
  • 通过 AppModule 上的 ConsensusVersion() 方法获取模块新的 ConsensusVersion(记为 N)。
  • 如果 N>M,则按顺序运行该模块所有已注册的迁移:M -> M+1 -> M+2...,直到 N。
    • 有一种特殊情况:如果该模块没有对应的 ConsensusVersion,说明该模块是在升级过程中新增的。此时不会运行任何迁移函数,而是将该模块当前的 ConsensusVersion 保存到 x/upgrade 的存储中。
如果缺少所需的迁移(例如未在 Configurator 中注册),那么 RunMigrations 函数将返回错误。 在实践中,RunMigrations 方法应当在 UpgradeHandler 内部调用。
app.UpgradeKeeper.SetUpgradeHandler("my-plan", func(ctx sdk.Context, plan upgradetypes.Plan, vm module.VersionMap) (module.VersionMap, error) {
    return app.mm.RunMigrations(ctx, vm)
})
假设一条链在区块 n 处升级,则流程应如下:
  • 旧二进制会在开始区块 N 时于 BeginBlock 中停机。在其存储中,保存着旧二进制各模块的 ConsensusVersion。
  • 新二进制会从区块 N 启动。UpgradeHandler 已在新二进制中设置,因此会在新二进制的 BeginBlock 中运行。在 x/upgrade 的 ApplyUpgrade 内部,将从存储(即旧二进制的存储)中取出 VersionMap,并传入 RunMigrations 函数,在各模块自己的 BeginBlock 运行前原地迁移所有模块的存储。

影响

向后兼容性

本 ADR 在 AppModule 上引入了一个新方法 ConsensusVersion(),所有模块都需要实现它。同时它也修改了 UpgradeHandler 的函数签名。因此,它不向后兼容。 虽然模块在提升 ConsensusVersion 时必须注册对应的迁移函数,但是否通过升级处理器运行这些脚本是可选的。应用完全可以选择不在其升级处理器中调用 RunMigrations,而继续使用传统的 JSON 迁移路径。

正面影响

  • 执行链升级时无需再操作 JSON 文件。
  • 虽然目前尚无基准测试,但原地存储迁移很可能比 JSON 迁移耗时更少。支持这一判断的主要原因是:旧二进制中的 simd export 命令和新二进制中的 InitChain 函数都将被跳过。

负面影响

  • 模块开发者必须正确跟踪其模块中的破坏共识变更。如果某个模块引入了破坏共识的变更,却没有同步提升对应的 ConsensusVersion(),那么 RunMigrations 函数将无法检测到该迁移,链升级可能会失败。文档应当明确反映这一点。

中性影响

  • Cosmos SDK 将继续通过现有的 simd export 和 simd genesis migrate 命令支持 JSON 迁移。
  • 当前 ADR 不支持创建、重命名或删除存储,只支持修改现有存储键和值。Cosmos SDK 已经提供 StoreLoader 来执行这些操作。

后续讨论

参考资料

  • 初始讨论:Link
  • ConsensusVersion 和 RunMigrations 的实现:Link
  • 讨论 x/upgrade 设计的议题:Link

Changelog

  • 17.02.2021: Initial Draft

Status

Accepted

Abstract

This ADR introduces a mechanism to perform in-place state store migrations during chain software upgrades.

Context

When a chain upgrade introduces state-breaking changes inside modules, the current procedure consists of exporting the whole state into a JSON file (via the simd export command), running migration scripts on the JSON file (simd genesis migrate command), clearing the stores (simd unsafe-reset-all command), and starting a new chain with the migrated JSON file as new genesis (optionally with a custom initial block height). An example of such a procedure can be seen in the Cosmos Hub 3->4 migration guide. This procedure is cumbersome for multiple reasons:
  • The procedure takes time. It can take hours to run the export command, plus some additional hours to run InitChain on the fresh chain using the migrated JSON.
  • The exported JSON file can be heavy (~100MB-1GB), making it difficult to view, edit and transfer, which in turn introduces additional work to solve these problems (such as streaming genesis).

Decision

We propose a migration procedure based on modifying the KV store in-place without involving the JSON export-process-import flow described above.

Module ConsensusVersion

We introduce a new method on the AppModule interface:
type AppModule interface {
    // --snip--
    ConsensusVersion()

uint64
}
This methods returns an uint64 which serves as state-breaking version of the module. It MUST be incremented on each consensus-breaking change introduced by the module. To avoid potential errors with default values, the initial version of a module MUST be set to 1. In the Cosmos SDK, version 1 corresponds to the modules in the v0.41 series.

Module-Specific Migration Functions

For each consensus-breaking change introduced by the module, a migration script from ConsensusVersion N to version N+1 MUST be registered in the Configurator using its newly-added RegisterMigration method. All modules receive a reference to the configurator in their RegisterServices method on AppModule, and this is where the migration functions should be registered. The migration functions should be registered in increasing order.
func (am AppModule)

RegisterServices(cfg module.Configurator) {
    // --snip--
    cfg.RegisterMigration(types.ModuleName, 1, func(ctx sdk.Context)

error {
        // Perform in-place store migrations from ConsensusVersion 1 to 2.
})

cfg.RegisterMigration(types.ModuleName, 2, func(ctx sdk.Context)

error {
        // Perform in-place store migrations from ConsensusVersion 2 to 3.
})
    // etc.
}
For example, if the new ConsensusVersion of a module is N , then N-1 migration functions MUST be registered in the configurator. In the Cosmos SDK, the migration functions are handled by each module’s keeper, because the keeper holds the sdk.StoreKey used to perform in-place store migrations. To not overload the keeper, a Migrator wrapper is used by each module to handle the migration functions:
// Migrator is a struct for handling in-place store migrations.
type Migrator struct {
    BaseKeeper
}
Migration functions should live inside the migrations/ folder of each module, and be called by the Migrator’s methods. We propose the format Migrate{M}to{N} for method names.
// Migrate1to2 migrates from version 1 to 2.
func (m Migrator)

Migrate1to2(ctx sdk.Context)

error {
    return v2bank.MigrateStore(ctx, m.keeper.storeKey) // v043bank is package `x/bank/migrations/v2`.
}
Each module’s migration functions are specific to the module’s store evolutions, and are not described in this ADR. An example of x/bank store key migrations after the introduction of ADR-028 length-prefixed addresses can be seen in this store.go code.

Tracking Module Versions in x/upgrade

We introduce a new prefix store in x/upgrade’s store. This store will track each module’s current version, it can be modelized as a map[string]uint64 of module name to module ConsensusVersion, and will be used when running the migrations (see next section for details). The key prefix used is 0x1, and the key/value format is:
0x2 | {bytes(module_name)} => BigEndian(module_consensus_version)
The initial state of the store is set from app.go’s InitChainer method. The UpgradeHandler signature needs to be updated to take a VersionMap, as well as return an upgraded VersionMap and an error:
- type UpgradeHandler func(ctx sdk.Context, plan Plan)
+ type UpgradeHandler func(ctx sdk.Context, plan Plan, versionMap VersionMap) (VersionMap, error)
To apply an upgrade, we query the VersionMap from the x/upgrade store and pass it into the handler. The handler runs the actual migration functions (see next section), and if successful, returns an updated VersionMap to be stored in state.
func (k UpgradeKeeper) ApplyUpgrade(ctx sdk.Context, plan types.Plan) {
    // --snip--
-   handler(ctx, plan)
+   updatedVM, err := handler(ctx, plan, k.GetModuleVersionMap(ctx)) // k.GetModuleVersionMap() fetches the VersionMap stored in state.
+   if err != nil {
+       return err
+   }
+
+   // Set the updated consensus versions to state
+   k.SetModuleVersionMap(ctx, updatedVM)
}
A gRPC query endpoint to query the VersionMap stored in x/upgrade’s state will also be added, so that app developers can double-check the VersionMap before the upgrade handler runs.

Running Migrations

Once all the migration handlers are registered inside the configurator (which happens at startup), running migrations can happen by calling the RunMigrations method on module.Manager. This function will loop through all modules, and for each module:
  • Get the old ConsensusVersion of the module from its VersionMap argument (let’s call it M).
  • Fetch the new ConsensusVersion of the module from the ConsensusVersion() method on AppModule (call it N).
  • If N>M, run all registered migrations for the module sequentially M -> M+1 -> M+2... until N.
    • There is a special case where there is no ConsensusVersion for the module, as this means that the module has been newly added during the upgrade. In this case, no migration function is run, and the module’s current ConsensusVersion is saved to x/upgrade’s store.
If a required migration is missing (e.g. if it has not been registered in the Configurator), then the RunMigrations function will error. In practice, the RunMigrations method should be called from inside an UpgradeHandler.
app.UpgradeKeeper.SetUpgradeHandler("my-plan", func(ctx sdk.Context, plan upgradetypes.Plan, vm module.VersionMap)  (module.VersionMap, error) {
    return app.mm.RunMigrations(ctx, vm)
})
Assuming a chain upgrades at block n, the procedure should run as follows:
  • the old binary will halt in BeginBlock when starting block N. In its store, the ConsensusVersions of the old binary’s modules are stored.
  • the new binary will start at block N. The UpgradeHandler is set in the new binary, so will run at BeginBlock of the new binary. Inside x/upgrade’s ApplyUpgrade, the VersionMap will be retrieved from the (old binary’s) store, and passed into the RunMigrations functon, migrating all module stores in-place before the modules’ own BeginBlocks.

Consequences

Backwards Compatibility

This ADR introduces a new method ConsensusVersion() on AppModule, which all modules need to implement. It also alters the UpgradeHandler function signature. As such, it is not backwards-compatible. While modules MUST register their migration functions when bumping ConsensusVersions, running those scripts using an upgrade handler is optional. An application may perfectly well decide to not call the RunMigrations inside its upgrade handler, and continue using the legacy JSON migration path.

Positive

  • Perform chain upgrades without manipulating JSON files.
  • While no benchmark has been made yet, it is probable that in-place store migrations will take less time than JSON migrations. The main reason supporting this claim is that both the simd export command on the old binary and the InitChain function on the new binary will be skipped.

Negative

  • Module developers MUST correctly track consensus-breaking changes in their modules. If a consensus-breaking change is introduced in a module without its corresponding ConsensusVersion() bump, then the RunMigrations function won’t detect the migration, and the chain upgrade might be unsuccessful. Documentation should clearly reflect this.

Neutral

  • The Cosmos SDK will continue to support JSON migrations via the existing simd export and simd genesis migrate commands.
  • The current ADR does not allow creating, renaming or deleting stores, only modifying existing store keys and values. The Cosmos SDK already has the StoreLoader for those operations.

Further Discussions

References

  • Initial discussion: Link
  • Implementation of ConsensusVersion and RunMigrations: Link
  • Issue discussing x/upgrade design: Link