变更日志

  • 2022-04-27:第一版草稿

状态

草案

摘要

为了将 Cosmos SDK 迁移到一种彼此解耦、遵循语义化版本控制的模块体系,这些模块可以按不同组合方式进行编排(例如 staking v3 搭配 bank v1 和 distribution v2),我们需要重新评估模块 API 暴露面的组织方式,以避免 Go 语义化导入版本控制以及循环依赖带来的问题。本文 ADR 将探讨我们可以采用的多种方案来解决这些问题。

背景

社区一直有相当多的诉求,希望 SDK 支持语义化版本控制,同时也在积极推动将 SDK 模块拆分为独立的 Go 模块。理想情况下,这两项工作都能让生态系统迭代得更快,因为我们不再需要等待所有依赖同步更新。比如,我们可以拥有 3 个核心 SDK 版本,它们兼容最新的 2 个 CosmWasm 发行版,以及 4 个不同版本的 staking。这种模式能让早期采用者更激进地集成新版本,同时也让更保守的用户可以自行选择准备好采用哪些版本。 为了实现这一点,我们需要解决以下问题:
  1. 由于 Go 语义化导入版本控制(SIV)的工作方式,如果直接、天真地迁移到 SIV,实际上会让这些目标更难达成
  2. 模块之间的循环依赖必须被打破,SDK 中的许多模块才能真正独立发布
  3. 即使 protobuf schema 是按正确方式演进的,如果没有正确的未知字段过滤,仍会引入隐蔽的次版本不兼容问题
请注意,下面所有讨论都假设一个模块的 proto 文件版本控制与状态机版本控制是彼此独立的,具体体现在:
  • proto 文件以非破坏性方式维护(使用类似 buf breaking 的工具来确保所有变更都向后兼容)
  • proto 文件版本的提升频率要低得多,也就是说,我们可能会在 bank 模块状态机经历多个版本期间持续维护 cosmos.bank.v1
  • 状态机的破坏性变更更为常见,而理想情况下,我们希望通过 Go 模块对这部分做语义化版本控制,例如 x/bank/v2、x/bank/v3 等

问题 1:与语义化导入版本控制的兼容性

假设我们有一个模块 foo,它定义了如下 MsgDoSomething,并且我们已将其状态机作为 Go 模块 example.com/foo 发布:
package foo.v1;

message MsgDoSomething {
  string sender = 1;
  uint64 amount = 2;
}

service Msg {
  DoSomething(MsgDoSomething) returns (MsgDoSomethingResponse);
}
现在再考虑,我们对这个模块做了一次修订:为 MsgDoSomething 新增了一个 condition 字段,同时新增了一条针对 amount 的校验规则,要求它必须非零;并且遵循 Go 语义化版本控制的约定,将 foo 的下一个状态机版本发布为 example.com/foo/v2。
// Revision 1
package foo.v1;

message MsgDoSomething {
  string sender = 1;
  
  // amount must be a non-zero integer.
  uint64 amount = 2;
  
  // condition is an optional condition on doing the thing.
  //
  // Since: Revision 1
  Condition condition = 3;
}
如果采用最直接的做法,我们会将 foo 初始版本的 protobuf 类型生成到 example.com/foo/types 中,并将第二个版本的 protobuf 类型生成到 example.com/foo/v2/types 中。 现在假设我们有一个模块 bar,它通过 foo 提供的这个 keeper 接口与 foo 通信:
type FooKeeper interface {
    DoSomething(MsgDoSomething)

error
}

场景 A:向后兼容:较新的 Foo,较旧的 Bar

设想我们有一条链同时使用 foo 和 bar,并且希望升级到 foo/v2,但 bar 模块还没有升级到 foo/v2。 在这种情况下,这条链将无法升级到 foo/v2,除非 bar 先把它对 example.com/foo/types.MsgDoSomething 的引用升级为 example.com/foo/v2/types.MsgDoSomething。 即使 bar 对 MsgDoSomething 的使用方式完全没有变化,这个升级依然不可能完成,因为在 Go 类型系统中,example.com/foo/types.MsgDoSomething 和 example.com/foo/v2/types.MsgDoSomething 本质上是两种不同且互不兼容的结构体。

场景 B:向前兼容:较旧的 Foo,较新的 Bar

现在考虑相反的场景:bar 通过将 MsgDoSomething 的引用改成 example.com/foo/v2/types.MsgDoSomething 来升级到 foo/v2,并将其连同链所需的其他一些变更一起作为 bar/v2 发布。然而,这条链认为 foo/v2 中的变更风险过高,因此更希望继续停留在 foo 的初始版本上。 在这种场景下,即使 bar/v2 除了把 MsgDoSomething 的导入路径改掉之外,与 foo 的配合本来可以 100% 正常工作(也就是说,bar/v2 实际上并没有使用 foo/v2 的任何新特性),也不可能在不升级到 foo/v2 的前提下升级到 bar/v2。 由于 Go 语义化导入版本控制的工作方式,我们最终被锁定在只能使用 foo 和 bar,或者 foo/v2 和 bar/v2 这两种组合之一。我们无法得到 foo + bar/v2,也无法得到 foo/v2 + bar。即使这些模块的两个版本在其他方面彼此兼容,Go 类型系统也不允许这样做。

朴素缓解方案

一种朴素的修复思路是:不要在 example.com/foo/v2/types 中重新生成 protobuf 类型,而是直接更新 example.com/foo/types,让它反映 v2 所需的变更(新增 condition 并要求 amount 非零)。然后我们可以发布一个包含这些更新的 example.com/foo/types 补丁版本,并将其用于 foo/v2。但这个变更对 v1 来说会破坏状态机兼容性。它要求修改 ValidateBasic 方法,以拒绝 amount 为零的情况;同时它还新增了 condition 字段,而依据ADR 020 未知字段过滤,这个字段本应被拒绝。因此,把这些变更作为 v1 的补丁实际上违背了语义化版本控制。那些希望继续停留在 foo 的 v1 上的链,不应该导入这些变更,因为它们对于 v1 来说是不正确的。

问题 2:循环依赖

如果 foo 和 bar 因某些原因以不同方式互相依赖,那么上述任何方案都无法让它们成为彼此独立的模块。比如,我们不能让 foo 导入 bar/types,同时又让 bar 导入 foo/types。 SDK 中已经存在多种模块循环依赖的情况(例如 staking、distribution 和 slashing),从状态机角度看,这些依赖是合理的。如果不以某种方式把 API 类型拆分出去,那么在没有其他缓解手段的前提下,就无法对这些模块进行彼此独立的语义化版本控制。

问题 3:处理次版本不兼容

假设我们已经解决了前两个问题,但现在出现这样一种情况:bar/v2 希望能够选择性地使用只有 foo/v2 才支持的 MsgDoSomething.condition。如果 bar/v2 在与 foo v1 配合时,把 condition 设置为某个非 nil 值,那么 foo 会静默忽略这个字段,从而导致一种静默的逻辑错误,这种错误可能具有危险性。假如 bar/v2 能够检查 foo 当前运行的是 v1 还是 v2,并动态地仅在 foo/v2 可用时才使用 condition,它就可以避免这个问题。然而,即便 bar/v2 能执行这项检查,我们又如何确保它总能正确执行?如果没有某种框架层面的未知字段过滤,就很难判断这些隐蔽且难以检测的缺陷是否已经进入我们的应用。而要实现这一点,可能需要类似ADR 033:模块间通信这样的客户端-服务端层。

解决方案

方案 A)分离 API 模块与状态机模块

一种方案(最早在链接中提出)是将所有 protobuf 生成代码从状态机模块中拆分出来,放入一个单独的模块中。这意味着我们可以有状态机 Go 模块 foo 和 foo/v2,它们共同使用一个类型或 API Go 模块,比如 foo/api。这个 foo/api Go 模块将始终停留在 v1.x,并且只接受非破坏性变更。这样一来,只要模块间 API 仅依赖 foo/api 中的类型,其他模块就可以同时兼容 foo 或 foo/v2。这也会让 foo 和 bar 两个模块能够互相依赖,因为它们都可以依赖 foo/api 和 bar/api,而不需要 foo 直接依赖 bar,反之亦然。 这与前面描述的朴素缓解方案类似,不同之处在于它将类型拆分到了独立的 Go 模块中,而这本身就可以用来打破模块之间的循环依赖。除此之外,它仍然存在与朴素方案相同的问题,我们可以通过以下方式进行修正:
  1. 从 API 模块中移除所有会破坏状态机兼容性的代码(例如 ValidateBasic 以及其他任何接口方法)
  2. 在二进制中嵌入用于未知字段过滤的正确文件描述符

将 API 类型上的所有接口方法迁移为处理器

为了解决第 1 点,我们需要移除生成类型上的所有接口实现,转而使用处理器模式。其本质是,对于类型 X,我们提供某种解析器,用于解析该类型对应的接口实现(例如 sdk.Msg 或 authz.Authorization)。例如:
func (k Keeper)

DoSomething(msg MsgDoSomething)

error {
    var validateBasicHandler ValidateBasicHandler
    err := k.resolver.Resolve(&validateBasic, msg)
    if err != nil {
    return err
}

err = validateBasicHandler.ValidateBasic()
	...
}
对于 sdk.Msg 上的某些方法,我们可以用声明式注解来替代。例如,GetSigners 已经可以由 protobuf 注解 cosmos.msg.v1.signer 替代。未来,我们还可以考虑某种 protobuf 校验框架(类似 链接,但更贴合 Cosmos 的需求)来替代 ValidateBasic。

固定的 FileDescriptor

为了解决第 2 点,状态机模块必须能够指定其构建时所依赖的 protobuf 文件版本。比如,如果 foo 的 API 模块升级到 foo/v2,原始的 foo 模块仍然需要保留一份它构建时所使用的原始 protobuf 文件副本,这样 ADR 020 的未知字段过滤才能在 condition 被设置时拒绝 MsgDoSomething。 最简单的方式,可能是将 protobuf 的 FileDescriptor 直接嵌入模块本身,以便运行时使用这些 FileDescriptor,而不是使用可能已经发生变化、内置在 foo/api 中的那一份。借助 buf build、go embed 以及构建脚本,我们大概率可以设计出一种相当直接的方案,把 FileDescriptor 嵌入模块之中。

生成代码的潜在限制

这种方法的一个挑战在于,它对 API 模块中可以包含的内容施加了严格限制,并且要求其中大部分内容都属于会破坏状态机的变更。API 模块中的全部或大部分代码都会从 protobuf 文件生成,因此我们大概可以通过控制代码生成方式来约束这一点,但这仍然是一个需要注意的风险。 例如,我们会为 ORM 做代码生成,而未来其中可能包含会破坏状态机的优化。我们要么需要非常谨慎地确保这些优化在生成代码中实际上不会破坏状态机,要么将这部分生成代码从 API 模块中拆分出来,放入状态机模块。这两种缓解方式都可能可行,但 API 模块方案确实需要额外谨慎,以避免这类问题。

次版本不兼容

这种方法本身几乎无法解决潜在的次版本不兼容问题,以及所需的未知字段过滤。很可能需要某种执行此检查的客户端-服务端路由层,例如 ADR 033:模块间通信,才能确保这件事被正确处理。这样一来,我们就可以允许模块在运行时基于 MsgClient 执行检查,例如:
func (k Keeper)

CallFoo()

error {
    if k.interModuleClient.MinorRevision(k.fooMsgClient) >= 2 {
    k.fooMsgClient.DoSomething(&MsgDoSomething{
    Condition: ...
})
}

else {
        ...
}
}
要实现未知字段过滤本身,ADR 033 路由器需要使用 protoreflect API 来确保不会设置接收模块未知的字段。根据这部分逻辑的复杂程度,这可能会带来不理想的性能损耗。

方案 B) 修改生成代码

解决版本问题的另一种方法是改变 protobuf 代码的生成方式,并让模块大部分或完全转向 ADR 033 中描述的模块间通信方向。在这种范式下,一个模块可以在内部生成它所需的所有类型,包括其他模块的 API 类型,并通过客户端-服务端边界与其他模块通信。例如,如果 bar 需要与 foo 通信,它可以将自己的 MsgDoSomething 版本生成为 bar/internal/foo/v1.MsgDoSomething,然后将其传递给模块间路由器,由后者以某种方式将其转换为 foo 所需的版本(例如 foo/internal.MsgDoSomething)。 当前,在同一个 go 二进制中,如果没有特殊的构建标志,就不能同时存在针对同一个 protobuf 类型生成的两个结构体(见链接)。对此,一个相对简单的缓解方案是设置 protobuf 代码,使其在 internal/ 包中生成时不全局注册 protobuf 类型。这将要求模块在应用级 protobuf 注册表中手动注册其类型,这与模块当前对 InterfaceRegistry 和 amino codec 所做的事情类似。 如果模块只通过 ADR 033 进行消息传递,那么将 bar/internal/foo/v1.MsgDoSomething 转换为 foo/internal.MsgDoSomething 的一种朴素且性能不佳的方案,就是在 ADR 033 路由器中进行编组和反编组。如果我们需要在 Keeper 接口中暴露 protobuf 类型,这种做法就行不通了,因为这里的核心目标正是尽量将这些类型保留在 internal/ 中,从而避免前面描述的各种导入版本不兼容问题。不过,考虑到次版本不兼容的问题以及未知字段过滤的需求,一开始就继续坚持 Keeper 范式而不是 ADR 033,可能并不可行。 一种性能更好的方案是(也许还能调整为适用于 Keeper 接口)只暴露生成类型的 getter 和 setter,并在内部将数据存储在内存缓冲区中,这样就可以以零拷贝的方式在不同实现之间传递。 例如,假设针对 MsgSend 暴露了下面这个仅包含 getter 和 setter 的 protobuf API:
type MsgSend interface {
    proto.Message
	GetFromAddress()

string
	GetToAddress()

string
	GetAmount() []v1beta1.Coin
    SetFromAddress(string)

SetToAddress(string)

SetAmount([]v1beta1.Coin)
}

func NewMsgSend()

MsgSend {
    return &msgSendImpl{
    memoryBuffers: ...
} 
}
在底层,MsgSend 可以像 Cap’n Proto 和 FlatBuffers 那样基于某种原始内存缓冲区实现,这样我们就可以在一个版本的 MsgSend 与另一个版本之间进行无序列化转换(即零拷贝)。这种方法还有额外的好处,即允许将消息以零拷贝方式传递给使用其他语言(如 Rust)编写并通过 VM 或 FFI 访问的模块。如果我们要求所有新字段都按顺序添加,它还可以简化模块间通信中的未知字段过滤,例如只需检查是否设置了 > 5 的字段。 此外,我们也不会遇到生成类型上的状态机破坏性代码问题,因为状态机中使用的所有生成代码实际上都会存在于状态机模块本身。不过,这取决于接口类型和 protobuf Any 在其他语言中的使用方式,方案 A 中描述的 handler 方法可能仍然是可取的。无论哪种方式,实现接口的类型仍然需要像现在一样注册到 InterfaceRegistry 中,因为无法通过全局注册表获取它们。 为了简化使用 ADR 033 访问其他模块的方式,可以让客户端模块使用一个公共 API 模块(甚至可能是一个由 Buf 远程生成的模块),而不是要求在内部生成所有客户端类型。 这种方法的主要缺点是,它需要对 protobuf 类型的使用方式做出重大改变,并且需要对 protobuf 代码生成器进行大量重写。不过,这种新的生成代码仍然可以与 google.golang.org/protobuf/reflect/protoreflect API 保持兼容,以便与所有标准 golang protobuf 工具链协同工作。 如果认为修改代码生成器过于复杂,那么在 ADR 033 路由器中进行编组/反编组的朴素方案也可能是一个可接受的过渡方案。不过,既然在这种方法下所有模块大概率最终都需要迁移到 ADR 033,那么最好还是一次性完成。

方案 C) 不处理这些问题

如果认为上述方案过于复杂,我们也可以选择不做任何显式工作来提升模块版本兼容性,也不去打破循环依赖。 在这种情况下,当开发者遇到上述问题时,可以要求依赖保持同步更新(也就是我们现在的做法),或者尝试某些临时性的、可能比较 hack 的方案。 一种做法是彻底放弃 go 的语义化导入版本控制(SIV)。有人评论说,go 的 SIV(即将导入路径改为 foo/v2、foo/v3 等)限制太强,应该是可选的。golang 维护者并不同意,并且官方只支持语义化导入版本控制。不过,我们也可以采取相反的立场,通过基本上永久使用 0.x 版本策略来获得更大的灵活性。 这样,模块版本兼容性就可以通过 go.mod 中的 replace 指令来实现,将依赖固定到特定的兼容 0.x 版本。例如,如果我们知道 foo 0.2 和 0.3 都与 bar 0.3 和 0.4 兼容,那么就可以在 go.mod 中使用 replace 指令,将 foo 和 bar 固定到我们想要的版本。只要 foo 和 bar 的作者避免在这些模块之间引入不兼容的破坏性变更,这种方法就是可行的。 或者,如果开发者选择使用语义化导入版本控制,他们可以尝试前面描述的朴素方案,同时也需要使用特殊标签和 replace 指令,以确保模块被固定到正确的版本。 但需要注意的是,所有这些临时方案都会受到上文所述次版本兼容性问题的影响,除非未知字段过滤得到正确处理。

方案 D)避免在公共 API 中使用 protobuf 生成代码

另一种方案是在公共模块 API 中避免使用 protobuf 生成代码。这有助于避免在模块到模块的边界上,状态机版本与客户端 API 版本之间出现不一致。它意味着我们不会采用基于 ADR 033 的模块间消息传递,而是继续沿用现有的 keeper 方式,并进一步在 keeper 接口方法中完全避免任何 protobuf 生成代码。 采用这种方案时,我们的 foo.Keeper.DoSomething 方法将不再接收生成出来的 MsgDoSomething 结构体(它来自 protobuf API),而是改用位置参数。这样一来,为了让 foo/v2 支持 foo/v1 keeper,它只需要同时实现 v1 和 v2 的 keeper API 即可。v2 中的 DoSomething 方法可以新增 condition 参数,而 v1 中则完全不会有这个参数,因此不会出现客户端在该参数不可用时误设它的风险。 因此,这种方案可以避开次版本不兼容的问题,因为现有模块 keeper API 不会在 protobuf 文件新增字段时也跟着新增字段。 不过,采用这种方案很可能需要将所有 protobuf 生成代码都设为 internal,以防它泄漏到 keeper API 中。这意味着我们仍然需要修改 protobuf 代码生成器,使其不要把 internal/ 代码注册到全局注册表中,并且仍然需要手动注册 protobuf 的 FileDescriptor(这一点大概在所有场景下都成立)。不过,它或许可以避免把生成类型上的接口方法重构为 handler。 此外,这种方案也没有解决在模块仍然希望使用消息路由器的场景下该怎么办。无论如何,我们大概仍然希望有一种安全地把消息从一个模块传到另一个模块路由器的方式,即使只是为了 x/gov、x/authz、CosmWasm 等这类用例。这仍然需要采用方案(B)中列出的绝大多数机制,尽管我们可以建议模块在与其他模块通信时优先使用 keeper。 这种方案最大的缺点,可能在于它要求对 keeper 接口进行严格重构,以避免生成代码泄漏到 API 中。这可能导致一些情况:我们不得不复制 proto 文件中已经定义过的类型,然后再编写在 golang 版本与 protobuf 版本之间转换的方法。最终可能产生大量不必要的样板代码,而这可能会打击模块实际采用该方案、从而实现有效版本兼容的积极性。方案(A)和(B)虽然在初期更为强硬,但目标是提供一套系统,一旦采用,开发者基本就能以极少样板代码“免费”获得版本兼容性。方案(D)可能无法提供这样直接的系统,因为它要求在 protobuf API 旁边再定义一套 golang API,而这意味着类型复制,以及两套不同的设计原则(protobuf API 鼓励增量式新增变更,而 golang API 则会禁止这样做)。 这种方案的其他缺点包括:
  • 没有明确路线图来支持 Rust 等其他语言编写的模块
  • 无法让我们更接近真正的对象能力安全模型(这是 ADR 033 的目标之一)
  • 对于确实需要 ADR 033 的那部分用例,无论如何都还是必须把 ADR 033 正确落地

决策

最新的 草案 提议是:
  1. 我们已就采用 ADR 033 达成一致,这不仅是对框架的补充,而且是要从根本上替代 keeper 范式。
  2. ADR 033 的模块间路由器将根据以下规则,兼容方案(A)或(B)的任意变体: a. 如果客户端类型与服务端类型相同,则直接透传; b. 如果客户端和服务端都使用零拷贝生成代码包装器(这部分仍需定义),则把内存缓冲区从一个包装器传递到另一个包装器;或者 c. 在客户端和服务端之间对类型进行 marshal/unmarshal。
这种方案既能实现最大程度的正确性,也为未来支持使用其他语言编写、并可能运行在 WASM VM 内的模块提供了清晰路径。

次版本 API 修订

为了声明 proto 文件的次版本 API 修订,我们提议采用以下准则(这些准则此前已经记录在 cosmos.app.v1alpha module options 中):
  • 从初始版本(视为修订版本 0)开始发生修订的 proto package,应在某个 .proto 文件中包含一个 package
  • 注释,并且该注释某一行的开头应包含文本 Revision N,其中 N 是当前修订版本号。
  • 在初始修订之后的版本中新增的所有字段、消息等,都应在注释某一行的开头加入形如 Since: Revision N 的注释,其中 N 是其被引入时对应的非零修订版本号。
建议一个状态机模块与一组带版本的 proto 文件保持 1:1 对应关系,而这组 proto 文件既可以作为 buf module 进行版本管理,也可以作为 go API module 进行版本管理,或者两者兼有。如果使用 buf schema registry,那么该 buf module 的版本应始终为 1.N,其中 N 对应 package 的修订号。只有在仅更新文档注释时,才应使用补丁版本发布。在同一个按 1.N 版本管理的 buf module 中包含名为 v2、v3 等的 proto package(例如 cosmos.bank.v2)也是可以的,前提是这些 proto package 全都构成一个单一 API,并且该 API 预期由单个 SDK 模块提供服务。

次版本 API 修订的自省

为了让模块能够自省其对等模块的次版本 API 修订号,我们提议在 cosmossdk.io/core/intermodule.Client 中加入以下方法:
ServiceRevision(ctx context.Context, serviceName string)

uint64
模块可以使用 go grpc 代码生成器静态生成的服务名来调用它:
intermoduleClient.ServiceRevision(ctx, bankv1beta1.Msg_ServiceDesc.ServiceName)
未来,我们可能会决定扩展用于 protobuf service 的代码生成器,在客户端类型上增加一个字段,以更简洁地完成这项检查,例如:
package bankv1beta1

type MsgClient interface {
    Send(context.Context, MsgSend) (MsgSendResponse, error)

ServiceRevision(context.Context)

uint64
}

未知字段过滤

为了正确执行未知字段过滤,模块间路由器可以采用以下任一方式:
  • 对支持 protoreflect API 的消息使用该 API
  • 对 gogo proto 消息,执行 marshal,并使用现有的 codec/unknownproto 代码
  • 对零拷贝消息,简单检查已设置的最高字段编号(前提是我们能够要求字段必须按递增顺序连续添加)

FileDescriptor 注册

由于单个 go 二进制中可能包含同一份 protobuf 生成代码的不同版本,我们不能依赖全局 protobuf 注册表中保存正确的 FileDescriptor。由于 appconfig 模块配置本身也是用 protobuf 编写的,我们希望在加载模块本身之前,先加载该模块的 FileDescriptor。因此,我们将提供在模块实例化之前、模块注册阶段注册 FileDescriptor 的方式。针对 FileDescriptor 可能的几种打包方式,我们提议提供以下 cosmossdk.io/core/appmodule.Option 构造函数:
package appmodule

// this can be used when we are using google.golang.org/protobuf compatible generated code
// Ex:
//   ProtoFiles(bankv1beta1.File_cosmos_bank_v1beta1_module_proto)

func ProtoFiles(file []protoreflect.FileDescriptor)

Option {
}

// this can be used when we are using gogo proto generated code.
func GzippedProtoFiles(file [][]byte)

Option {
}

// this can be used when we are using buf build to generated a pinned file descriptor
func ProtoImage(protoImage []byte)

Option {
}
这种方式使我们能够支持 protobuf 文件的多种生成模式:
  • 在模块内部生成的 proto 文件(使用 ProtoFiles)
  • 带有固定 file descriptor 的 API module 方案(使用 ProtoImage)
  • gogo proto(使用 GzippedProtoFiles)

模块依赖声明

ADR 033 的一个风险是,运行时可能会调用那些并不存在于已加载 SDK 模块集合中的依赖。
此外,我们还希望模块能够定义它所需依赖 API 的最低修订版本。因此,所有模块都应预先声明自己的依赖集合。这些依赖可以在模块实例化时定义,但更理想的情况是,我们在实例化之前就知道这些依赖,并且能够静态检查某个 app config,以判断模块集合是否满足要求。例如,如果 bar 依赖 foo 的修订版本 >= 1,那么当我们创建一个同时包含两个版本 bar 和 foo 的 app config 时,就应该能够提前知道这一点。
我们提议在模块配置对象本身的 proto option 中定义这些依赖。

接口注册

我们还需要定义:对于那些会被序列化为 google.protobuf.Any 的类型,其接口方法应如何定义。考虑到希望支持其他语言编写的模块,我们可能需要思考能够适配其他语言的方案,例如 ADR 033 中简要提到的插件方案。

测试

为了确保模块确实能够兼容其依赖的多个版本,我们计划提供专门的单元测试和集成测试基础设施,以自动测试依赖的多个版本。

单元测试

单元测试应在 SDK 模块内部通过模拟其依赖来进行。在完整的 ADR 033 场景中, 这意味着与其他模块的所有交互都通过模块间路由器完成,因此对依赖的模拟 就是模拟它们的 msg 和 query server 实现。我们将同时提供测试运行器和夹具,以便让这一过程 更加顺畅。为了测试兼容性,测试运行器需要做的关键事情是测试所有依赖 API 修订版本的组合。这可以通过获取依赖的文件描述符、解析其中的注释以确定各个元素是在哪个修订版本中添加的, 然后通过减去后续才添加的元素,为每个修订版本创建合成文件描述符来完成。 下面是为单元测试运行器和夹具提议的 API:
package moduletesting

import (
    
	"context"
    "testing"
    "cosmossdk.io/core/intermodule"
    "cosmossdk.io/depinject"
    "google.golang.org/grpc"
    "google.golang.org/protobuf/proto"
    "google.golang.org/protobuf/reflect/protodesc"
)

type TestFixture interface {
    context.Context
	intermodule.Client // for making calls to the module we're testing
	BeginBlock()

EndBlock()
}

type UnitTestFixture interface {
    TestFixture
	grpc.ServiceRegistrar // for registering mock service implementations
}

type UnitTestConfig struct {
    ModuleConfig              proto.Message    // the module's config object
	DepinjectConfig           depinject.Config // optional additional depinject config options
	DependencyFileDescriptors []protodesc.FileDescriptorProto // optional dependency file descriptors to use instead of the global registry
}

// Run runs the test function for all combinations of dependency API revisions.
func (cfg UnitTestConfig)

Run(t *testing.T, f func(t *testing.T, f UnitTestFixture)) {
	// ...
}
下面是一个测试 bar 调用 foo 的示例,它利用了预期 mock 参数中的条件服务修订版本:
func TestBar(t *testing.T) {
    UnitTestConfig{
    ModuleConfig: &foomodulev1.Module{
}}.Run(t, func (t *testing.T, f moduletesting.UnitTestFixture) {
    ctrl := gomock.NewController(t)
    mockFooMsgServer := footestutil.NewMockMsgServer()

foov1.RegisterMsgServer(f, mockFooMsgServer)
    barMsgClient := barv1.NewMsgClient(f)
    if f.ServiceRevision(foov1.Msg_ServiceDesc.ServiceName) >= 1 {
    mockFooMsgServer.EXPECT().DoSomething(gomock.Any(), &foov1.MsgDoSomething{
				...,
    Condition: ..., // condition is expected in revision >= 1
}).Return(&foov1.MsgDoSomethingResponse{
}, nil)
}

else {
    mockFooMsgServer.EXPECT().DoSomething(gomock.Any(), &foov1.MsgDoSomething{...
}).Return(&foov1.MsgDoSomethingResponse{
}, nil)
}

res, err := barMsgClient.CallFoo(f, &MsgCallFoo{
})
        ...
})
}
单元测试运行器会确保依赖 mock 不会返回对于当前被测试服务修订版本无效的参数, 从而保证模块不会错误地依赖某个修订版本中并不存在的功能。

集成测试

我们还会提供一个集成测试运行器和夹具,它不使用 mock,而是测试各种组合下的真实模块依赖。下面是提议的 API:
type IntegrationTestFixture interface {
    TestFixture
}

type IntegrationTestConfig struct {
    ModuleConfig     proto.Message    // the module's config object
    DependencyMatrix map[string][]proto.Message // all the dependent module configs
}

// Run runs the test function for all combinations of dependency modules.
func (cfg IntegationTestConfig)

Run(t *testing.T, f func (t *testing.T, f IntegrationTestFixture)) {
    // ...
}
下面是一个包含 foo 和 bar 的示例:
func TestBarIntegration(t *testing.T) {
    IntegrationTestConfig{
    ModuleConfig: &barmodulev1.Module{
},
    DependencyMatrix: map[string][]proto.Message{
            "runtime": []proto.Message{ // test against two versions of runtime
                &runtimev1.Module{
},
                &runtimev2.Module{
},
},
            "foo": []proto.Message{ // test against three versions of foo
                &foomodulev1.Module{
},
                &foomodulev2.Module{
},
                &foomodulev3.Module{
},
}
 
} 
}.Run(t, func (t *testing.T, f moduletesting.IntegrationTestFixture) {
    barMsgClient := barv1.NewMsgClient(f)

res, err := barMsgClient.CallFoo(f, &MsgCallFoo{
})
        ...
})
}
与单元测试不同,集成测试会实际引入其他模块依赖。为了让模块能够在不直接依赖其他模块的情况下编写, 并且因为 golang 没有开发依赖这一概念,集成测试应写在单独的 go 模块中,例如 example.com/bar/v2/test。由于这种范式使用了 go 语义化版本控制,因此可以构建一个单独的 go 模块,同时导入 3 个版本的 bar 和 2 个版本的 runtime, 并在这六种不同的依赖组合下统一进行测试。

影响

向后兼容性

完全迁移到 ADR 033 的模块将无法与使用 keeper 范式的现有模块兼容。 作为临时变通方案,我们可能会创建一些包装类型来模拟当前的 keeper 接口,以尽量降低 迁移开销。

正面影响

  • 我们将能够交付可互操作、遵循语义化版本控制的模块,这应会显著提升 Cosmos SDK 生态系统迭代新功能的能力
  • 未来不久将有可能使用其他语言编写 Cosmos SDK 模块

负面影响

  • 所有模块都需要进行幅度相当大的重构

中性影响

  • cosmossdk.io/core/appconfig 框架将在模块定义方式上扮演更核心的角色,这 总体上很可能是件好事,但也意味着希望继续沿用 depinject 之前那种模块装配方式的用户 需要做额外改动
  • 由于完整 ADR 033 方法的存在,depinject 的必要性会有所下降,甚至可能变得不再需要。如果我们采纳 链接 中提出的 core API,那么模块大概率总是会通过 ProvideModule(appmodule.Service) (appmodule.AppModule, error) 方法来实例化自身。在这种场景下不存在 keeper 依赖的复杂装配问题,依赖注入的使用场景 可能会大幅减少,甚至完全消失。

后续讨论

上述决策目前仍被视为草案状态,尚待团队和关键利益相关方的最终认可。 如果我们确实采用这一方向,当前仍待讨论的关键问题包括:
  • 模块客户端如何探查依赖模块的 API 修订版本
  • 模块如何确定次要版本级别的依赖模块 API 修订要求
  • 模块如何恰当地测试与不同依赖版本的兼容性
  • 如何注册和解析接口实现
  • 模块如何根据其生成代码所采用的方法来注册 protobuf 文件描述符( API 模块方法仍可能作为一种受支持的策略而可行,并且会需要固定的文件描述符)

参考资料


Changelog

  • 2022-04-27: First draft

Status

DRAFT

Abstract

In order to move the Cosmos SDK to a system of decoupled semantically versioned modules which can be composed in different combinations (ex. staking v3 with bank v1 and distribution v2), we need to reassess how we organize the API surface of modules to avoid problems with go semantic import versioning and circular dependencies. This ADR explores various approaches we can take to addressing these issues.

Context

There has been a fair amount of desire in the community for semantic versioning in the SDK and there has been significant movement to splitting SDK modules into standalone go modules. Both of these will ideally allow the ecosystem to move faster because we won’t be waiting for all dependencies to update synchronously. For instance, we could have 3 versions of the core SDK compatible with the latest 2 releases of CosmWasm as well as 4 different versions of staking . This sort of setup would allow early adopters to aggressively integrate new versions, while allowing more conservative users to be selective about which versions they’re ready for. In order to achieve this, we need to solve the following problems:
  1. because of the way go semantic import versioning (SIV) works, moving to SIV naively will actually make it harder to achieve these goals
  2. circular dependencies between modules need to be broken to actually release many modules in the SDK independently
  3. pernicious minor version incompatibilities introduced through correctly evolving protobuf schemas without correct unknown field filtering
Note that all the following discussion assumes that the proto file versioning and state machine versioning of a module are distinct in that:
  • proto files are maintained in a non-breaking way (using something like buf breaking to ensure all changes are backwards compatible)
  • proto file versions get bumped much less frequently, i.e. we might maintain cosmos.bank.v1 through many versions of the bank module state machine
  • state machine breaking changes are more common and ideally this is what we’d want to semantically version with go modules, ex. x/bank/v2, x/bank/v3, etc.

Problem 1: Semantic Import Versioning Compatibility

Consider we have a module foo which defines the following MsgDoSomething and that we’ve released its state machine in go module example.com/foo:
package foo.v1;

message MsgDoSomething {
  string sender = 1;
  uint64 amount = 2;
}

service Msg {
  DoSomething(MsgDoSomething) returns (MsgDoSomethingResponse);
}
Now consider that we make a revision to this module and add a new condition field to MsgDoSomething and also add a new validation rule on amount requiring it to be non-zero, and that following go semantic versioning we release the next state machine version of foo as example.com/foo/v2.
// Revision 1
package foo.v1;

message MsgDoSomething {
  string sender = 1;
  
  // amount must be a non-zero integer.
  uint64 amount = 2;
  
  // condition is an optional condition on doing the thing.
  //
  // Since: Revision 1
  Condition condition = 3;
}
Approaching this naively, we would generate the protobuf types for the initial version of foo in example.com/foo/types and we would generate the protobuf types for the second version in example.com/foo/v2/types. Now let’s say we have a module bar which talks to foo using this keeper interface which foo provides:
type FooKeeper interface {
    DoSomething(MsgDoSomething)

error
}

Scenario A: Backward Compatibility: Newer Foo, Older Bar

Imagine we have a chain which uses both foo and bar and wants to upgrade to foo/v2, but the bar module has not upgraded to foo/v2. In this case, the chain will not be able to upgrade to foo/v2 until bar has upgraded its references to example.com/foo/types.MsgDoSomething to example.com/foo/v2/types.MsgDoSomething. Even if bar’s usage of MsgDoSomething has not changed at all, the upgrade will be impossible without this change because example.com/foo/types.MsgDoSomething and example.com/foo/v2/types.MsgDoSomething are fundamentally different incompatible structs in the go type system.

Scenario B: Forward Compatibility: Older Foo, Newer Bar

Now let’s consider the reverse scenario, where bar upgrades to foo/v2 by changing the MsgDoSomething reference to example.com/foo/v2/types.MsgDoSomething and releases that as bar/v2 with some other changes that a chain wants. The chain, however, has decided that it thinks the changes in foo/v2 are too risky and that it’d prefer to stay on the initial version of foo. In this scenario, it is impossible to upgrade to bar/v2 without upgrading to foo/v2 even if bar/v2 would have worked 100% fine with foo other than changing the import path to MsgDoSomething (meaning that bar/v2 doesn’t actually use any new features of foo/v2). Now because of the way go semantic import versioning works, we are locked into either using foo and bar OR foo/v2 and bar/v2. We cannot have foo + bar/v2 OR foo/v2 + bar. The go type system doesn’t allow this even if both versions of these modules are otherwise compatible with each other.

Naive Mitigation

A naive approach to fixing this would be to not regenerate the protobuf types in example.com/foo/v2/types but instead just update example.com/foo/types to reflect the changes needed for v2 (adding condition and requiring amount to be non-zero). Then we could release a patch of example.com/foo/types with this update and use that for foo/v2. But this change is state machine breaking for v1. It requires changing the ValidateBasic method to reject the case where amount is zero, and it adds the condition field which should be rejected based on ADR 020 unknown field filtering. So adding these changes as a patch on v1 is actually incorrect based on semantic versioning. Chains that want to stay on v1 of foo should not be importing these changes because they are incorrect for v1.

Problem 2: Circular dependencies

None of the above approaches allow foo and bar to be separate modules if for some reason foo and bar depend on each other in different ways. For instance, we can’t have foo import bar/types while bar imports foo/types. We have several cases of circular module dependencies in the SDK (ex. staking, distribution and slashing) that are legitimate from a state machine perspective. Without separating the API types out somehow, there would be no way to independently semantically version these modules without some other mitigation.

Problem 3: Handling Minor Version Incompatibilities

Imagine that we solve the first two problems but now have a scenario where bar/v2 wants the option to use MsgDoSomething.condition which only foo/v2 supports. If bar/v2 works with foo v1 and sets condition to some non-nil value, then foo will silently ignore this field resulting in a silent logic possibly dangerous logic error. If bar/v2 were able to check whether foo was on v1 or v2 and dynamically, it could choose to only use condition when foo/v2 is available. Even if bar/v2 were able to perform this check, however, how do we know that it is always performing the check properly. Without some sort of framework-level unknown field filtering, it is hard to know whether these pernicious hard to detect bugs are getting into our app and a client-server layer such as ADR 033: Inter-Module Communication may be needed to do this.

Solutions

Approach A) Separate API and State Machine Modules

One solution (first proposed in Link) is to isolate all protobuf generated code into a separate module from the state machine module. This would mean that we could have state machine go modules foo and foo/v2 which could use a types or API go module say foo/api. This foo/api go module would be perpetually on v1.x and only accept non-breaking changes. This would then allow other modules to be compatible with either foo or foo/v2 as long as the inter-module API only depends on the types in foo/api. It would also allow modules foo and bar to depend on each other in that both of them could depend on foo/api and bar/api without foo directly depending on bar and vice versa. This is similar to the naive mitigation described above except that it separates the types into separate go modules which in and of itself could be used to break circular module dependencies. It has the same problems as the naive solution, otherwise, which we could rectify by:
  1. removing all state machine breaking code from the API module (ex. ValidateBasic and any other interface methods)
  2. embedding the correct file descriptors for unknown field filtering in the binary

Migrate all interface methods on API types to handlers

To solve 1), we need to remove all interface implementations from generated types and instead use a handler approach which essentially means that given a type X, we have some sort of resolver which allows us to resolve interface implementations for that type (ex. sdk.Msg or authz.Authorization). For example:
func (k Keeper)

DoSomething(msg MsgDoSomething)

error {
    var validateBasicHandler ValidateBasicHandler
    err := k.resolver.Resolve(&validateBasic, msg)
    if err != nil {
    return err
}

err = validateBasicHandler.ValidateBasic()
	...
}
In the case of some methods on sdk.Msg, we could replace them with declarative annotations. For instance, GetSigners can already be replaced by the protobuf annotation cosmos.msg.v1.signer. In the future, we may consider some sort of protobuf validation framework (like Link but more Cosmos-specific) to replace ValidateBasic.

Pinned FileDescriptor’s

To solve 2), state machine modules must be able to specify what the version of the protobuf files was that they were built against. For instance if the API module for foo upgrades to foo/v2, the original foo module still needs a copy of the original protobuf files it was built with so that ADR 020 unknown field filtering will reject MsgDoSomething when condition is set. The simplest way to do this may be to embed the protobuf FileDescriptors into the module itself so that these FileDescriptors are used at runtime rather than the ones that are built into the foo/api which may be different. Using buf build, go embed, and a build script we can probably come up with a solution for embedding FileDescriptors into modules that is fairly straightforward.

Potential limitations to generated code

One challenge with this approach is that it places heavy restrictions on what can go in API modules and requires that most of this is state machine breaking. All or most of the code in the API module would be generated from protobuf files, so we can probably control this with how code generation is done, but it is a risk to be aware of. For instance, we do code generation for the ORM that in the future could contain optimizations that are state machine breaking. We would either need to ensure very carefully that the optimizations aren’t actually state machine breaking in generated code or separate this generated code out from the API module into the state machine module. Both of these mitigations are potentially viable but the API module approach does require an extra level of care to avoid these sorts of issues.

Minor Version Incompatibilities

This approach in and of itself does little to address any potential minor version incompatibilities and the requisite unknown field filtering. Likely some sort of client-server routing layer which does this check such as ADR 033: Inter-Module communication is required to make sure that this is done properly. We could then allow modules to perform a runtime check given a MsgClient, ex:
func (k Keeper)

CallFoo()

error {
    if k.interModuleClient.MinorRevision(k.fooMsgClient) >= 2 {
    k.fooMsgClient.DoSomething(&MsgDoSomething{
    Condition: ...
})
}

else {
        ...
}
}
To do the unknown field filtering itself, the ADR 033 router would need to use the protoreflect API to ensure that no fields unknown to the receiving module are set. This could result in an undesirable performance hit depending on how complex this logic is.

Approach B) Changes to Generated Code

An alternate approach to solving the versioning problem is to change how protobuf code is generated and move modules mostly or completely in the direction of inter-module communication as described in ADR 033. In this paradigm, a module could generate all the types it needs internally - including the API types of other modules - and talk to other modules via a client-server boundary. For instance, if bar needs to talk to foo, it could generate its own version of MsgDoSomething as bar/internal/foo/v1.MsgDoSomething and just pass this to the inter-module router which would somehow convert it to the version which foo needs (ex. foo/internal.MsgDoSomething). Currently, two generated structs for the same protobuf type cannot exist in the same go binary without special build flags (see Link). A relatively simple mitigation to this issue would be to set up the protobuf code to not register protobuf types globally if they are generated in an internal/ package. This will require modules to register their types manually with the app-level level protobuf registry, this is similar to what modules already do with the InterfaceRegistry and amino codec. If modules only do ADR 033 message passing then a naive and non-performant solution for converting bar/internal/foo/v1.MsgDoSomething to foo/internal.MsgDoSomething would be marshaling and unmarshaling in the ADR 033 router. This would break down if we needed to expose protobuf types in Keeper interfaces because the whole point is to try to keep these types internal/ so that we don’t end up with all the import version incompatibilities we’ve described above. However, because of the issue with minor version incompatibilities and the need for unknown field filtering, sticking with the Keeper paradigm instead of ADR 033 may be unviable to begin with. A more performant solution (that could maybe be adapted to work with Keeper interfaces) would be to only expose getters and setters for generated types and internally store data in memory buffers which could be passed from one implementation to another in a zero-copy way. For example, imagine this protobuf API with only getters and setters is exposed for MsgSend:
type MsgSend interface {
    proto.Message
	GetFromAddress()

string
	GetToAddress()

string
	GetAmount() []v1beta1.Coin
    SetFromAddress(string)

SetToAddress(string)

SetAmount([]v1beta1.Coin)
}

func NewMsgSend()

MsgSend {
    return &msgSendImpl{
    memoryBuffers: ...
} 
}
Under the hood, MsgSend could be implemented based on some raw memory buffer in the same way that Cap’n Proto and FlatBuffers so that we could convert between one version of MsgSend and another without serialization (i.e. zero-copy). This approach would have the added benefits of allowing zero-copy message passing to modules written in other languages such as Rust and accessed through a VM or FFI. It could also make unknown field filtering in inter-module communication simpler if we require that all new fields are added in sequential order, ex. just checking that no field > 5 is set. Also, we wouldn’t have any issues with state machine breaking code on generated types because all the generated code used in the state machine would actually live in the state machine module itself. Depending on how interface types and protobuf Anys are used in other languages, however, it may still be desirable to take the handler approach described in approach A. Either way, types implementing interfaces would still need to be registered with an InterfaceRegistry as they are now because there would be no way to retrieve them via the global registry. In order to simplify access to other modules using ADR 033, a public API module (maybe even one remotely generated by Buf) could be used by client modules instead of requiring to generate all client types internally. The big downsides of this approach are that it requires big changes to how people use protobuf types and would be a substantial rewrite of the protobuf code generator. This new generated code, however, could still be made compatible with the google.golang.org/protobuf/reflect/protoreflect API in order to work with all standard golang protobuf tooling. It is possible that the naive approach of marshaling/unmarshaling in the ADR 033 router is an acceptable intermediate solution if the changes to the code generator are seen as too complex. However, since all modules would likely need to migrate to ADR 033 anyway with this approach, it might be better to do this all at once.

Approach C) Don’t address these issues

If the above solutions are seen as too complex, we can also decide not to do anything explicit to enable better module version compatibility, and break circular dependencies. In this case, when developers are confronted with the issues described above they can require dependencies to update in sync (what we do now) or attempt some ad-hoc potentially hacky solution. One approach is to ditch go semantic import versioning (SIV) altogether. Some people have commented that go’s SIV (i.e. changing the import path to foo/v2, foo/v3, etc.) is too restrictive and that it should be optional. The golang maintainers disagree and only officially support semantic import versioning. We could, however, take the contrarian perspective and get more flexibility by using 0.x-based versioning basically forever. Module version compatibility could then be achieved using go.mod replace directives to pin dependencies to specific compatible 0.x versions. For instance if we knew foo 0.2 and 0.3 were both compatible with bar 0.3 and 0.4, we could use replace directives in our go.mod to stick to the versions of foo and bar we want. This would work as long as the authors of foo and bar avoid incompatible breaking changes between these modules. Or, if developers choose to use semantic import versioning, they can attempt the naive solution described above and would also need to use special tags and replace directives to make sure that modules are pinned to the correct versions. Note, however, that all of these ad-hoc approaches, would be vulnerable to the minor version compatibility issues described above unless unknown field filtering is properly addressed.

Approach D) Avoid protobuf generated code in public APIs

An alternative approach would be to avoid protobuf generated code in public module APIs. This would help avoid the discrepancy between state machine versions and client API versions at the module to module boundaries. It would mean that we wouldn’t do inter-module message passing based on ADR 033, but rather stick to the existing keeper approach and take it one step further by avoiding any protobuf generated code in the keeper interface methods. Using this approach, our foo.Keeper.DoSomething method wouldn’t have the generated MsgDoSomething struct (which comes from the protobuf API), but instead positional parameters. Then in order for foo/v2 to support the foo/v1 keeper it would simply need to implement both the v1 and v2 keeper APIs. The DoSomething method in v2 could have the additional condition parameter, but this wouldn’t be present in v1 at all so there would be no danger of a client accidentally setting this when it isn’t available. So this approach would avoid the challenge around minor version incompatibilities because the existing module keeper API would not get new fields when they are added to protobuf files. Taking this approach, however, would likely require making all protobuf generated code internal in order to prevent it from leaking into the keeper API. This means we would still need to modify the protobuf code generator to not register internal/ code with the global registry, and we would still need to manually register protobuf FileDescriptors (this is probably true in all scenarios). It may, however, be possible to avoid needing to refactor interface methods on generated types to handlers. Also, this approach doesn’t address what would be done in scenarios where modules still want to use the message router. Either way, we probably still want a way to pass messages from one module to another router safely even if it’s just for use cases like x/gov, x/authz, CosmWasm, etc. That would still require most of the things outlined in approach (B), although we could advise modules to prefer keepers for communicating with other modules. The biggest downside of this approach is probably that it requires a strict refactoring of keeper interfaces to avoid generated code leaking into the API. This may result in cases where we need to duplicate types that are already defined in proto files and then write methods for converting between the golang and protobuf version. This may end up in a lot of unnecessary boilerplate and that may discourage modules from actually adopting it and achieving effective version compatibility. Approaches (A) and (B), although heavy handed initially, aim to provide a system which once adopted more or less gives the developer version compatibility for free with minimal boilerplate. Approach (D) may not be able to provide such a straightforward system since it requires a golang API to be defined alongside a protobuf API in a way that requires duplication and differing sets of design principles (protobuf APIs encourage additive changes while golang APIs would forbid it). Other downsides to this approach are:
  • no clear roadmap to supporting modules in other languages like Rust
  • doesn’t get us any closer to proper object capability security (one of the goals of ADR 033)
  • ADR 033 needs to be done properly anyway for the set of use cases which do need it

Decision

The latest DRAFT proposal is:
  1. we are alignment on adopting ADR 033 not just as an addition to the framework, but as a core replacement to the keeper paradigm entirely.
  2. the ADR 033 inter-module router will accommodate any variation of approach (A) or (B) given the following rules: a. if the client type is the same as the server type then pass it directly through, b. if both client and server use the zero-copy generated code wrappers (which still need to be defined), then pass the memory buffers from one wrapper to the other, or c. marshal/unmarshal types between client and server.
This approach will allow for both maximal correctness and enable a clear path to enabling modules within in other languages, possibly executed within a WASM VM.

Minor API Revisions

To declare minor API revisions of proto files, we propose the following guidelines (which were already documented in cosmos.app.v1alpha module options):
  • proto packages which are revised from their initial version (considered revision 0) should include a package
  • comment in some .proto file containing the test Revision N at the start of a comment line where N is the current revision number.
  • all fields, messages, etc. added in a version beyond the initial revision should add a comment at the start of a comment line of the form Since: Revision N where N is the non-zero revision it was added.
It is advised that there is a 1:1 correspondence between a state machine module and versioned set of proto files which are versioned either as a buf module a go API module or both. If the buf schema registry is used, the version of this buf module should always be 1.N where N corresponds to the package revision. Patch releases should be used when only documentation comments are updated. It is okay to include proto packages named v2, v3, etc. in this same 1.N versioned buf module (ex. cosmos.bank.v2) as long as all these proto packages consist of a single API intended to be served by a single SDK module.

Introspecting Minor API Revisions

In order for modules to introspect the minor API revision of peer modules, we propose adding the following method to cosmossdk.io/core/intermodule.Client:
ServiceRevision(ctx context.Context, serviceName string)

uint64
Modules could all this using the service name statically generated by the go grpc code generator:
intermoduleClient.ServiceRevision(ctx, bankv1beta1.Msg_ServiceDesc.ServiceName)
In the future, we may decide to extend the code generator used for protobuf services to add a field to client types which does this check more concisely, ex:
package bankv1beta1

type MsgClient interface {
    Send(context.Context, MsgSend) (MsgSendResponse, error)

ServiceRevision(context.Context)

uint64
}

Unknown Field Filtering

To correctly perform unknown field filtering, the inter-module router can do one of the following:
  • use the protoreflect API for messages which support that
  • for gogo proto messages, marshal and use the existing codec/unknownproto code
  • for zero-copy messages, do a simple check on the highest set field number (assuming we can require that fields are adding consecutively in increasing order)

FileDescriptor Registration

Because a single go binary may contain different versions of the same generated protobuf code, we cannot rely on the global protobuf registry to contain the correct FileDescriptors. Because appconfig module configuration is itself written in protobuf, we would like to load the FileDescriptors for a module before loading a module itself. So we will provide ways to register FileDescriptors at module registration time before instantiation. We propose the following cosmossdk.io/core/appmodule.Option constructors for the various cases of how FileDescriptors may be packaged:
package appmodule

// this can be used when we are using google.golang.org/protobuf compatible generated code
// Ex:
//   ProtoFiles(bankv1beta1.File_cosmos_bank_v1beta1_module_proto)

func ProtoFiles(file []protoreflect.FileDescriptor)

Option {
}

// this can be used when we are using gogo proto generated code.
func GzippedProtoFiles(file [][]byte)

Option {
}

// this can be used when we are using buf build to generated a pinned file descriptor
func ProtoImage(protoImage []byte)

Option {
}
This approach allows us to support several ways protobuf files might be generated:
  • proto files generated internally to a module (use ProtoFiles)
  • the API module approach with pinned file descriptors (use ProtoImage)
  • gogo proto (use GzippedProtoFiles)

Module Dependency Declaration

One risk of ADR 033 is that dependencies are called at runtime which are not present in the loaded set of SDK modules.
Also we want modules to have a way to define a minimum dependency API revision that they require. Therefore, all modules should declare their set of dependencies upfront. These dependencies could be defined when a module is instantiated, but ideally we know what the dependencies are before instantiation and can statically look at an app config and determine whether the set of modules. For example, if bar requires foo revision >= 1, then we should be able to know this when creating an app config with two versions of bar and foo.
We propose defining these dependencies in the proto options of the module config object itself.

Interface Registration

We will also need to define how interface methods are defined on types that are serialized as google.protobuf.Any’s. In light of the desire to support modules in other languages, we may want to think of solutions that will accommodate other languages such as plugins described briefly in ADR 033.

Testing

In order to ensure that modules are indeed with multiple versions of their dependencies, we plan to provide specialized unit and integration testing infrastructure that automatically tests multiple versions of dependencies.

Unit Testing

Unit tests should be conducted inside SDK modules by mocking their dependencies. In a full ADR 033 scenario, this means that all interaction with other modules is done via the inter-module router, so mocking of dependencies means mocking their msg and query server implementations. We will provide both a test runner and fixture to make this streamlined. The key thing that the test runner should do to test compatibility is to test all combinations of dependency API revisions. This can be done by taking the file descriptors for the dependencies, parsing their comments to determine the revisions various elements were added, and then created synthetic file descriptors for each revision by subtracting elements that were added later. Here is a proposed API for the unit test runner and fixture:
package moduletesting

import (
    
	"context"
    "testing"
    "cosmossdk.io/core/intermodule"
    "cosmossdk.io/depinject"
    "google.golang.org/grpc"
    "google.golang.org/protobuf/proto"
    "google.golang.org/protobuf/reflect/protodesc"
)

type TestFixture interface {
    context.Context
	intermodule.Client // for making calls to the module we're testing
	BeginBlock()

EndBlock()
}

type UnitTestFixture interface {
    TestFixture
	grpc.ServiceRegistrar // for registering mock service implementations
}

type UnitTestConfig struct {
    ModuleConfig              proto.Message    // the module's config object
	DepinjectConfig           depinject.Config // optional additional depinject config options
	DependencyFileDescriptors []protodesc.FileDescriptorProto // optional dependency file descriptors to use instead of the global registry
}

// Run runs the test function for all combinations of dependency API revisions.
func (cfg UnitTestConfig)

Run(t *testing.T, f func(t *testing.T, f UnitTestFixture)) {
	// ...
}
Here is an example for testing bar calling foo which takes advantage of conditional service revisions in the expected mock arguments:
func TestBar(t *testing.T) {
    UnitTestConfig{
    ModuleConfig: &foomodulev1.Module{
}}.Run(t, func (t *testing.T, f moduletesting.UnitTestFixture) {
    ctrl := gomock.NewController(t)
    mockFooMsgServer := footestutil.NewMockMsgServer()

foov1.RegisterMsgServer(f, mockFooMsgServer)
    barMsgClient := barv1.NewMsgClient(f)
    if f.ServiceRevision(foov1.Msg_ServiceDesc.ServiceName) >= 1 {
    mockFooMsgServer.EXPECT().DoSomething(gomock.Any(), &foov1.MsgDoSomething{
				...,
    Condition: ..., // condition is expected in revision >= 1
}).Return(&foov1.MsgDoSomethingResponse{
}, nil)
}

else {
    mockFooMsgServer.EXPECT().DoSomething(gomock.Any(), &foov1.MsgDoSomething{...
}).Return(&foov1.MsgDoSomethingResponse{
}, nil)
}

res, err := barMsgClient.CallFoo(f, &MsgCallFoo{
})
        ...
})
}
The unit test runner would make sure that no dependency mocks return arguments which are invalid for the service revision being tested to ensure that modules don’t incorrectly depend on functionality not present in a given revision.

Integration Testing

An integration test runner and fixture would also be provided which instead of using mocks would test actual module dependencies in various combinations. Here is the proposed API:
type IntegrationTestFixture interface {
    TestFixture
}

type IntegrationTestConfig struct {
    ModuleConfig     proto.Message    // the module's config object
    DependencyMatrix map[string][]proto.Message // all the dependent module configs
}

// Run runs the test function for all combinations of dependency modules.
func (cfg IntegationTestConfig)

Run(t *testing.T, f func (t *testing.T, f IntegrationTestFixture)) {
    // ...
}
And here is an example with foo and bar:
func TestBarIntegration(t *testing.T) {
    IntegrationTestConfig{
    ModuleConfig: &barmodulev1.Module{
},
    DependencyMatrix: map[string][]proto.Message{
            "runtime": []proto.Message{ // test against two versions of runtime
                &runtimev1.Module{
},
                &runtimev2.Module{
},
},
            "foo": []proto.Message{ // test against three versions of foo
                &foomodulev1.Module{
},
                &foomodulev2.Module{
},
                &foomodulev3.Module{
},
}
 
} 
}.Run(t, func (t *testing.T, f moduletesting.IntegrationTestFixture) {
    barMsgClient := barv1.NewMsgClient(f)

res, err := barMsgClient.CallFoo(f, &MsgCallFoo{
})
        ...
})
}
Unlike unit tests, integration tests actually pull in other module dependencies. So that modules can be written without direct dependencies on other modules and because golang has no concept of development dependencies, integration tests should be written in separate go modules, ex. example.com/bar/v2/test. Because this paradigm uses go semantic versioning, it is possible to build a single go module which imports 3 versions of bar and 2 versions of runtime and can test these all together in the six various combinations of dependencies.

Consequences

Backwards Compatibility

Modules which migrate fully to ADR 033 will not be compatible with existing modules which use the keeper paradigm. As a temporary workaround we may create some wrapper types that emulate the current keeper interface to minimize the migration overhead.

Positive

  • we will be able to deliver interoperable semantically versioned modules which should dramatically increase the ability of the Cosmos SDK ecosystem to iterate on new features
  • it will be possible to write Cosmos SDK modules in other languages in the near future

Negative

  • all modules will need to be refactored somewhat dramatically

Neutral

  • the cosmossdk.io/core/appconfig framework will play a more central role in terms of how modules are defined, this is likely generally a good thing but does mean additional changes for users wanting to stick to the pre-depinject way of wiring up modules
  • depinject is somewhat less needed or maybe even obviated because of the full ADR 033 approach. If we adopt the core API proposed in Link, then a module would probably always instantiate itself with a method ProvideModule(appmodule.Service) (appmodule.AppModule, error). There is no complex wiring of keeper dependencies in this scenario and dependency injection may not have as much of (or any) use case.

Further Discussions

The decision described above is considered in draft mode and is pending final buy-in from the team and key stakeholders. Key outstanding discussions if we do adopt that direction are:
  • how do module clients introspect dependency module API revisions
  • how do modules determine a minor dependency module API revision requirement
  • how do modules appropriately test compatibility with different dependency versions
  • how to register and resolve interface implementations
  • how do modules register their protobuf file descriptors depending on the approach they take to generated code (the API module approach may still be viable as a supported strategy and would need pinned file descriptors)

References