变更日志
- 2022-04-27:第一版草稿
状态
草案摘要
为了将 Cosmos SDK 迁移到一种彼此解耦、遵循语义化版本控制的模块体系,这些模块可以按不同组合方式进行编排(例如 staking v3 搭配 bank v1 和 distribution v2),我们需要重新评估模块 API 暴露面的组织方式,以避免 Go 语义化导入版本控制以及循环依赖带来的问题。本文 ADR 将探讨我们可以采用的多种方案来解决这些问题。背景
社区一直有相当多的诉求,希望 SDK 支持语义化版本控制,同时也在积极推动将 SDK 模块拆分为独立的 Go 模块。理想情况下,这两项工作都能让生态系统迭代得更快,因为我们不再需要等待所有依赖同步更新。比如,我们可以拥有 3 个核心 SDK 版本,它们兼容最新的 2 个 CosmWasm 发行版,以及 4 个不同版本的 staking。这种模式能让早期采用者更激进地集成新版本,同时也让更保守的用户可以自行选择准备好采用哪些版本。 为了实现这一点,我们需要解决以下问题:- 由于 Go 语义化导入版本控制(SIV)的工作方式,如果直接、天真地迁移到 SIV,实际上会让这些目标更难达成
- 模块之间的循环依赖必须被打破,SDK 中的许多模块才能真正独立发布
- 即使 protobuf schema 是按正确方式演进的,如果没有正确的未知字段过滤,仍会引入隐蔽的次版本不兼容问题
- proto 文件以非破坏性方式维护(使用类似 buf breaking 的工具来确保所有变更都向后兼容)
- proto 文件版本的提升频率要低得多,也就是说,我们可能会在 bank 模块状态机经历多个版本期间持续维护
cosmos.bank.v1 - 状态机的破坏性变更更为常见,而理想情况下,我们希望通过 Go 模块对这部分做语义化版本控制,例如
x/bank/v2、x/bank/v3等
问题 1:与语义化导入版本控制的兼容性
假设我们有一个模块foo,它定义了如下 MsgDoSomething,并且我们已将其状态机作为 Go 模块 example.com/foo 发布:
MsgDoSomething 新增了一个 condition 字段,同时新增了一条针对 amount 的校验规则,要求它必须非零;并且遵循 Go 语义化版本控制的约定,将 foo 的下一个状态机版本发布为 example.com/foo/v2。
foo 初始版本的 protobuf 类型生成到 example.com/foo/types 中,并将第二个版本的 protobuf 类型生成到 example.com/foo/v2/types 中。
现在假设我们有一个模块 bar,它通过 foo 提供的这个 keeper 接口与 foo 通信:
场景 A:向后兼容:较新的 Foo,较旧的 Bar
设想我们有一条链同时使用foo 和 bar,并且希望升级到 foo/v2,但 bar 模块还没有升级到 foo/v2。
在这种情况下,这条链将无法升级到 foo/v2,除非 bar 先把它对 example.com/foo/types.MsgDoSomething 的引用升级为 example.com/foo/v2/types.MsgDoSomething。
即使 bar 对 MsgDoSomething 的使用方式完全没有变化,这个升级依然不可能完成,因为在 Go 类型系统中,example.com/foo/types.MsgDoSomething 和 example.com/foo/v2/types.MsgDoSomething 本质上是两种不同且互不兼容的结构体。
场景 B:向前兼容:较旧的 Foo,较新的 Bar
现在考虑相反的场景:bar 通过将 MsgDoSomething 的引用改成 example.com/foo/v2/types.MsgDoSomething 来升级到 foo/v2,并将其连同链所需的其他一些变更一起作为 bar/v2 发布。然而,这条链认为 foo/v2 中的变更风险过高,因此更希望继续停留在 foo 的初始版本上。
在这种场景下,即使 bar/v2 除了把 MsgDoSomething 的导入路径改掉之外,与 foo 的配合本来可以 100% 正常工作(也就是说,bar/v2 实际上并没有使用 foo/v2 的任何新特性),也不可能在不升级到 foo/v2 的前提下升级到 bar/v2。
由于 Go 语义化导入版本控制的工作方式,我们最终被锁定在只能使用 foo 和 bar,或者 foo/v2 和 bar/v2 这两种组合之一。我们无法得到 foo + bar/v2,也无法得到 foo/v2 + bar。即使这些模块的两个版本在其他方面彼此兼容,Go 类型系统也不允许这样做。
朴素缓解方案
一种朴素的修复思路是:不要在example.com/foo/v2/types 中重新生成 protobuf 类型,而是直接更新 example.com/foo/types,让它反映 v2 所需的变更(新增 condition 并要求 amount 非零)。然后我们可以发布一个包含这些更新的 example.com/foo/types 补丁版本,并将其用于 foo/v2。但这个变更对 v1 来说会破坏状态机兼容性。它要求修改 ValidateBasic 方法,以拒绝 amount 为零的情况;同时它还新增了 condition 字段,而依据ADR 020 未知字段过滤,这个字段本应被拒绝。因此,把这些变更作为 v1 的补丁实际上违背了语义化版本控制。那些希望继续停留在 foo 的 v1 上的链,不应该导入这些变更,因为它们对于 v1 来说是不正确的。
问题 2:循环依赖
如果foo 和 bar 因某些原因以不同方式互相依赖,那么上述任何方案都无法让它们成为彼此独立的模块。比如,我们不能让 foo 导入 bar/types,同时又让 bar 导入 foo/types。
SDK 中已经存在多种模块循环依赖的情况(例如 staking、distribution 和 slashing),从状态机角度看,这些依赖是合理的。如果不以某种方式把 API 类型拆分出去,那么在没有其他缓解手段的前提下,就无法对这些模块进行彼此独立的语义化版本控制。
问题 3:处理次版本不兼容
假设我们已经解决了前两个问题,但现在出现这样一种情况:bar/v2 希望能够选择性地使用只有 foo/v2 才支持的 MsgDoSomething.condition。如果 bar/v2 在与 foo v1 配合时,把 condition 设置为某个非 nil 值,那么 foo 会静默忽略这个字段,从而导致一种静默的逻辑错误,这种错误可能具有危险性。假如 bar/v2 能够检查 foo 当前运行的是 v1 还是 v2,并动态地仅在 foo/v2 可用时才使用 condition,它就可以避免这个问题。然而,即便 bar/v2 能执行这项检查,我们又如何确保它总能正确执行?如果没有某种框架层面的未知字段过滤,就很难判断这些隐蔽且难以检测的缺陷是否已经进入我们的应用。而要实现这一点,可能需要类似ADR 033:模块间通信这样的客户端-服务端层。
解决方案
方案 A)分离 API 模块与状态机模块
一种方案(最早在链接中提出)是将所有 protobuf 生成代码从状态机模块中拆分出来,放入一个单独的模块中。这意味着我们可以有状态机 Go 模块foo 和 foo/v2,它们共同使用一个类型或 API Go 模块,比如 foo/api。这个 foo/api Go 模块将始终停留在 v1.x,并且只接受非破坏性变更。这样一来,只要模块间 API 仅依赖 foo/api 中的类型,其他模块就可以同时兼容 foo 或 foo/v2。这也会让 foo 和 bar 两个模块能够互相依赖,因为它们都可以依赖 foo/api 和 bar/api,而不需要 foo 直接依赖 bar,反之亦然。
这与前面描述的朴素缓解方案类似,不同之处在于它将类型拆分到了独立的 Go 模块中,而这本身就可以用来打破模块之间的循环依赖。除此之外,它仍然存在与朴素方案相同的问题,我们可以通过以下方式进行修正:
- 从 API 模块中移除所有会破坏状态机兼容性的代码(例如
ValidateBasic以及其他任何接口方法) - 在二进制中嵌入用于未知字段过滤的正确文件描述符
将 API 类型上的所有接口方法迁移为处理器
为了解决第 1 点,我们需要移除生成类型上的所有接口实现,转而使用处理器模式。其本质是,对于类型X,我们提供某种解析器,用于解析该类型对应的接口实现(例如 sdk.Msg 或 authz.Authorization)。例如:
sdk.Msg 上的某些方法,我们可以用声明式注解来替代。例如,GetSigners 已经可以由 protobuf 注解 cosmos.msg.v1.signer 替代。未来,我们还可以考虑某种 protobuf 校验框架(类似 链接,但更贴合 Cosmos 的需求)来替代 ValidateBasic。
固定的 FileDescriptor
为了解决第 2 点,状态机模块必须能够指定其构建时所依赖的 protobuf 文件版本。比如,如果foo 的 API 模块升级到 foo/v2,原始的 foo 模块仍然需要保留一份它构建时所使用的原始 protobuf 文件副本,这样 ADR 020 的未知字段过滤才能在 condition 被设置时拒绝 MsgDoSomething。
最简单的方式,可能是将 protobuf 的 FileDescriptor 直接嵌入模块本身,以便运行时使用这些 FileDescriptor,而不是使用可能已经发生变化、内置在 foo/api 中的那一份。借助 buf build、go embed 以及构建脚本,我们大概率可以设计出一种相当直接的方案,把 FileDescriptor 嵌入模块之中。
生成代码的潜在限制
这种方法的一个挑战在于,它对 API 模块中可以包含的内容施加了严格限制,并且要求其中大部分内容都属于会破坏状态机的变更。API 模块中的全部或大部分代码都会从 protobuf 文件生成,因此我们大概可以通过控制代码生成方式来约束这一点,但这仍然是一个需要注意的风险。 例如,我们会为 ORM 做代码生成,而未来其中可能包含会破坏状态机的优化。我们要么需要非常谨慎地确保这些优化在生成代码中实际上不会破坏状态机,要么将这部分生成代码从 API 模块中拆分出来,放入状态机模块。这两种缓解方式都可能可行,但 API 模块方案确实需要额外谨慎,以避免这类问题。次版本不兼容
这种方法本身几乎无法解决潜在的次版本不兼容问题,以及所需的未知字段过滤。很可能需要某种执行此检查的客户端-服务端路由层,例如 ADR 033:模块间通信,才能确保这件事被正确处理。这样一来,我们就可以允许模块在运行时基于MsgClient 执行检查,例如:
方案 B) 修改生成代码
解决版本问题的另一种方法是改变 protobuf 代码的生成方式,并让模块大部分或完全转向 ADR 033 中描述的模块间通信方向。在这种范式下,一个模块可以在内部生成它所需的所有类型,包括其他模块的 API 类型,并通过客户端-服务端边界与其他模块通信。例如,如果bar 需要与 foo 通信,它可以将自己的 MsgDoSomething 版本生成为 bar/internal/foo/v1.MsgDoSomething,然后将其传递给模块间路由器,由后者以某种方式将其转换为 foo 所需的版本(例如 foo/internal.MsgDoSomething)。
当前,在同一个 go 二进制中,如果没有特殊的构建标志,就不能同时存在针对同一个 protobuf 类型生成的两个结构体(见链接)。对此,一个相对简单的缓解方案是设置 protobuf 代码,使其在 internal/ 包中生成时不全局注册 protobuf 类型。这将要求模块在应用级 protobuf 注册表中手动注册其类型,这与模块当前对 InterfaceRegistry 和 amino codec 所做的事情类似。
如果模块只通过 ADR 033 进行消息传递,那么将 bar/internal/foo/v1.MsgDoSomething 转换为 foo/internal.MsgDoSomething 的一种朴素且性能不佳的方案,就是在 ADR 033 路由器中进行编组和反编组。如果我们需要在 Keeper 接口中暴露 protobuf 类型,这种做法就行不通了,因为这里的核心目标正是尽量将这些类型保留在 internal/ 中,从而避免前面描述的各种导入版本不兼容问题。不过,考虑到次版本不兼容的问题以及未知字段过滤的需求,一开始就继续坚持 Keeper 范式而不是 ADR 033,可能并不可行。
一种性能更好的方案是(也许还能调整为适用于 Keeper 接口)只暴露生成类型的 getter 和 setter,并在内部将数据存储在内存缓冲区中,这样就可以以零拷贝的方式在不同实现之间传递。
例如,假设针对 MsgSend 暴露了下面这个仅包含 getter 和 setter 的 protobuf API:
MsgSend 可以像 Cap’n Proto 和 FlatBuffers 那样基于某种原始内存缓冲区实现,这样我们就可以在一个版本的 MsgSend 与另一个版本之间进行无序列化转换(即零拷贝)。这种方法还有额外的好处,即允许将消息以零拷贝方式传递给使用其他语言(如 Rust)编写并通过 VM 或 FFI 访问的模块。如果我们要求所有新字段都按顺序添加,它还可以简化模块间通信中的未知字段过滤,例如只需检查是否设置了 > 5 的字段。
此外,我们也不会遇到生成类型上的状态机破坏性代码问题,因为状态机中使用的所有生成代码实际上都会存在于状态机模块本身。不过,这取决于接口类型和 protobuf Any 在其他语言中的使用方式,方案 A 中描述的 handler 方法可能仍然是可取的。无论哪种方式,实现接口的类型仍然需要像现在一样注册到 InterfaceRegistry 中,因为无法通过全局注册表获取它们。
为了简化使用 ADR 033 访问其他模块的方式,可以让客户端模块使用一个公共 API 模块(甚至可能是一个由 Buf 远程生成的模块),而不是要求在内部生成所有客户端类型。
这种方法的主要缺点是,它需要对 protobuf 类型的使用方式做出重大改变,并且需要对 protobuf 代码生成器进行大量重写。不过,这种新的生成代码仍然可以与 google.golang.org/protobuf/reflect/protoreflect API 保持兼容,以便与所有标准 golang protobuf 工具链协同工作。
如果认为修改代码生成器过于复杂,那么在 ADR 033 路由器中进行编组/反编组的朴素方案也可能是一个可接受的过渡方案。不过,既然在这种方法下所有模块大概率最终都需要迁移到 ADR 033,那么最好还是一次性完成。
方案 C) 不处理这些问题
如果认为上述方案过于复杂,我们也可以选择不做任何显式工作来提升模块版本兼容性,也不去打破循环依赖。 在这种情况下,当开发者遇到上述问题时,可以要求依赖保持同步更新(也就是我们现在的做法),或者尝试某些临时性的、可能比较 hack 的方案。 一种做法是彻底放弃 go 的语义化导入版本控制(SIV)。有人评论说,go 的 SIV(即将导入路径改为foo/v2、foo/v3 等)限制太强,应该是可选的。golang 维护者并不同意,并且官方只支持语义化导入版本控制。不过,我们也可以采取相反的立场,通过基本上永久使用 0.x 版本策略来获得更大的灵活性。
这样,模块版本兼容性就可以通过 go.mod 中的 replace 指令来实现,将依赖固定到特定的兼容 0.x 版本。例如,如果我们知道 foo 0.2 和 0.3 都与 bar 0.3 和 0.4 兼容,那么就可以在 go.mod 中使用 replace 指令,将 foo 和 bar 固定到我们想要的版本。只要 foo 和 bar 的作者避免在这些模块之间引入不兼容的破坏性变更,这种方法就是可行的。
或者,如果开发者选择使用语义化导入版本控制,他们可以尝试前面描述的朴素方案,同时也需要使用特殊标签和 replace 指令,以确保模块被固定到正确的版本。
但需要注意的是,所有这些临时方案都会受到上文所述次版本兼容性问题的影响,除非未知字段过滤得到正确处理。
方案 D)避免在公共 API 中使用 protobuf 生成代码
另一种方案是在公共模块 API 中避免使用 protobuf 生成代码。这有助于避免在模块到模块的边界上,状态机版本与客户端 API 版本之间出现不一致。它意味着我们不会采用基于 ADR 033 的模块间消息传递,而是继续沿用现有的 keeper 方式,并进一步在 keeper 接口方法中完全避免任何 protobuf 生成代码。 采用这种方案时,我们的foo.Keeper.DoSomething 方法将不再接收生成出来的 MsgDoSomething 结构体(它来自 protobuf API),而是改用位置参数。这样一来,为了让 foo/v2 支持 foo/v1 keeper,它只需要同时实现 v1 和 v2 的 keeper API 即可。v2 中的 DoSomething 方法可以新增 condition 参数,而 v1 中则完全不会有这个参数,因此不会出现客户端在该参数不可用时误设它的风险。
因此,这种方案可以避开次版本不兼容的问题,因为现有模块 keeper API 不会在 protobuf 文件新增字段时也跟着新增字段。
不过,采用这种方案很可能需要将所有 protobuf 生成代码都设为 internal,以防它泄漏到 keeper API 中。这意味着我们仍然需要修改 protobuf 代码生成器,使其不要把 internal/ 代码注册到全局注册表中,并且仍然需要手动注册 protobuf 的 FileDescriptor(这一点大概在所有场景下都成立)。不过,它或许可以避免把生成类型上的接口方法重构为 handler。
此外,这种方案也没有解决在模块仍然希望使用消息路由器的场景下该怎么办。无论如何,我们大概仍然希望有一种安全地把消息从一个模块传到另一个模块路由器的方式,即使只是为了 x/gov、x/authz、CosmWasm 等这类用例。这仍然需要采用方案(B)中列出的绝大多数机制,尽管我们可以建议模块在与其他模块通信时优先使用 keeper。
这种方案最大的缺点,可能在于它要求对 keeper 接口进行严格重构,以避免生成代码泄漏到 API 中。这可能导致一些情况:我们不得不复制 proto 文件中已经定义过的类型,然后再编写在 golang 版本与 protobuf 版本之间转换的方法。最终可能产生大量不必要的样板代码,而这可能会打击模块实际采用该方案、从而实现有效版本兼容的积极性。方案(A)和(B)虽然在初期更为强硬,但目标是提供一套系统,一旦采用,开发者基本就能以极少样板代码“免费”获得版本兼容性。方案(D)可能无法提供这样直接的系统,因为它要求在 protobuf API 旁边再定义一套 golang API,而这意味着类型复制,以及两套不同的设计原则(protobuf API 鼓励增量式新增变更,而 golang API 则会禁止这样做)。
这种方案的其他缺点包括:
- 没有明确路线图来支持 Rust 等其他语言编写的模块
- 无法让我们更接近真正的对象能力安全模型(这是 ADR 033 的目标之一)
- 对于确实需要 ADR 033 的那部分用例,无论如何都还是必须把 ADR 033 正确落地
决策
最新的 草案 提议是:- 我们已就采用 ADR 033 达成一致,这不仅是对框架的补充,而且是要从根本上替代 keeper 范式。
- ADR 033 的模块间路由器将根据以下规则,兼容方案(A)或(B)的任意变体: a. 如果客户端类型与服务端类型相同,则直接透传; b. 如果客户端和服务端都使用零拷贝生成代码包装器(这部分仍需定义),则把内存缓冲区从一个包装器传递到另一个包装器;或者 c. 在客户端和服务端之间对类型进行 marshal/unmarshal。
次版本 API 修订
为了声明 proto 文件的次版本 API 修订,我们提议采用以下准则(这些准则此前已经记录在 cosmos.app.v1alpha module options 中):- 从初始版本(视为修订版本
0)开始发生修订的 proto package,应在某个 .proto 文件中包含一个package - 注释,并且该注释某一行的开头应包含文本
Revision N,其中N是当前修订版本号。 - 在初始修订之后的版本中新增的所有字段、消息等,都应在注释某一行的开头加入形如
Since: Revision N的注释,其中N是其被引入时对应的非零修订版本号。
1.N,其中 N 对应 package 的修订号。只有在仅更新文档注释时,才应使用补丁版本发布。在同一个按 1.N 版本管理的 buf module 中包含名为 v2、v3 等的 proto package(例如 cosmos.bank.v2)也是可以的,前提是这些 proto package 全都构成一个单一 API,并且该 API 预期由单个 SDK 模块提供服务。
次版本 API 修订的自省
为了让模块能够自省其对等模块的次版本 API 修订号,我们提议在cosmossdk.io/core/intermodule.Client 中加入以下方法:
未知字段过滤
为了正确执行未知字段过滤,模块间路由器可以采用以下任一方式:- 对支持
protoreflectAPI 的消息使用该 API - 对 gogo proto 消息,执行 marshal,并使用现有的
codec/unknownproto代码 - 对零拷贝消息,简单检查已设置的最高字段编号(前提是我们能够要求字段必须按递增顺序连续添加)
FileDescriptor 注册
由于单个 go 二进制中可能包含同一份 protobuf 生成代码的不同版本,我们不能依赖全局 protobuf 注册表中保存正确的 FileDescriptor。由于 appconfig 模块配置本身也是用 protobuf 编写的,我们希望在加载模块本身之前,先加载该模块的 FileDescriptor。因此,我们将提供在模块实例化之前、模块注册阶段注册 FileDescriptor 的方式。针对 FileDescriptor 可能的几种打包方式,我们提议提供以下 cosmossdk.io/core/appmodule.Option 构造函数:
- 在模块内部生成的 proto 文件(使用
ProtoFiles) - 带有固定 file descriptor 的 API module 方案(使用
ProtoImage) - gogo proto(使用
GzippedProtoFiles)
模块依赖声明
ADR 033 的一个风险是,运行时可能会调用那些并不存在于已加载 SDK 模块集合中的依赖。此外,我们还希望模块能够定义它所需依赖 API 的最低修订版本。因此,所有模块都应预先声明自己的依赖集合。这些依赖可以在模块实例化时定义,但更理想的情况是,我们在实例化之前就知道这些依赖,并且能够静态检查某个 app config,以判断模块集合是否满足要求。例如,如果
bar 依赖 foo 的修订版本 >= 1,那么当我们创建一个同时包含两个版本 bar 和 foo 的 app config 时,就应该能够提前知道这一点。
我们提议在模块配置对象本身的 proto option 中定义这些依赖。
接口注册
我们还需要定义:对于那些会被序列化为google.protobuf.Any 的类型,其接口方法应如何定义。考虑到希望支持其他语言编写的模块,我们可能需要思考能够适配其他语言的方案,例如 ADR 033 中简要提到的插件方案。
测试
为了确保模块确实能够兼容其依赖的多个版本,我们计划提供专门的单元测试和集成测试基础设施,以自动测试依赖的多个版本。单元测试
单元测试应在 SDK 模块内部通过模拟其依赖来进行。在完整的 ADR 033 场景中, 这意味着与其他模块的所有交互都通过模块间路由器完成,因此对依赖的模拟 就是模拟它们的 msg 和 query server 实现。我们将同时提供测试运行器和夹具,以便让这一过程 更加顺畅。为了测试兼容性,测试运行器需要做的关键事情是测试所有依赖 API 修订版本的组合。这可以通过获取依赖的文件描述符、解析其中的注释以确定各个元素是在哪个修订版本中添加的, 然后通过减去后续才添加的元素,为每个修订版本创建合成文件描述符来完成。 下面是为单元测试运行器和夹具提议的 API:集成测试
我们还会提供一个集成测试运行器和夹具,它不使用 mock,而是测试各种组合下的真实模块依赖。下面是提议的 API:example.com/bar/v2/test。由于这种范式使用了 go 语义化版本控制,因此可以构建一个单独的 go 模块,同时导入 3 个版本的 bar 和 2 个版本的 runtime,
并在这六种不同的依赖组合下统一进行测试。
影响
向后兼容性
完全迁移到 ADR 033 的模块将无法与使用 keeper 范式的现有模块兼容。 作为临时变通方案,我们可能会创建一些包装类型来模拟当前的 keeper 接口,以尽量降低 迁移开销。正面影响
- 我们将能够交付可互操作、遵循语义化版本控制的模块,这应会显著提升 Cosmos SDK 生态系统迭代新功能的能力
- 未来不久将有可能使用其他语言编写 Cosmos SDK 模块
负面影响
- 所有模块都需要进行幅度相当大的重构
中性影响
cosmossdk.io/core/appconfig框架将在模块定义方式上扮演更核心的角色,这 总体上很可能是件好事,但也意味着希望继续沿用 depinject 之前那种模块装配方式的用户 需要做额外改动- 由于完整 ADR 033 方法的存在,
depinject的必要性会有所下降,甚至可能变得不再需要。如果我们采纳 链接 中提出的 core API,那么模块大概率总是会通过ProvideModule(appmodule.Service) (appmodule.AppModule, error)方法来实例化自身。在这种场景下不存在 keeper 依赖的复杂装配问题,依赖注入的使用场景 可能会大幅减少,甚至完全消失。
后续讨论
上述决策目前仍被视为草案状态,尚待团队和关键利益相关方的最终认可。 如果我们确实采用这一方向,当前仍待讨论的关键问题包括:- 模块客户端如何探查依赖模块的 API 修订版本
- 模块如何确定次要版本级别的依赖模块 API 修订要求
- 模块如何恰当地测试与不同依赖版本的兼容性
- 如何注册和解析接口实现
- 模块如何根据其生成代码所采用的方法来注册 protobuf 文件描述符( API 模块方法仍可能作为一种受支持的策略而可行,并且会需要固定的文件描述符)
参考资料
Changelog
- 2022-04-27: First draft
Status
DRAFTAbstract
In order to move the Cosmos SDK to a system of decoupled semantically versioned modules which can be composed in different combinations (ex. staking v3 with bank v1 and distribution v2), we need to reassess how we organize the API surface of modules to avoid problems with go semantic import versioning and circular dependencies. This ADR explores various approaches we can take to addressing these issues.Context
There has been a fair amount of desire in the community for semantic versioning in the SDK and there has been significant movement to splitting SDK modules into standalone go modules. Both of these will ideally allow the ecosystem to move faster because we won’t be waiting for all dependencies to update synchronously. For instance, we could have 3 versions of the core SDK compatible with the latest 2 releases of CosmWasm as well as 4 different versions of staking . This sort of setup would allow early adopters to aggressively integrate new versions, while allowing more conservative users to be selective about which versions they’re ready for. In order to achieve this, we need to solve the following problems:- because of the way go semantic import versioning (SIV) works, moving to SIV naively will actually make it harder to achieve these goals
- circular dependencies between modules need to be broken to actually release many modules in the SDK independently
- pernicious minor version incompatibilities introduced through correctly evolving protobuf schemas without correct unknown field filtering
- proto files are maintained in a non-breaking way (using something like buf breaking to ensure all changes are backwards compatible)
- proto file versions get bumped much less frequently, i.e. we might maintain
cosmos.bank.v1through many versions of the bank module state machine - state machine breaking changes are more common and ideally this is what we’d want to semantically version with
go modules, ex.
x/bank/v2,x/bank/v3, etc.
Problem 1: Semantic Import Versioning Compatibility
Consider we have a modulefoo which defines the following MsgDoSomething and that we’ve released its state
machine in go module example.com/foo:
condition field to MsgDoSomething and also
add a new validation rule on amount requiring it to be non-zero, and that following go semantic versioning we
release the next state machine version of foo as example.com/foo/v2.
foo in example.com/foo/types and we would generate the protobuf
types for the second version in example.com/foo/v2/types.
Now let’s say we have a module bar which talks to foo using this keeper
interface which foo provides:
Scenario A: Backward Compatibility: Newer Foo, Older Bar
Imagine we have a chain which uses bothfoo and bar and wants to upgrade to
foo/v2, but the bar module has not upgraded to foo/v2.
In this case, the chain will not be able to upgrade to foo/v2 until bar
has upgraded its references to example.com/foo/types.MsgDoSomething to
example.com/foo/v2/types.MsgDoSomething.
Even if bar’s usage of MsgDoSomething has not changed at all, the upgrade
will be impossible without this change because example.com/foo/types.MsgDoSomething
and example.com/foo/v2/types.MsgDoSomething are fundamentally different
incompatible structs in the go type system.
Scenario B: Forward Compatibility: Older Foo, Newer Bar
Now let’s consider the reverse scenario, wherebar upgrades to foo/v2
by changing the MsgDoSomething reference to example.com/foo/v2/types.MsgDoSomething
and releases that as bar/v2 with some other changes that a chain wants.
The chain, however, has decided that it thinks the changes in foo/v2 are too
risky and that it’d prefer to stay on the initial version of foo.
In this scenario, it is impossible to upgrade to bar/v2 without upgrading
to foo/v2 even if bar/v2 would have worked 100% fine with foo other
than changing the import path to MsgDoSomething (meaning that bar/v2
doesn’t actually use any new features of foo/v2).
Now because of the way go semantic import versioning works, we are locked
into either using foo and bar OR foo/v2 and bar/v2. We cannot have
foo + bar/v2 OR foo/v2 + bar. The go type system doesn’t allow this
even if both versions of these modules are otherwise compatible with each
other.
Naive Mitigation
A naive approach to fixing this would be to not regenerate the protobuf types inexample.com/foo/v2/types but instead just update example.com/foo/types
to reflect the changes needed for v2 (adding condition and requiring
amount to be non-zero). Then we could release a patch of example.com/foo/types
with this update and use that for foo/v2. But this change is state machine
breaking for v1. It requires changing the ValidateBasic method to reject
the case where amount is zero, and it adds the condition field which
should be rejected based
on ADR 020 unknown field filtering.
So adding these changes as a patch on v1 is actually incorrect based on semantic
versioning. Chains that want to stay on v1 of foo should not
be importing these changes because they are incorrect for v1.
Problem 2: Circular dependencies
None of the above approaches allowfoo and bar to be separate modules
if for some reason foo and bar depend on each other in different ways.
For instance, we can’t have foo import bar/types while bar imports
foo/types.
We have several cases of circular module dependencies in the SDK
(ex. staking, distribution and slashing) that are legitimate from a state machine
perspective. Without separating the API types out somehow, there would be
no way to independently semantically version these modules without some other
mitigation.
Problem 3: Handling Minor Version Incompatibilities
Imagine that we solve the first two problems but now have a scenario wherebar/v2 wants the option to use MsgDoSomething.condition which only foo/v2
supports. If bar/v2 works with foo v1 and sets condition to some non-nil
value, then foo will silently ignore this field resulting in a silent logic
possibly dangerous logic error. If bar/v2 were able to check whether foo was
on v1 or v2 and dynamically, it could choose to only use condition when
foo/v2 is available. Even if bar/v2 were able to perform this check, however,
how do we know that it is always performing the check properly. Without
some sort of
framework-level unknown field filtering,
it is hard to know whether these pernicious hard to detect bugs are getting into
our app and a client-server layer such as ADR 033: Inter-Module Communication
may be needed to do this.
Solutions
Approach A) Separate API and State Machine Modules
One solution (first proposed in Link) is to isolate all protobuf generated code into a separate module from the state machine module. This would mean that we could have state machine go modulesfoo and foo/v2 which could use a types or API go module say
foo/api. This foo/api go module would be perpetually on v1.x and only
accept non-breaking changes. This would then allow other modules to be
compatible with either foo or foo/v2 as long as the inter-module API only
depends on the types in foo/api. It would also allow modules foo and bar
to depend on each other in that both of them could depend on foo/api and
bar/api without foo directly depending on bar and vice versa.
This is similar to the naive mitigation described above except that it separates
the types into separate go modules which in and of itself could be used to
break circular module dependencies. It has the same problems as the naive solution,
otherwise, which we could rectify by:
- removing all state machine breaking code from the API module (ex.
ValidateBasicand any other interface methods) - embedding the correct file descriptors for unknown field filtering in the binary
Migrate all interface methods on API types to handlers
To solve 1), we need to remove all interface implementations from generated types and instead use a handler approach which essentially means that given a typeX, we have some sort of resolver which allows us to resolve interface
implementations for that type (ex. sdk.Msg or authz.Authorization). For
example:
sdk.Msg, we could replace them with declarative
annotations. For instance, GetSigners can already be replaced by the protobuf
annotation cosmos.msg.v1.signer. In the future, we may consider some sort
of protobuf validation framework (like Link
but more Cosmos-specific) to replace ValidateBasic.
Pinned FileDescriptor’s
To solve 2), state machine modules must be able to specify what the version of the protobuf files was that they were built against. For instance if the API module forfoo upgrades to foo/v2, the original foo module still needs
a copy of the original protobuf files it was built with so that ADR 020
unknown field filtering will reject MsgDoSomething when condition is
set.
The simplest way to do this may be to embed the protobuf FileDescriptors into
the module itself so that these FileDescriptors are used at runtime rather
than the ones that are built into the foo/api which may be different. Using
buf build, go embed,
and a build script we can probably come up with a solution for embedding
FileDescriptors into modules that is fairly straightforward.
Potential limitations to generated code
One challenge with this approach is that it places heavy restrictions on what can go in API modules and requires that most of this is state machine breaking. All or most of the code in the API module would be generated from protobuf files, so we can probably control this with how code generation is done, but it is a risk to be aware of. For instance, we do code generation for the ORM that in the future could contain optimizations that are state machine breaking. We would either need to ensure very carefully that the optimizations aren’t actually state machine breaking in generated code or separate this generated code out from the API module into the state machine module. Both of these mitigations are potentially viable but the API module approach does require an extra level of care to avoid these sorts of issues.Minor Version Incompatibilities
This approach in and of itself does little to address any potential minor version incompatibilities and the requisite unknown field filtering. Likely some sort of client-server routing layer which does this check such as ADR 033: Inter-Module communication is required to make sure that this is done properly. We could then allow modules to perform a runtime check given aMsgClient, ex:
Approach B) Changes to Generated Code
An alternate approach to solving the versioning problem is to change how protobuf code is generated and move modules mostly or completely in the direction of inter-module communication as described in ADR 033. In this paradigm, a module could generate all the types it needs internally - including the API types of other modules - and talk to other modules via a client-server boundary. For instance, ifbar needs to talk to foo, it could
generate its own version of MsgDoSomething as bar/internal/foo/v1.MsgDoSomething and just pass this to the
inter-module router which would somehow convert it to the version which foo needs (ex. foo/internal.MsgDoSomething).
Currently, two generated structs for the same protobuf type cannot exist in the same go binary without special
build flags (see Link).
A relatively simple mitigation to this issue would be to set up the protobuf code to not register protobuf types
globally if they are generated in an internal/ package. This will require modules to register their types manually
with the app-level level protobuf registry, this is similar to what modules already do with the InterfaceRegistry
and amino codec.
If modules only do ADR 033 message passing then a naive and non-performant solution for
converting bar/internal/foo/v1.MsgDoSomething
to foo/internal.MsgDoSomething would be marshaling and unmarshaling in the ADR 033 router. This would break down if
we needed to expose protobuf types in Keeper interfaces because the whole point is to try to keep these types
internal/ so that we don’t end up with all the import version incompatibilities we’ve described above. However,
because of the issue with minor version incompatibilities and the need
for unknown field filtering,
sticking with the Keeper paradigm instead of ADR 033 may be unviable to begin with.
A more performant solution (that could maybe be adapted to work with Keeper interfaces) would be to only expose
getters and setters for generated types and internally store data in memory buffers which could be passed from
one implementation to another in a zero-copy way.
For example, imagine this protobuf API with only getters and setters is exposed for MsgSend:
MsgSend could be implemented based on some raw memory buffer in the same way
that Cap’n Proto
and FlatBuffers so that we could convert between one version of MsgSend
and another without serialization (i.e. zero-copy). This approach would have the added benefits of allowing zero-copy
message passing to modules written in other languages such as Rust and accessed through a VM or FFI. It could also make
unknown field filtering in inter-module communication simpler if we require that all new fields are added in sequential
order, ex. just checking that no field > 5 is set.
Also, we wouldn’t have any issues with state machine breaking code on generated types because all the generated
code used in the state machine would actually live in the state machine module itself. Depending on how interface
types and protobuf Anys are used in other languages, however, it may still be desirable to take the handler
approach described in approach A. Either way, types implementing interfaces would still need to be registered
with an InterfaceRegistry as they are now because there would be no way to retrieve them via the global registry.
In order to simplify access to other modules using ADR 033, a public API module (maybe even one
remotely generated by Buf) could be used by client modules instead
of requiring to generate all client types internally.
The big downsides of this approach are that it requires big changes to how people use protobuf types and would be a
substantial rewrite of the protobuf code generator. This new generated code, however, could still be made compatible
with
the google.golang.org/protobuf/reflect/protoreflect
API in order to work with all standard golang protobuf tooling.
It is possible that the naive approach of marshaling/unmarshaling in the ADR 033 router is an acceptable intermediate
solution if the changes to the code generator are seen as too complex. However, since all modules would likely need
to migrate to ADR 033 anyway with this approach, it might be better to do this all at once.
Approach C) Don’t address these issues
If the above solutions are seen as too complex, we can also decide not to do anything explicit to enable better module version compatibility, and break circular dependencies. In this case, when developers are confronted with the issues described above they can require dependencies to update in sync (what we do now) or attempt some ad-hoc potentially hacky solution. One approach is to ditch go semantic import versioning (SIV) altogether. Some people have commented that go’s SIV (i.e. changing the import path tofoo/v2, foo/v3, etc.) is too restrictive and that it should be optional. The
golang maintainers disagree and only officially support semantic import versioning. We could, however, take the
contrarian perspective and get more flexibility by using 0.x-based versioning basically forever.
Module version compatibility could then be achieved using go.mod replace directives to pin dependencies to specific
compatible 0.x versions. For instance if we knew foo 0.2 and 0.3 were both compatible with bar 0.3 and 0.4, we
could use replace directives in our go.mod to stick to the versions of foo and bar we want. This would work as
long as the authors of foo and bar avoid incompatible breaking changes between these modules.
Or, if developers choose to use semantic import versioning, they can attempt the naive solution described above
and would also need to use special tags and replace directives to make sure that modules are pinned to the correct
versions.
Note, however, that all of these ad-hoc approaches, would be vulnerable to the minor version compatibility issues
described above unless unknown field filtering
is properly addressed.
Approach D) Avoid protobuf generated code in public APIs
An alternative approach would be to avoid protobuf generated code in public module APIs. This would help avoid the discrepancy between state machine versions and client API versions at the module to module boundaries. It would mean that we wouldn’t do inter-module message passing based on ADR 033, but rather stick to the existing keeper approach and take it one step further by avoiding any protobuf generated code in the keeper interface methods. Using this approach, ourfoo.Keeper.DoSomething method wouldn’t have the generated MsgDoSomething struct (which
comes from the protobuf API), but instead positional parameters. Then in order for foo/v2 to support the foo/v1
keeper it would simply need to implement both the v1 and v2 keeper APIs. The DoSomething method in v2 could have the
additional condition parameter, but this wouldn’t be present in v1 at all so there would be no danger of a client
accidentally setting this when it isn’t available.
So this approach would avoid the challenge around minor version incompatibilities because the existing module keeper
API would not get new fields when they are added to protobuf files.
Taking this approach, however, would likely require making all protobuf generated code internal in order to prevent
it from leaking into the keeper API. This means we would still need to modify the protobuf code generator to not
register internal/ code with the global registry, and we would still need to manually register protobuf
FileDescriptors (this is probably true in all scenarios). It may, however, be possible to avoid needing to refactor
interface methods on generated types to handlers.
Also, this approach doesn’t address what would be done in scenarios where modules still want to use the message router.
Either way, we probably still want a way to pass messages from one module to another router safely even if it’s just for
use cases like x/gov, x/authz, CosmWasm, etc. That would still require most of the things outlined in approach (B),
although we could advise modules to prefer keepers for communicating with other modules.
The biggest downside of this approach is probably that it requires a strict refactoring of keeper interfaces to avoid
generated code leaking into the API. This may result in cases where we need to duplicate types that are already defined
in proto files and then write methods for converting between the golang and protobuf version. This may end up in a lot
of unnecessary boilerplate and that may discourage modules from actually adopting it and achieving effective version
compatibility. Approaches (A) and (B), although heavy handed initially, aim to provide a system which once adopted
more or less gives the developer version compatibility for free with minimal boilerplate. Approach (D) may not be able
to provide such a straightforward system since it requires a golang API to be defined alongside a protobuf API in a
way that requires duplication and differing sets of design principles (protobuf APIs encourage additive changes
while golang APIs would forbid it).
Other downsides to this approach are:
- no clear roadmap to supporting modules in other languages like Rust
- doesn’t get us any closer to proper object capability security (one of the goals of ADR 033)
- ADR 033 needs to be done properly anyway for the set of use cases which do need it
Decision
The latest DRAFT proposal is:- we are alignment on adopting ADR 033 not just as an addition to the framework, but as a core replacement to the keeper paradigm entirely.
- the ADR 033 inter-module router will accommodate any variation of approach (A) or (B) given the following rules: a. if the client type is the same as the server type then pass it directly through, b. if both client and server use the zero-copy generated code wrappers (which still need to be defined), then pass the memory buffers from one wrapper to the other, or c. marshal/unmarshal types between client and server.
Minor API Revisions
To declare minor API revisions of proto files, we propose the following guidelines (which were already documented in cosmos.app.v1alpha module options):- proto packages which are revised from their initial version (considered revision
0) should include apackage - comment in some .proto file containing the test
Revision Nat the start of a comment line whereNis the current revision number. - all fields, messages, etc. added in a version beyond the initial revision should add a comment at the start of a
comment line of the form
Since: Revision NwhereNis the non-zero revision it was added.
1.N where N corresponds to the package revision. Patch releases should be used when
only documentation comments are updated. It is okay to include proto packages named v2, v3, etc. in this same
1.N versioned buf module (ex. cosmos.bank.v2) as long as all these proto packages consist of a single API intended
to be served by a single SDK module.
Introspecting Minor API Revisions
In order for modules to introspect the minor API revision of peer modules, we propose adding the following method tocosmossdk.io/core/intermodule.Client:
Unknown Field Filtering
To correctly perform unknown field filtering, the inter-module router can do one of the following:- use the
protoreflectAPI for messages which support that - for gogo proto messages, marshal and use the existing
codec/unknownprotocode - for zero-copy messages, do a simple check on the highest set field number (assuming we can require that fields are adding consecutively in increasing order)
FileDescriptor Registration
Because a single go binary may contain different versions of the same generated protobuf code, we cannot rely on the
global protobuf registry to contain the correct FileDescriptors. Because appconfig module configuration is itself
written in protobuf, we would like to load the FileDescriptors for a module before loading a module itself. So we
will provide ways to register FileDescriptors at module registration time before instantiation. We propose the
following cosmossdk.io/core/appmodule.Option constructors for the various cases of how FileDescriptors may be
packaged:
- proto files generated internally to a module (use
ProtoFiles) - the API module approach with pinned file descriptors (use
ProtoImage) - gogo proto (use
GzippedProtoFiles)
Module Dependency Declaration
One risk of ADR 033 is that dependencies are called at runtime which are not present in the loaded set of SDK modules.Also we want modules to have a way to define a minimum dependency API revision that they require. Therefore, all modules should declare their set of dependencies upfront. These dependencies could be defined when a module is instantiated, but ideally we know what the dependencies are before instantiation and can statically look at an app config and determine whether the set of modules. For example, if
bar requires foo revision >= 1, then we
should be able to know this when creating an app config with two versions of bar and foo.
We propose defining these dependencies in the proto options of the module config object itself.
Interface Registration
We will also need to define how interface methods are defined on types that are serialized asgoogle.protobuf.Any’s.
In light of the desire to support modules in other languages, we may want to think of solutions that will accommodate
other languages such as plugins described briefly in ADR 033.
Testing
In order to ensure that modules are indeed with multiple versions of their dependencies, we plan to provide specialized unit and integration testing infrastructure that automatically tests multiple versions of dependencies.Unit Testing
Unit tests should be conducted inside SDK modules by mocking their dependencies. In a full ADR 033 scenario, this means that all interaction with other modules is done via the inter-module router, so mocking of dependencies means mocking their msg and query server implementations. We will provide both a test runner and fixture to make this streamlined. The key thing that the test runner should do to test compatibility is to test all combinations of dependency API revisions. This can be done by taking the file descriptors for the dependencies, parsing their comments to determine the revisions various elements were added, and then created synthetic file descriptors for each revision by subtracting elements that were added later. Here is a proposed API for the unit test runner and fixture:Integration Testing
An integration test runner and fixture would also be provided which instead of using mocks would test actual module dependencies in various combinations. Here is the proposed API:example.com/bar/v2/test. Because this paradigm uses go semantic
versioning, it is possible to build a single go module which imports 3 versions of bar and 2 versions of runtime and
can test these all together in the six various combinations of dependencies.
Consequences
Backwards Compatibility
Modules which migrate fully to ADR 033 will not be compatible with existing modules which use the keeper paradigm. As a temporary workaround we may create some wrapper types that emulate the current keeper interface to minimize the migration overhead.Positive
- we will be able to deliver interoperable semantically versioned modules which should dramatically increase the ability of the Cosmos SDK ecosystem to iterate on new features
- it will be possible to write Cosmos SDK modules in other languages in the near future
Negative
- all modules will need to be refactored somewhat dramatically
Neutral
- the
cosmossdk.io/core/appconfigframework will play a more central role in terms of how modules are defined, this is likely generally a good thing but does mean additional changes for users wanting to stick to the pre-depinject way of wiring up modules depinjectis somewhat less needed or maybe even obviated because of the full ADR 033 approach. If we adopt the core API proposed in Link, then a module would probably always instantiate itself with a methodProvideModule(appmodule.Service) (appmodule.AppModule, error). There is no complex wiring of keeper dependencies in this scenario and dependency injection may not have as much of (or any) use case.
Further Discussions
The decision described above is considered in draft mode and is pending final buy-in from the team and key stakeholders. Key outstanding discussions if we do adopt that direction are:- how do module clients introspect dependency module API revisions
- how do modules determine a minor dependency module API revision requirement
- how do modules appropriately test compatibility with different dependency versions
- how to register and resolve interface implementations
- how do modules register their protobuf file descriptors depending on the approach they take to generated code (the API module approach may still be viable as a supported strategy and would need pinned file descriptors)