变更记录

  • 2022-04-27:初稿

状态

已接受 已实现

摘要

为了让开发者更容易构建 Cosmos SDK 模块,并让客户端能够针对状态数据进行查询、索引和证明验证,我们为 Cosmos SDK 实现了一层 ORM(对象关系映射)。

背景

从历史上看,Cosmos SDK 中的模块一直都是直接使用键值存储,并手写各种函数来管理键格式以及构造二级索引。这会在构建模块时消耗大量时间,而且容易出错。由于键格式并不标准化,有时文档记录也不充分,并且还可能发生变化,因此客户端很难以通用方式对状态数据进行索引、查询和 Merkle 证明验证。 Cosmos 生态中已知最早的 “ORM” 实例出现在 weave 中。随后,regen-ledger 为 group 模块构建了一个后续版本,后来又为该用途移植到 SDK。 尽管这些早期设计显著降低了状态机编写难度,但它们仍然需要大量手动配置,无法将状态格式直接暴露给客户端,而且对不同类型的索引键、复合键和范围查询的支持也有限。 围绕该设计的讨论持续在 链接 中推进,随后又在 链接 和 链接 中创建了更复杂的概念验证实现。

决策

这些先前的工作最终促成了 Cosmos SDK orm Go 模块的诞生。该模块使用 protobuf 注解来定义 ORM 表结构。这个 ORM 基于新的 google.golang.org/protobuf/reflect/protoreflect API,并支持:
  • 适用于所有简单 protobuf 类型的有序索引(bytes、enum、float、double 除外),以及 Timestamp 和 Duration
  • 无序的 bytes 和 enum 索引
  • 复合主键和复合二级键
  • 唯一索引
  • 自动递增的 uint64 主键
  • 复杂前缀查询和范围查询
  • 分页查询
  • 对 KV-store 数据的完整逻辑解码
几乎所有直接解码状态所需的信息都定义在 .proto 文件中。每个表定义都指定了一个在对应 .proto 文件内唯一的 ID,而表中的每个索引在该表内也都是唯一的。因此,客户端只需要知道模块名称,以及该模块内某个特定 .proto 文件对应的 ORM 数据前缀,就能够直接解码状态数据。这些附加信息将通过应用配置直接暴露,具体方式会在未来与应用接线相关的 ADR 中说明。 ORM 还通过避免在存储主键记录时于键值中重复主键值,来优化存储空间。例如,如果对象 {"a":0,"b":1} 的主键是 a,那么它会以 Key: '0', Value: {"b":1} 的形式存储在键值存储中(实际会使用更高效的 protobuf 二进制编码)。此外,链接 生成的代码也会围绕 google.golang.org/protobuf/reflect/protoreflect API 做性能优化。 ORM 内置了一个代码生成器,可为 ORM 的动态 Table 实现生成类型安全的包装器,这也是模块使用 ORM 的推荐方式。 ORM 测试提供了一个简化版 bank 模块示例,用于说明:

影响

向后兼容性

采用 ORM 的状态机代码需要进行迁移,因为其状态布局通常与旧版本不向后兼容。这些状态机也需要迁移到 链接,至少状态数据部分需要如此。

正面影响

  • 更容易构建模块
  • 更容易为状态添加二级索引
  • 可以为 ORM 状态编写通用索引器
  • 更容易编写执行状态证明的客户端
  • 可以自动生成查询层,而不必手动实现 gRPC 查询

负面影响

中性影响

进一步讨论

进一步讨论将在 Cosmos SDK Framework Working Group 内开展。目前已规划和正在进行的工作包括:
  • 自动生成面向客户端的查询层
  • 在客户端侧提供可透明验证轻客户端证明的查询库
  • 将 ORM 数据索引到 SQL 数据库
  • 通过以下方式提升性能:
    • 优化现有基于反射的代码,在对简单表执行删除和更新时避免不必要的读取
    • 进行更复杂的代码生成,例如让快速路径反射更快(避免 switch 语句),甚至完全生成与手写实现性能相当的代码

参考资料


Changelog

  • 2022-04-27: First draft

Status

ACCEPTED Implemented

Abstract

In order to make it easier for developers to build Cosmos SDK modules and for clients to query, index and verify proofs against state data, we have implemented an ORM (object-relational mapping) layer for the Cosmos SDK.

Context

Historically modules in the Cosmos SDK have always used the key-value store directly and created various handwritten functions for managing key format as well as constructing secondary indexes. This consumes a significant amount of time when building a module and is error-prone. Because key formats are non-standard, sometimes poorly documented, and subject to change, it is hard for clients to generically index, query and verify merkle proofs against state data. The known first instance of an “ORM” in the Cosmos ecosystem was in weave. A later version was built for regen-ledger for use in the group module and later ported to the SDK just for that purpose. While these earlier designs made it significantly easier to write state machines, they still required a lot of manual configuration, didn’t expose state format directly to clients, and were limited in their support of different types of index keys, composite keys, and range queries. Discussions about the design continued in Link and more sophisticated proofs of concept were created in Link and Link.

Decision

These prior efforts culminated in the creation of the Cosmos SDK orm go module which uses protobuf annotations for specifying ORM table definitions. This ORM is based on the new google.golang.org/protobuf/reflect/protoreflect API and supports:
  • sorted indexes for all simple protobuf types (except bytes, enum, float, double) as well as Timestamp and Duration
  • unsorted bytes and enum indexes
  • composite primary and secondary keys
  • unique indexes
  • auto-incrementing uint64 primary keys
  • complex prefix and range queries
  • paginated queries
  • complete logical decoding of KV-store data
Almost all the information needed to decode state directly is specified in .proto files. Each table definition specifies an ID which is unique per .proto file and each index within a table is unique within that table. Clients then only need to know the name of a module and the prefix ORM data for a specific .proto file within that module in order to decode state data directly. This additional information will be exposed directly through app configs which will be explained in a future ADR related to app wiring. The ORM makes optimizations around storage space by not repeating values in the primary key in the key value when storing primary key records. For example, if the object {"a":0,"b":1} has the primary key a, it will be stored in the key value store as Key: '0', Value: {"b":1} (with more efficient protobuf binary encoding). Also, the generated code from Link does optimizations around the google.golang.org/protobuf/reflect/protoreflect API to improve performance. A code generator is included with the ORM which creates type safe wrappers around the ORM’s dynamic Table implementation and is the recommended way for modules to use the ORM. The ORM tests provide a simplified bank module demonstration which illustrates:

Consequences

Backwards Compatibility

State machine code that adopts the ORM will need migrations as the state layout is generally backwards incompatible. These state machines will also need to migrate to Link at least for state data.

Positive

  • easier to build modules
  • easier to add secondary indexes to state
  • possible to write a generic indexer for ORM state
  • easier to write clients that do state proofs
  • possible to automatically write query layers rather than needing to manually implement gRPC queries

Negative

  • worse performance than handwritten keys (for now). See Further Discussions for potential improvements

Neutral

Further Discussions

Further discussions will happen within the Cosmos SDK Framework Working Group. Current planned and ongoing work includes:
  • automatically generate client-facing query layer
  • client-side query libraries that transparently verify light client proofs
  • index ORM data to SQL databases
  • improve performance by:
    • optimizing existing reflection based code to avoid unnecessary gets when doing deletes & updates of simple tables
    • more sophisticated code generation such as making fast path reflection even faster (avoiding switch statements), or even fully generating code that equals handwritten performance

References