变更记录

  • 2023 年 2 月 14 日:初始草案(@alexanderbez)

状态

草案

摘要

Cosmos SDK 应用长期以来所使用的存储与状态原语,自首个 Cosmos Hub 发布以来基本没有发生变化。从开发者体验和客户端用户体验两个视角来看,Cosmos SDK 应用的诉求与需求已经不断演进,并且早已超出这些原语最初引入时所能覆盖的生态范围。 随着这些应用逐步获得大规模采用,Cosmos SDK 的状态与存储原语中暴露出了许多关键性的不足与缺陷。 为了跟上客户端与开发者不断演进的需求,有必要对这些原语进行一次重大重构。

背景

Cosmos SDK 为应用开发者提供了多种用于处理应用状态的存储原语。具体来说,每个模块都包含其自身的默克尔承诺数据结构,即一棵 IAVL 树。在这一数据结构中,模块可以存储和检索键值对,并附带这些键值对的 Merkle 承诺,也就是证明,用来表明它们在全局应用状态中存在或不存在。该数据结构就是基础层 KVStore。 此外,SDK 还在这一 Merkle 数据结构之上提供了抽象层。也就是说,根多存储(RMS)是由各模块 KVStore 组成的集合。通过 RMS,应用不仅可以响应查询并向客户端提供证明,还可以借助 StoreKey(一种 OCAP 原语)让模块访问其自身唯一的 KVStore。 在 RMS 与底层 IAVL KVStore 之间,还存在更多抽象层。GasKVStore 负责跟踪状态机读写带来的 gas IO 消耗。CacheKVStore 则负责提供读取缓存与写入缓冲机制,以实现状态转换的原子性,例如交易执行或治理提案执行。 这些抽象层以及 Cosmos SDK 整体存储设计存在一些关键缺点:
  • 由于每个模块都有各自的 IAVL KVStore,承诺不是原子的
    • 注意,我们仍然可以允许模块拥有各自的 IAVL KVStore,但 IAVL 库需要支持在多个 IAVL API 中将 DB 实例作为参数传入。
  • 由于 IAVL 同时负责状态存储和承诺,随着磁盘空间呈指数级增长,运行归档节点的成本会越来越高。
  • 随着网络规模扩大,查询性能、网络升级、状态迁移以及整体应用性能等多个方面都会开始出现性能瓶颈。
  • 开发者体验较差,因为它不允许应用开发者尝试不同的存储与承诺方案,同时还要面对上文提到的多层抽象所带来的复杂性。
更多信息请参见存储讨论。

备选方案

此前已有一次对存储层进行重构的尝试,见 ADR-040。 不过,这一方案主要源于 IAVL 的缺陷及其周边的多种性能问题。尽管 ADR-040 曾有过一个(部分)实现,但由于多种原因它从未被采用,例如依赖 SMT,而 SMT 当时仍更接近研究阶段;以及一些未能完全达成一致的设计选择,例如会导致状态体积大幅膨胀的快照机制。

决策

我们提议在 ADR-040 引入的一些优秀思想基础上继续推进,同时在底层实现上保持更高灵活性,并尽量减少侵入性。具体来说,我们提议:
  • 将共识所需的状态承诺(SC)与状态机和客户端所需的状态存储(SS)分离。
  • 减少 RMS 与底层存储之间所需的抽象层数。
  • 通过向核心 IAVL API 提供批量数据库对象,实现模块存储承诺的原子性。
  • 降低 CacheKVStore 实现的复杂度,同时提升性能[3]。
此外,当前阶段我们仍将 IAVL 保留为底层的承诺存储。尽管从长期来看我们未必最终会继续使用 IAVL,但目前并没有足够有力的经验数据表明存在更好的替代方案。考虑到 SDK 为存储提供了接口,如果未来出现足够证据证明存在更优替代方案,那么替换底层承诺存储应当是可行的。不过,围绕 IAVL 目前已有一些很有前景的工作,预计将带来显著的性能提升 [1,2]。

分离 SS 与 SC

通过分离 SS 与 SC,我们可以针对状态的主要使用场景和访问模式进行优化。具体而言,SS 层将负责以(key, value)对形式直接访问数据,而 SC 层(IAVL)将负责对数据进行承诺并提供 Merkle 证明。 需要注意的是,SS 层与 SC 层底层使用的物理存储数据库将是同一个。因此,为了避免(key, value)对发生冲突,这两层都将使用命名空间进行隔离。

状态承诺(SC)

鉴于当前现有方案同时承担 SS 和 SC 两种职责,我们可以直接将其重新定位为仅承担 SC 层职责,而无需对访问模式或行为做出重大变更。换句话说,现有整套由 IAVL 支撑的模块 KVStore 集合将作为 SC 层。 不过,为了让 SC 层保持轻量,并避免复制 SS 层中绝大多数数据,我们建议节点运营者采用较为严格的裁剪策略。

状态存储(SS)

在 RMS 中,我们将暴露一个由与 SC 层相同物理数据库支撑的单一 KVStore。这个 KVStore 将显式使用命名空间以避免冲突,并作为(key, value)对的主要存储。 虽然我们很可能会继续使用 cosmos-db 或某种本地接口,以便在研究和基准测试持续推进的过程中,保留对首选物理存储后端的灵活性和迭代空间,但我们提议将 RocksDB 硬编码为主要的物理存储后端。 由于 SS 层将以 KVStore 的形式实现,因此它将支持以下功能:
  • 范围查询
  • CRUD 操作
  • 历史查询与版本控制
  • 裁剪
RMS 将通过为每个 StoreKey 设置专用的内部 MemoryListener 来跟踪所有缓冲写入。对于每个区块高度,在执行 Commit 时,SS 层会将所有缓冲的(key, value)对写入带有 RocksDB 用户自定义时间戳列族中,并使用区块高度作为时间戳,该时间戳为无符号整数。这样一来,客户端就可以获取历史高度和当前高度下的(key, value)对,同时由于时间戳是键后缀,迭代和范围查询也能保持相对较好的性能。 需要注意的是,我们没有选择更通用的方案,即允许使用任意嵌入式键值数据库(如 LevelDB 或 PebbleDB),并通过以高度为前缀的键来实现状态版本化,因为这些数据库大多使用变长键,这会使迭代和范围查询等操作的性能变差。 由于运营者可能希望 SS 的裁剪策略与 SC 不同,例如 SC 采用非常严格的裁剪策略,而 SS 采用较宽松的裁剪策略,因此我们提议新增一套额外的裁剪配置,其参数与当前 SDK 中已有的配置保持一致,并允许运营者独立控制 SS 层的裁剪策略,而不受 SC 层影响。 需要注意的是,SC 的裁剪策略必须与运营者的状态同步配置保持一致。这样才能确保状态同步快照能够成功执行,否则可能会在 SC 中不存在的高度触发快照。

状态同步

状态同步流程基本不会因 SC 层与 SS 层的分离而受到影响。不过,如果节点通过状态同步完成同步,那么该节点的 SS 层将不会包含状态同步对应高度的数据,因为当前 IAVL 导入流程的设计方式并不便于直接插入键值对。为了让状态同步高度在 SS 层中可用,必须对 IAVL 导入流程进行修改。 需要注意的是,这对状态机本身并不会造成问题,因为在发起查询时,RMS 会自动将查询正确路由(参见查询)。

查询

为了统一 SC 层与 SS 层之间的查询路由,我们提议在 RMS 中引入“查询路由器”的概念。该查询路由器会提供给每个 KVStore 实现。查询路由器将基于若干参数,把查询路由到 SC 层或 SS 层。如果 prove: true,那么查询必须路由到 SC 层。否则,如果查询高度在 SS 层中可用,则由 SS 层提供查询结果;如果不可用,则回退到 SC 层。 如果没有提供高度,SS 层将默认使用最新高度。SS 层会维护一个反向索引,用于查找 LatestVersion -> timestamp(version),该索引会在 Commit 时设置。

证明

由于 SS 层天然只是一个存储层,并不对(key, value)对做任何承诺,因此它无法在查询期间向客户端提供 Merkle 证明。 由于针对 SC 层的裁剪策略由运营者配置,因此只要对应版本存在且 prove: true,RMS 就可以将查询路由到 SC 层。否则,查询将回退到 SS 层,并且不附带证明。 我们也可以探索这样一种思路:利用状态快照,针对查询中给定版本附近的某个版本,实时重建一棵内存中的 IAVL 树。不过,目前尚不清楚这种方法在性能上的影响会如何。

原子承诺

我们提议修改现有 IAVL API,使其接受一个批量 DB 对象,而不是依赖 nodeDB 内部的批量对象。由于 SC 层中的每个底层 IAVL KVStore 共享同一个 DB,这将使提交具备原子性。 具体来说,我们提议:
  • 从 nodeDB 中移除 dbm.Batch 字段
  • 更新 IAVL MutableTree 类型的 SaveVersion 方法,使其接收批量对象
  • 更新 CommitKVStore 接口的 Commit 方法,使其接收批量对象
  • 在 RMS 执行 Commit 时创建批量对象,并将该对象传递给每个 KVStore
  • 在所有存储都成功提交之后,再统一写入数据库批量操作
需要注意的是,这将要求 IAVL 更新实现,不再在 SaveVersion 期间依赖或假设某个批量对象必然存在。

结果

引入新的 store V2 包之后,由于职责分离,我们预计查询和交易的性能都会得到提升。我们也预计开发者在尝试不同承诺方案与存储后端以进一步优化性能时,将获得更好的开发体验;同时,由于围绕 KVStore 的抽象层减少,缓存和状态分支等操作也会变得更加直观。 不过,由于所提议的设计,在为历史查询提供状态证明方面会存在一些缺点。

向后兼容性

本 ADR 提议通过一个全新的 package 对 Cosmos SDK 中的存储实现进行变更。接口可以借用并扩展 store 中现有类型的设计,但不会破坏或修改任何现有实现或接口。

正面影响

  • 提升独立 SS 层和 SC 层的性能
  • 减少抽象层级,使存储原语更易于理解
  • 为 SC 提供原子提交
  • 重新设计存储类型和接口将允许开展更多实验, 例如为不同的应用模块采用不同的物理存储后端和不同的承诺方案

负面影响

  • 为历史状态提供证明具有挑战性

中性影响

  • 继续将 IAVL 作为主要的承诺数据结构,尽管其性能正在获得显著提升

进一步讨论

模块存储控制

许多模块会存储二级索引,这些索引通常仅用于支持客户端查询,但实际上并不是状态机状态转换所必需的。这意味着,从技术上讲,这些索引根本没有理由存在于 SC 层中,因为它们会占用不必要的空间。值得探索一种 API 形式,使模块能够指明它们希望哪些 (key, value) 对持久化到 SC 层中,这也隐含表示会持久化到 SS 层,而不是仅将 (key, value) 对持久化到 SS 层。

历史状态证明

目前尚不清楚,在社区内部,为历史状态提供承诺证明的重要性或需求究竟有多高。虽然可以设计出一些方案,例如基于状态快照动态重建树结构,但这类方案会带来怎样的性能影响仍不明确。

物理 DB 后端

本 ADR 提议使用 RocksDB,以利用用户定义时间戳作为版本控制机制。不过,当前也存在其他可用的物理 DB 后端,它们可能提供替代性的版本实现方式,同时在性能上优于 RocksDB。例如,PebbleDB 同样支持 MVCC 时间戳,但我们仍需要进一步探索 PebbleDB 如何处理压缩以及状态随时间增长的问题。

参考资料


Changelog

  • Feb 14, 2023: Initial Draft (@alexanderbez)

Status

DRAFT

Abstract

The storage and state primitives that Cosmos SDK based applications have used have by and large not changed since the launch of the inaugural Cosmos Hub. The demands and needs of Cosmos SDK based applications, from both developer and client UX perspectives, have evolved and outgrown the ecosystem since these primitives were first introduced. Over time as these applications have gained significant adoption, many critical shortcomings and flaws have been exposed in the state and storage primitives of the Cosmos SDK. In order to keep up with the evolving demands and needs of both clients and developers, a major overhaul to these primitives are necessary.

Context

The Cosmos SDK provides application developers with various storage primitives for dealing with application state. Specifically, each module contains its own merkle commitment data structure — an IAVL tree. In this data structure, a module can store and retrieve key-value pairs along with Merkle commitments, i.e. proofs, to those key-value pairs indicating that they do or do not exist in the global application state. This data structure is the base layer KVStore. In addition, the SDK provides abstractions on top of this Merkle data structure. Namely, a root multi-store (RMS) is a collection of each module’s KVStore. Through the RMS, the application can serve queries and provide proofs to clients in addition to provide a module access to its own unique KVStore though the use of StoreKey, which is an OCAP primitive. There are further layers of abstraction that sit between the RMS and the underlying IAVL KVStore. A GasKVStore is responsible for tracking gas IO consumption for state machine reads and writes. A CacheKVStore is responsible for providing a way to cache reads and buffer writes to make state transitions atomic, e.g. transaction execution or governance proposal execution. There are a few critical drawbacks to these layers of abstraction and the overall design of storage in the Cosmos SDK:
  • Since each module has its own IAVL KVStore, commitments are not atomic
    • Note, we can still allow modules to have their own IAVL KVStore, but the IAVL library will need to support the ability to pass a DB instance as an argument to various IAVL APIs.
  • Since IAVL is responsible for both state storage and commitment, running an archive node becomes increasingly expensive as disk space grows exponentially.
  • As the size of a network increases, various performance bottlenecks start to emerge in many areas such as query performance, network upgrades, state migrations, and general application performance.
  • Developer UX is poor as it does not allow application developers to experiment with different types of approaches to storage and commitments, along with the complications of many layers of abstractions referenced above.
See the Storage Discussion for more information.

Alternatives

There was a previous attempt to refactor the storage layer described in ADR-040. However, this approach mainly stems on the short comings of IAVL and various performance issues around it. While there was a (partial) implementation of ADR-040, it was never adopted for a variety of reasons, such as the reliance on using an SMT, which was more in a research phase, and some design choices that couldn’t be fully agreed upon, such as the snap-shotting mechanism that would result in massive state bloat.

Decision

We propose to build upon some of the great ideas introduced in ADR-040, while being a bit more flexible with the underlying implementations and overall less intrusive. Specifically, we propose to:
  • Separate the concerns of state commitment (SC), needed for consensus, and state storage (SS), needed for state machine and clients.
  • Reduce layers of abstractions necessary between the RMS and underlying stores.
  • Provide atomic module store commitments by providing a batch database object to core IAVL APIs.
  • Reduce complexities in the CacheKVStore implementation while also improving performance[3].
Furthermore, we will keep the IAVL is the backing commitment store for the time being. While we might not fully settle on the use of IAVL in the long term, we do not have strong empirical evidence to suggest a better alternative. Given that the SDK provides interfaces for stores, it should be sufficient to change the backing commitment store in the future should evidence arise to warrant a better alternative. However there is promising work being done to IAVL that should result in significant performance improvement [1,2].

Separating SS and SC

By separating SS and SC, it will allow for us to optimize against primary use cases and access patterns to state. Specifically, The SS layer will be responsible for direct access to data in the form of (key, value) pairs, whereas the SC layer (IAVL) will be responsible for committing to data and providing Merkle proofs. Note, the underlying physical storage database will be the same between both the SS and SC layers. So to avoid collisions between (key, value) pairs, both layers will be namespaced.

State Commitment (SC)

Given that the existing solution today acts as both SS and SC, we can simply repurpose it to act solely as the SC layer without any significant changes to access patterns or behavior. In other words, the entire collection of existing IAVL-backed module KVStores will act as the SC layer. However, in order for the SC layer to remain lightweight and not duplicate a majority of the data held in the SS layer, we encourage node operators to keep tight pruning strategies.

State Storage (SS)

In the RMS, we will expose a single KVStore backed by the same physical database that backs the SC layer. This KVStore will be explicitly namespaced to avoid collisions and will act as the primary storage for (key, value) pairs. While we most likely will continue the use of cosmos-db, or some local interface, to allow for flexibility and iteration over preferred physical storage backends as research and benchmarking continues. However, we propose to hardcode the use of RocksDB as the primary physical storage backend. Since the SS layer will be implemented as a KVStore, it will support the following functionality:
  • Range queries
  • CRUD operations
  • Historical queries and versioning
  • Pruning
The RMS will keep track of all buffered writes using a dedicated and internal MemoryListener for each StoreKey. For each block height, upon Commit, the SS layer will write all buffered (key, value) pairs under a RocksDB user-defined timestamp column family using the block height as the timestamp, which is an unsigned integer. This will allow a client to fetch (key, value) pairs at historical and current heights along with making iteration and range queries relatively performant as the timestamp is the key suffix. Note, we choose not to use a more general approach of allowing any embedded key/value database, such as LevelDB or PebbleDB, using height key-prefixed keys to effectively version state because most of these databases use variable length keys which would effectively make actions likes iteration and range queries less performant. Since operators might want pruning strategies to differ in SS compared to SC, e.g. having a very tight pruning strategy in SC while having a looser pruning strategy for SS, we propose to introduce an additional pruning configuration, with parameters that are identical to what exists in the SDK today, and allow operators to control the pruning strategy of the SS layer independently of the SC layer. Note, the SC pruning strategy must be congruent with the operator’s state sync configuration. This is so as to allow state sync snapshots to execute successfully, otherwise, a snapshot could be triggered on a height that is not available in SC.

State Sync

The state sync process should be largely unaffected by the separation of the SC and SS layers. However, if a node syncs via state sync, the SS layer of the node will not have the state synced height available, since the IAVL import process is not setup in way to easily allow direct key/value insertion. A modification of the IAVL import process would be necessary to facilitate having the state sync height available. Note, this is not problematic for the state machine itself because when a query is made, the RMS will automatically direct the query correctly (see Queries).

Queries

To consolidate the query routing between both the SC and SS layers, we propose to have a notion of a “query router” that is constructed in the RMS. This query router will be supplied to each KVStore implementation. The query router will route queries to either the SC layer or the SS layer based on a few parameters. If prove: true, then the query must be routed to the SC layer. Otherwise, if the query height is available in the SS layer, the query will be served from the SS layer. Otherwise, we fall back on the SC layer. If no height is provided, the SS layer will assume the latest height. The SS layer will store a reverse index to lookup LatestVersion -> timestamp(version) which is set on Commit.

Proofs

Since the SS layer is naturally a storage layer only, without any commitments to (key, value) pairs, it cannot provide Merkle proofs to clients during queries. Since the pruning strategy against the SC layer is configured by the operator, we can therefore have the RMS route the query SC layer if the version exists and prove: true. Otherwise, the query will fall back to the SS layer without a proof. We could explore the idea of using state snapshots to rebuild an in-memory IAVL tree in real time against a version closest to the one provided in the query. However, it is not clear what the performance implications will be of this approach.

Atomic Commitment

We propose to modify the existing IAVL APIs to accept a batch DB object instead of relying on an internal batch object in nodeDB. Since each underlying IAVL KVStore shares the same DB in the SC layer, this will allow commits to be atomic. Specifically, we propose to:
  • Remove the dbm.Batch field from nodeDB
  • Update the SaveVersion method of the MutableTree IAVL type to accept a batch object
  • Update the Commit method of the CommitKVStore interface to accept a batch object
  • Create a batch object in the RMS during Commit and pass this object to each KVStore
  • Write the database batch after all stores have committed successfully
Note, this will require IAVL to be updated to not rely or assume on any batch being present during SaveVersion.

Consequences

As a result of a new store V2 package, we should expect to see improved performance for queries and transactions due to the separation of concerns. We should also expect to see improved developer UX around experimentation of commitment schemes and storage backends for further performance, in addition to a reduced amount of abstraction around KVStores making operations such as caching and state branching more intuitive. However, due to the proposed design, there are drawbacks around providing state proofs for historical queries.

Backwards Compatibility

This ADR proposes changes to the storage implementation in the Cosmos SDK through an entirely new package. Interfaces may be borrowed and extended from existing types that exist in store, but no existing implementations or interfaces will be broken or modified.

Positive

  • Improved performance of independent SS and SC layers
  • Reduced layers of abstraction making storage primitives easier to understand
  • Atomic commitments for SC
  • Redesign of storage types and interfaces will allow for greater experimentation such as different physical storage backends and different commitment schemes for different application modules

Negative

  • Providing proofs for historical state is challenging

Neutral

  • Keeping IAVL as the primary commitment data structure, although drastic performance improvements are being made

Further Discussions

Module Storage Control

Many modules store secondary indexes that are typically solely used to support client queries, but are actually not needed for the state machine’s state transitions. What this means is that these indexes technically have no reason to exist in the SC layer at all, as they take up unnecessary space. It is worth exploring what an API would look like to allow modules to indicate what (key, value) pairs they want to be persisted in the SC layer, implicitly indicating the SS layer as well, as opposed to just persisting the (key, value) pair only in the SS layer.

Historical State Proofs

It is not clear what the importance or demand is within the community of providing commitment proofs for historical state. While solutions can be devised such as rebuilding trees on the fly based on state snapshots, it is not clear what the performance implications are for such solutions.

Physical DB Backends

This ADR proposes usage of RocksDB to utilize user-defined timestamps as a versioning mechanism. However, other physical DB backends are available that may offer alternative ways to implement versioning while also providing performance improvements over RocksDB. E.g. PebbleDB supports MVCC timestamps as well, but we’ll need to explore how PebbleDB handles compaction and state growth over time.

References