变更记录

  • 2022-08-02:初始草案
  • 2023-03-02:补充集成测试的精确定义
  • 2023-03-23:补充 E2E 测试的精确定义

状态

已提议,部分已实现

摘要

SDK 近期为拆分单体根 Go 模块所做的工作,暴露了我们测试范式中的不足与不一致之处。本 ADR 明确了讨论测试范围时应使用的通用语言,并提出了各个范围下测试的理想状态。

背景

ADR-053:Go 模块重构 表达了我们希望将 SDK 组成多个可独立版本化的 Go 模块,而 ADR-057:应用组装 提供了一种通过依赖注入拆分模块间依赖的方法。正如 EPIC:将所有 SDK 模块拆分为独立 Go 模块 中所描述的,模块依赖在测试阶段尤其复杂,因为 simapp 被用作设置和运行测试的关键测试夹具。显然,要成功完成该 EPIC 的第 3 和第 4 阶段,必须先解决这一依赖问题。 在 EPIC:通过 Mock 对模块进行单元测试 中,我们曾认为可以通过在每个模块的测试阶段对所有依赖进行 mock 来解开这个戈尔迪之结,但随着这些重构变成对测试套件的彻底重写,关于现有集成测试命运的讨论也开始出现。一种观点认为应当将其全部丢弃,另一种观点则认为集成测试本身仍有其价值,并且在 SDK 的测试体系中应当占有一席之地。 另一个令人困惑的问题是当前 CLI 测试套件的状态,例如 x/auth。在代码中,这些测试被称为集成测试,但实际上它们通过启动一个 Tendermint 节点和完整应用来充当端到端测试。EPIC:重写并简化 CLI 测试 说明了使用 mock 的 CLI 测试理想状态,但并未讨论端到端测试在 SDK 中可能扮演的角色。 接下来,我们识别出三种测试范围:单元、集成、E2E(端到端),并尝试定义各自的边界、它们的不足之处(无论是真实存在还是人为施加),以及它们在 SDK 中的理想状态。

单元测试

单元测试以与代码库其余部分隔离的方式,验证单个模块(例如 /x/bank)或包(例如 /client)中的代码。在这一范围内,我们识别出两类单元测试:说明性 与 旅程式。下面的定义大量借鉴了 The BDD Books - Formulation 第 1.3 节。 说明性 测试会在隔离环境中验证模块中的一个原子部分,在这种情况下,我们可能会对模块的其他部分进行夹具设置或 mock。 对整个模块功能进行验证、但其依赖均已被 mock 的测试属于 旅程式。这类测试有些类似集成测试,因为它们会同时覆盖许多内容,但仍然使用 mock。 示例 1,旅程式测试与说明性测试的对比:depinject 的 BDD 风格测试展示了我们如何在不需要太多代码的情况下,快速构建大量说明性用例来展示行为规则,同时保持较高层次的可读性。 示例 2:depinject 表驱动测试 示例 3:Bank keeper 测试 - 向 keeper 构造函数提供了 AccountKeeper 的一个 mock 实现。

局限性

某些模块即使超出测试阶段也依然紧密耦合。最近一份关于 bank -> auth 的依赖报告发现,bank 中总共有 274 处对 auth 的使用,其中 50 处位于生产代码中,224 处位于测试中。这种紧耦合可能意味着这些模块应该被合并,或者需要通过重构来抽象出将这些模块绑定在一起的核心类型引用。它也可能表明,除了基于 mock 的单元测试之外,这些模块还应在集成测试中一起测试。 在某些情况下,为一个拥有大量 mock 依赖的模块设置测试用例会相当繁琐,而得到的测试可能更多只是说明 mocking 框架按预期工作,而不是作为相互依赖模块行为的功能性测试。

集成测试

集成测试用于定义并验证任意数量模块和/或应用子系统之间的关系。 集成测试的组装由 depinject 提供,同时还有一些辅助代码用于启动一个正在运行的应用。随后便可以对该运行中应用的一部分进行测试。应用生命周期不同阶段中的某些输入,应当在不过度关注组件内部实现的前提下产生不变的输出。这种黑盒测试的范围比单元测试更大。 示例 1:client/grpc_query_test/TestGRPCQuery - 这个测试放在 /client 中并不合适,但它测试了(至少)runtime 和 bank 从启动、创世到查询阶段的生命周期。它还借助 QueryServiceTestHelper,在不通过网络传输字节的情况下验证了客户端和查询服务器的适配性。 示例 2:x/evidence Keeper 集成测试 - 启动了一个由 8 个模块组成的应用,并在集成测试套件中使用了 5 个 keeper。该套件中的一个测试验证了 HandleEquivocationEvidence,其中包含许多与 staking keeper 的交互。 示例 3:集成套件的应用配置也可以通过 golang 来指定,而不是像上面那样使用 YAML;既可以是静态配置,也可以是动态配置。

局限性

设置特定的输入状态可能会更具挑战,因为应用是从零状态启动的。其中一部分问题可以通过良好的测试夹具抽象以及对这些抽象自身的测试来缓解。测试也可能更脆弱,较大的重构可能会以难以理解的错误、出乎意料地影响应用初始化。另一方面,这也可以被视为一种好处;事实上,SDK 当前的集成测试就在早期 app-wiring 重构阶段帮助定位了逻辑错误。

模拟测试

模拟测试(也称生成式测试)是集成测试的一种特殊形式,其中会针对运行中的 simapp 执行确定性的随机模块操作,不断在链上构建区块,直到达到指定高度。对于模块操作导致的状态转换,不会做出特定断言,但任何错误都会中止并判定模拟失败。由于 simapp 中包含 crisis,并且模拟会在每个区块结束时运行 EndBlockers,因此任何模块不变量违规也会导致模拟失败。 模块必须实现 AppModuleSimulation.WeightedOperations 来定义其模拟操作。请注意,并非所有模块都实现了这一点,这可能表明当前模拟测试覆盖率存在缺口。 未返回模拟操作的模块:
  • auth
  • evidence
  • mint
  • params
一个独立的二进制程序 runsim 负责触发其中一部分测试并管理其生命周期。

局限性

  • 一次成功运行可能耗时很长,在 CI 中每次模拟需要 7 到 10 分钟。
  • 有时会在看似成功的情况下发生超时,而且没有任何原因说明。
  • CI 失败时不会提供有用的错误信息,因此开发者需要在本地运行模拟来复现问题。

E2E 测试

端到端测试以尽可能接近生产环境的方式,验证我们所理解的整个系统。目前这些测试位于 tests/e2e,并依赖 testutil/network 启动一个进程内 Tendermint 节点。 应用应尽可能以最小化方式构建,只覆盖目标功能。SDK 使用的应用仅包含测试所需的模块。建议应用开发者在 E2E 测试中使用自己的应用。

局限性

总体而言,端到端测试的局限在于编排复杂度和计算成本。为了启动并运行一个接近生产环境的环境,需要额外的脚手架,而且这一过程的启动和运行时间都比单元测试或集成测试长得多。 Tendermint 代码中的全局锁会导致有状态的启动/停止过程在 CI 环境中运行时有时卡住,或间歇性失败。 E2E 测试的范围还与命令行接口测试纠缠在一起。

决策

我们接受这些测试范围,并为每一种识别出以下决策点。
范围应用类型使用 Mock?
单元无是
集成integration helpers部分使用
模拟最小化应用否
E2E最小化应用否
上述决策适用于 SDK。应用开发者应使用自己的完整应用来测试其应用,而不是最小化应用。

单元测试

所有模块都必须具备基于 mock 的单元测试覆盖。 在单元测试中,说明性测试的数量应多于旅程式测试。 单元测试的数量应多于集成测试。 单元测试不得引入超出生产代码中已存在依赖之外的额外依赖。 当按照 EPIC:通过 mock 对模块进行单元测试 引入模块单元测试,导致某个集成测试套件几乎被完全重写时,该测试套件应被保留并移动到 /tests/integration。我们接受由此带来的测试逻辑重复,但建议通过增加说明性测试来改进单元测试套件。

集成测试

所有集成测试都应位于 /tests/integration,即使它们没有引入额外的模块依赖也是如此。 为帮助限制范围和复杂度,建议在应用启动时使用尽可能少的模块,也就是说,不要依赖 simapp。 集成测试的数量应多于 E2E 测试。

模拟测试

模拟测试应使用最小化应用(通常通过 app wiring),并位于 /x/{moduleName}/simulation 下。

E2E 测试

现有的 e2e 测试应迁移为集成测试,去除对测试网络和进程内 Tendermint 节点的依赖,以确保我们不会损失测试覆盖率。 E2E REST 运行器应从进程内 Tendermint 迁移为通过 dockertest 由 Docker 驱动的运行器。 应编写用于验证完整网络升级的 E2E 测试。 现有 e2e 测试中的 CLI 测试部分应使用 PR#12706 中展示的网络模拟方式重写。

后果

正面

  • 测试覆盖率提高
  • 测试组织得到改进
  • 模块中的依赖图规模减小
  • simapp 不再作为模块依赖
  • 测试代码中引入的模块间依赖被移除
  • 摆脱进程内 Tendermint 后,CI 运行时间缩短

负面

  • 在过渡期间,单元测试与集成测试之间会存在部分测试逻辑重复
  • 使用 dockertest 编写测试的开发体验可能会稍差一些

中性

  • e2e 向 dockertest 迁移需要进行一些探索

进一步讨论

如果测试套件既能以集成模式运行(使用模拟的 tendermint),也能配合 e2e 测试夹具运行(使用真实的 tendermint 和多个节点),可能会很有价值。集成测试夹具可用于更快的运行,e2e 测试夹具可用于更充分的实战化验证。 PR #12847 中已完成 x/gov 的一个 PoC;用于在单元测试中演示 BDD 的工作仍在推进中 [已拒绝]。 鉴于 BDD 规范的优势在于可读性,而其缺点在于编写和维护时的认知负担,目前的共识是将 BDD 保留给 SDK 中那些需要展示复杂规则和模块交互的场景。 更直接或更底层的测试用例将继续依赖 Go 表格测试。 关于在集成测试和 e2e 测试中进行网络模拟的层级,目前仍在持续推进和正式化中。

Changelog

  • 2022-08-02: Initial Draft
  • 2023-03-02: Add precision for integration tests
  • 2023-03-23: Add precision for E2E tests

Status

PROPOSED Partially Implemented

Abstract

Recent work in the SDK aimed at breaking apart the monolithic root go module has highlighted shortcomings and inconsistencies in our testing paradigm. This ADR clarifies a common language for talking about test scopes and proposes an ideal state of tests at each scope.

Context

ADR-053: Go Module Refactoring expresses our desire for an SDK composed of many independently versioned Go modules, and ADR-057: App Wiring offers a methodology for breaking apart inter-module dependencies through the use of dependency injection. As described in EPIC: Separate all SDK modules into standalone go modules, module dependencies are particularly complected in the test phase, where simapp is used as the key test fixture in setting up and running tests. It is clear that the successful completion of Phases 3 and 4 in that EPIC require the resolution of this dependency problem. In EPIC: Unit Testing of Modules via Mocks it was thought this Gordian knot could be unwound by mocking all dependencies in the test phase for each module, but seeing how these refactors were complete rewrites of test suites discussions began around the fate of the existing integration tests. One perspective is that they ought to be thrown out, another is that integration tests have some utility of their own and a place in the SDK’s testing story. Another point of confusion has been the current state of CLI test suites, x/auth for example. In code these are called integration tests, but in reality function as end to end tests by starting up a tendermint node and full application. EPIC: Rewrite and simplify CLI tests identifies the ideal state of CLI tests using mocks, but does not address the place end to end tests may have in the SDK. From here we identify three scopes of testing, unit, integration, e2e (end to end), seek to define the boundaries of each, their shortcomings (real and imposed), and their ideal state in the SDK.

Unit tests

Unit tests exercise the code contained in a single module (e.g. /x/bank) or package (e.g. /client) in isolation from the rest of the code base. Within this we identify two levels of unit tests, illustrative and journey. The definitions below lean heavily on The BDD Books - Formulation section 1.3. Illustrative tests exercise an atomic part of a module in isolation - in this case we might do fixture setup/mocking of other parts of the module. Tests which exercise a whole module’s function with dependencies mocked, are journeys. These are almost like integration tests in that they exercise many things together but still use mocks. Example 1 journey vs illustrative tests - depinject’s BDD style tests, show how we can rapidly build up many illustrative cases demonstrating behavioral rules without very much code while maintaining high level readability. Example 2 depinject table driven tests Example 3 Bank keeper tests - A mock implementation of AccountKeeper is supplied to the keeper constructor.

Limitations

Certain modules are tightly coupled beyond the test phase. A recent dependency report for bank -> auth found 274 total usages of auth in bank, 50 of which are in production code and 224 in test. This tight coupling may suggest that either the modules should be merged, or refactoring is required to abstract references to the core types tying the modules together. It could also indicate that these modules should be tested together in integration tests beyond mocked unit tests. In some cases setting up a test case for a module with many mocked dependencies can be quite cumbersome and the resulting test may only show that the mocking framework works as expected rather than working as a functional test of interdependent module behavior.

Integration tests

Integration tests define and exercise relationships between an arbitrary number of modules and/or application subsystems. Wiring for integration tests is provided by depinject and some helper code starts up a running application. A section of the running application may then be tested. Certain inputs during different phases of the application life cycle are expected to produce invariant outputs without too much concern for component internals. This type of black box testing has a larger scope than unit testing. Example 1 client/grpc_query_test/TestGRPCQuery - This test is misplaced in /client, but tests the life cycle of (at least) runtime and bank as they progress through startup, genesis and query time. It also exercises the fitness of the client and query server without putting bytes on the wire through the use of QueryServiceTestHelper. Example 2 x/evidence Keeper integration tests - Starts up an application composed of 8 modules with 5 keepers used in the integration test suite. One test in the suite exercises HandleEquivocationEvidence which contains many interactions with the staking keeper. Example 3 - Integration suite app configurations may also be specified via golang (not YAML as above) statically or dynamically.

Limitations

Setting up a particular input state may be more challenging since the application is starting from a zero state. Some of this may be addressed by good test fixture abstractions with testing of their own. Tests may also be more brittle, and larger refactors could impact application initialization in unexpected ways with harder to understand errors. This could also be seen as a benefit, and indeed the SDK’s current integration tests were helpful in tracking down logic errors during earlier stages of app-wiring refactors.

Simulations

Simulations (also called generative testing) are a special case of integration tests where deterministically random module operations are executed against a running simapp, building blocks on the chain until a specified height is reached. No specific assertions are made for the state transitions resulting from module operations but any error will halt and fail the simulation. Since crisis is included in simapp and the simulation runs EndBlockers at the end of each block any module invariant violations will also fail the simulation. Modules must implement AppModuleSimulation.WeightedOperations to define their simulation operations. Note that not all modules implement this which may indicate a gap in current simulation test coverage. Modules not returning simulation operations:
  • auth
  • evidence
  • mint
  • params
A separate binary, runsim, is responsible for kicking off some of these tests and managing their life cycle.

Limitations

  • A success may take a long time to run, 7-10 minutes per simulation in CI.
  • Timeouts sometimes occur on apparent successes without any indication why.
  • Useful error messages not provided on failure from CI, requiring a developer to run the simulation locally to reproduce.

E2E tests

End to end tests exercise the entire system as we understand it in as close an approximation to a production environment as is practical. Presently these tests are located at tests/e2e and rely on testutil/network to start up an in-process Tendermint node. An application should be built as minimally as possible to exercise the desired functionality. The SDK uses an application will only the required modules for the tests. The application developer is adviced to use its own application for e2e tests.

Limitations

In general the limitations of end to end tests are orchestration and compute cost. Scaffolding is required to start up and run a prod-like environment and the this process takes much longer to start and run than unit or integration tests. Global locks present in Tendermint code cause stateful starting/stopping to sometimes hang or fail intermittently when run in a CI environment. The scope of e2e tests has been complected with command line interface testing.

Decision

We accept these test scopes and identify the following decisions points for each.
ScopeApp TypeMocks?
UnitNoneYes
Integrationintegration helpersSome
Simulationminimal appNo
E2Eminimal appNo
The decision above is valid for the SDK. An application developer should test their application with their full application instead of the minimal app.

Unit Tests

All modules must have mocked unit test coverage. Illustrative tests should outnumber journeys in unit tests. Unit tests should outnumber integration tests. Unit tests must not introduce additional dependencies beyond those already present in production code. When module unit test introduction as per EPIC: Unit testing of modules via mocks results in a near complete rewrite of an integration test suite the test suite should be retained and moved to /tests/integration. We accept the resulting test logic duplication but recommend improving the unit test suite through the addition of illustrative tests.

Integration Tests

All integration tests shall be located in /tests/integration, even those which do not introduce extra module dependencies. To help limit scope and complexity, it is recommended to use the smallest possible number of modules in application startup, i.e. don’t depend on simapp. Integration tests should outnumber e2e tests.

Simulations

Simulations shall use a minimal application (usually via app wiring). They are located under /x/{moduleName}/simulation.

E2E Tests

Existing e2e tests shall be migrated to integration tests by removing the dependency on the test network and in-process Tendermint node to ensure we do not lose test coverage. The e2e rest runner shall transition from in process Tendermint to a runner powered by Docker via dockertest. E2E tests exercising a full network upgrade shall be written. The CLI testing aspect of existing e2e tests shall be rewritten using the network mocking demonstrated in PR#12706.

Consequences

Positive

  • test coverage is increased
  • test organization is improved
  • reduced dependency graph size in modules
  • simapp removed as a dependency from modules
  • inter-module dependencies introduced in test code are removed
  • reduced CI run time after transitioning away from in process Tendermint

Negative

  • some test logic duplication between unit and integration tests during transition
  • test written using dockertest DX may be a bit worse

Neutral

  • some discovery required for e2e transition to dockertest

Further Discussions

It may be useful if test suites could be run in integration mode (with mocked tendermint) or with e2e fixtures (with real tendermint and many nodes). Integration fixtures could be used for quicker runs, e2e fixures could be used for more battle hardening. A PoC x/gov was completed in PR #12847 is in progress for unit tests demonstrating BDD [Rejected]. Observing that a strength of BDD specifications is their readability, and a con is the cognitive load while writing and maintaining, current consensus is to reserve BDD use for places in the SDK where complex rules and module interactions are demonstrated. More straightforward or low level test cases will continue to rely on go table tests. Levels are network mocking in integration and e2e tests are still being worked on and formalized.