概述
Block-STM 通过乐观并发控制在 FinalizeBlock 期间实现交易并行执行,从而提升区块处理吞吐量。
Block-STM 是一种最早发表于 Block-STM 论文 并为 Aptos 区块链实现的算法。随后,Cronos 区块链的开发者在 go-block-stm 中使用 Go 为兼容 Cosmos SDK 的链实现了该算法。
该库随后被 fork 并直接集成到 Cosmos SDK 中,同时对 baseapp 和 store 包进行了配套修改。此后又在原始实现基础上进行了多项变更和改进,以进一步优化内存和时间性能。
算法概览
Block-STM 实现了一种乐观并发控制形式,以支持交易并行执行。它通过在 SDK 的 IAVL 存储层之上实现读集合与写集合跟踪来完成这一点。再结合区块提案提供的交易绝对顺序,在验证阶段判断任意两笔已执行交易是否存在冲突的存储访问。如果发生冲突的存储访问,该算法会根据提案中的顺序,对冲突交易进行重新执行和重新验证。
Block-STM 目前仅集成到 FinalizeBlock 执行阶段,这意味着在共识就区块达成一致之前,相关代码路径不会被访问。未来该算法可能会扩展以支持不同的执行模型,但截至目前,它要求输入完整区块,并在整个区块执行完成后返回结果。因此,预期 Block-STM 产生的结果应与串行执行完全一致。换言之,Block-STM 并行执行产生的 AppHash 应等于默认串行交易执行器产生的 AppHash。
安全部署实践
鉴于 Block-STM 执行器是一个通用的并行执行引擎,我们建议针对每个应用分别进行大量测试,并采用分阶段发布方式。
由于并行执行纯粹是一种性能优化,应用在使用 Block-STM 时计算出的 AppHash 应与通过默认 TxRunner 启用串行执行时相同。这使团队可以只在部分节点上启用并行执行,例如在 API 节点而非验证者节点上启用,或者仅在分布式验证者集群的一部分节点上启用。
长时间以并行和串行混合的节点集群运行,可在发生故障时将影响范围降到最低。
我们已经尽可能对核心 SDK 消息类型进行了测试,但由于 Cosmos SDK 允许任意创建消息,因此不可能验证所有现有消息类型及其所有工作流组合。每个在生产环境中集成 Block-STM 的团队,都应对自己的消息类型在正确性和性能两方面进行验证。
注意:我们明确__尚未__验证使用 Block-STM 运行 CosmWasm 消息类型的支持情况。如果你的链使用 CosmWasm,请在启用前自行完成验证。
Block-STM 通过 SDK 的 MultiStore 接口中的依赖跟踪机制工作。任何位于 store 之外、可能导致执行产生有状态变化的数据,都是引发非确定性的最大风险。总体上,我们建议避免使用缓存数据、内存 store,或在 Store 作用域之外持久化任何状态。
应用集成
并行执行的集成被抽象为两个接口:DeliverTxFunc 和 TxRunner。
// DeliverTxFunc is the function called for each transaction in order to produce
// a single ExecTxResult. `memTx` is an optional in-memory representation of
// the transaction, which can be used to avoid decoding the transaction.
type DeliverTxFunc func(
tx []byte,
memTx Tx,
ms storetypes.MultiStore,
txIndex int,
incarnationCache map[string]any,
) *abci.ExecTxResult
// TxRunner defines an interface for types which can be used to execute the
// DeliverTxFunc. It should return an array of *abci.ExecTxResult corresponding
// to the result of executing each transaction provided to the Run function.
type TxRunner interface {
Run(
ctx context.Context,
ms storetypes.MultiStore,
txs [][]byte,
deliverTx DeliverTxFunc,
) ([]*abci.ExecTxResult, error)
}
TxRunner 是开发者接入应用的主要接口。baseapp 包提供了设置它的方式:
func (app *BaseApp) SetBlockSTMTxRunner(txRunner sdk.TxRunner) {
app.txRunner = txRunner
}
Runner 实现
baseapp/txnrunner 包中提供了 TxRunner 的两种实现:
NewDefaultRunner(txDecoder sdk.TxDecoder) *DefaultRunner
NewSTMRunner(
txDecoder sdk.TxDecoder,
stores []storetypes.StoreKey,
workers int,
estimate bool,
coinDenom func(storetypes.MultiStore) string,
) *STMRunner
NewDefaultRunner 是 BaseApp 默认使用的实现,提供串行执行,而不会走 Block-STM 的代码路径。你无需显式接入它。
NewSTMRunner 会构造一个使用并行执行的 runner。其参数如下:
-
txDecoder — 标准的 sdk.TxDecoder,在任何 SDK 应用中都可直接获得。
-
stores — 应用中使用的全部 store key 列表。由于 Block-STM 需要跨交易跟踪 store 使用情况,因此必须传入所有模块级 store key。以下是摘自 Cosmos EVM 的示例:
keys := storetypes.NewKVStoreKeys(
authtypes.StoreKey, banktypes.StoreKey, stakingtypes.StoreKey,
minttypes.StoreKey, distrtypes.StoreKey, slashingtypes.StoreKey,
govtypes.StoreKey, consensusparamtypes.StoreKey,
upgradetypes.StoreKey, feegrant.StoreKey, evidencetypes.StoreKey,
authzkeeper.StoreKey,
// IBC keys
ibcexported.StoreKey, ibctransfertypes.StoreKey,
// Cosmos EVM store keys
evmtypes.StoreKey, feemarkettypes.StoreKey, erc20types.StoreKey,
)
oKeys := storetypes.NewObjectStoreKeys(
banktypes.ObjectStoreKey, evmtypes.ObjectKey,
)
var nonTransientKeys []storetypes.StoreKey
for _, k := range keys {
nonTransientKeys = append(nonTransientKeys, k)
}
for _, k := range oKeys {
nonTransientKeys = append(nonTransientKeys, k)
}
workers — 并行 worker 数量。实验表明,超过系统硬件并行度后,收益会逐渐下降。推荐值为:
import "runtime"
workers := min(runtime.GOMAXPROCS(0), runtime.NumCPU())
-
estimate — 控制系统是否应在执行前主动判断交易的读写冲突。在所有情况下都应将其设为 true。
-
coinDenom — 一个在运行时返回 staking coin denom 的函数。估算阶段会使用它推断收取手续费时 bank 模块中的哪些 key 会被修改。可以使用硬编码值;该值应为你的链的 bond denom。
完整接入示例
以下是摘自 Cosmos EVM evmd 应用的完整示例:
bApp.SetBlockSTMTxRunner(txnrunner.NewSTMRunner(
encodingConfig.TxConfig.TxDecoder(),
nonTransientKeys,
min(goruntime.GOMAXPROCS(0), goruntime.NumCPU()),
true,
func(ms storetypes.MultiStore) string { return sdk.DefaultBondDenom },
))
并行交易优化
接入 Block-STM 后,你一开始可能会注意到,大多数区块的执行速度反而比串行执行更慢。这是因为只要任意两笔交易存在冲突的读或写,就会产生重新执行交易的额外开销。要获得性能收益,你需要减少交易之间的存储访问冲突。
PR #26005 就是一个例子,其中新账户创建需要分配一个 ID,而该 ID 的值通过 x/auth 模块中的单个 key 读取并递增。该 PR 将账户 ID 生成改为使用确定性 UUID 生成,而不是依赖会产生冲突的存储位置。这样一来,同一区块中多笔各自创建新账户的交易将不再访问这个 key,因此可以并行运行而无需重新执行。
SDK 和 Cosmos EVM 已经针对常见交易类型(例如 bank 转账和 EVM gas 转账)完成了相关优化。以下步骤说明所需的额外配置。
启用虚拟费用归集(仅 EVM)
这会改变 EVM 交易的费用归集方式:不再在交易执行期间使用常规转账,而是在 EndBlocker 中将费用累计到 fee collector 模块。
app.EVMKeeper.EnableVirtualFeeCollection()
在 Bank Keeper 中设置 Object Store
这使 bank keeper 可以在 EndBlocker 中归集费用,而不再要求每笔交易都直接向 FeeCollector 模块账户发送费用。
app.BankKeeper = app.BankKeeper.WithObjStoreKey(oKeys[banktypes.ObjectStoreKey])
自定义模块
针对常见交易并行化的其他变更都已以无需配置的方式完成。
对于自定义交易类型或自定义模块,可能仍需要额外调整 KV store 的访问模式。目前还没有通用方法。对于必须访问同一存储 key 的功能,常见模式是将中间值写入 transient 或 object 存储,并在所有交易执行完成后通过 EndBlocker 统一汇总这些值。
基准测试
- 机器: Apple M3 Pro,11 核
- 操作系统: macOS (Darwin 25.3.0)
- 包:
github.com/cosmos/cosmos-sdk/internal/blockstm
随机负载(10k 笔交易,100 个 key)
| 工作线程数 | ns/op | B/op | allocs/op | 加速比 |
|---|
| 串行 | 1,169M | 9.9M | 220K | 1.0x |
| 1 | 1,208M | 36.3M | 544K | 0.97x |
| 5 | 324M | 37.1M | 552K | 3.6x |
| 10 | 218M | 43.7M | 621K | 5.4x |
| 15 | 211M | 77.4M | 975K | 5.5x |
| 20 | 226M | 78.0M | 982K | 5.2x |
无冲突负载(10k 笔交易)
| 工作线程数 | ns/op | B/op | allocs/op | 加速比 |
|---|
| 串行 | 1,381M | 11.7M | 221K | 1.0x |
| 1 | 1,358M | 80.3M | 1,095K | 1.0x |
| 5 | 291M | 81.1M | 1,103K | 4.8x |
| 10 | 209M | 83.5M | 1,131K | 6.6x |
| 15 | 200M | 83.7M | 1,135K | 6.9x |
| 20 | 204M | 83.8M | 1,136K | 6.8x |
最坏情况负载(完全冲突,10k 笔交易)
| 工作线程数 | ns/op | B/op | allocs/op | 加速比 |
|---|
| 串行 | 1,239M | 9.6M | 220K | 1.0x |
| 1 | 1,280M | 34.5M | 520K | 0.97x |
| 5 | 295M | 42.8M | 607K | 4.2x |
| 10 | 224M | 70.1M | 899K | 5.5x |
| 15 | 246M | 77.3M | 980K | 5.0x |
| 20 | 262M | 77.4M | 980K | 4.7x |
遍历负载(10k 笔交易,100 个 key)
| 工作线程数 | ns/op | B/op | allocs/op | 加速比 |
|---|
| 串行 | 1,286M | 16.5M | 290K | 1.0x |
| 1 | 1,332M | 75.2M | 843K | 0.97x |
| 5 | 288M | 76.5M | 855K | 4.5x |
| 10 | 252M | 123.1M | 1,280K | 5.1x |
| 15 | 317M | 359.5M | 3,405K | 4.1x |
| 20 | 319M | 363.5M | 3,419K | 4.0x |
关键结论
- 在无冲突工作负载下,15 个 worker 时峰值加速比达到 6.9x
- 超过 10 到 15 个 worker 后,收益开始递减,同时内存使用显著增加
- 即使在最差情况(完全冲突)下,10 个 worker 时仍可实现约 5x 的加速
iterate 工作负载在超过 10 个 worker 后出现性能下降,可能是由于范围读取上的竞争加剧所致(15 个及以上 worker 时,内存使用量跃升约 3x)
Synopsis
Block-STM enables parallel execution of transactions during FinalizeBlock, using optimistic concurrency control to improve block processing throughput.
Background
Block-STM is an algorithm originally published in the Block-STM paper and implemented for the Aptos blockchain. The algorithm was then written for Cosmos SDK compatible chains in Go by developers for the Cronos blockchain in go-block-stm.
This library was forked and directly integrated into the Cosmos SDK with accompanying changes to the baseapp and store packages. Subsequent changes and improvements have been made on top of the original implementation to further optimize performance in both memory and time.
Algorithm Overview
Block-STM implements a form of optimistic concurrency control to enable parallel execution of transactions. It does this by implementing read and write set tracking on top of the SDK’s IAVL storage layer. This, combined with the absolute ordering of transactions provided by the block proposal, is used in a validation phase which determines if any two executed transactions have conflicting storage access. In the case of conflicting storage access, the algorithm provides a means for re-execution and re-validation of the conflicting transactions based on the ordering in the proposal.
Block-STM is currently only integrated into the FinalizeBlock phase of execution, meaning the code path is never accessed until the block is agreed upon in consensus. It is possible that the algorithm may be extended in the future to support different execution models, but as of right now it expects a complete block and returns its result after the entire block has been executed. For this reason, Block-STM is expected to produce identical results to serial execution. In other words, the AppHash produced by Block-STM’s parallel execution should be equal to the AppHash produced by the default serial transaction runner.
Safe Deployment Practices
Given the Block-STM executor is a general purpose parallel execution engine, we recommend a phased rollout with
extensive
testing for each application individually.
Since parallel execution is purely a performance optimization, applications should expect to calculate the same
AppHash when using Block-STM as they would when serial execution is enabled via the default TxRunner. This allows teams
to turn on parallel execution for a fraction of their nodes—API nodes instead of validators or on a portion of a
distributed validator cluster for example.
Running with a mixed fleet of parallel and serial execution for an extended time should minimize blast radius in the
event that a failure occurs.
We have done testing on as many of the core SDK message types as possible, but given the Cosmos SDK allows arbitrary
message creation it will be impossible to validate all message types that exist and all combinations of workflows.
Each team integrating Block-STM in production should validate their own message types for both correctness and performance.
NOTE: We specifically have not validated support for CosmWasm message types run using Block-STM. Run independent
validation before enabling if your chain uses CosmWasm.
Block-STM works via dependency tracking within the SDK’s MultiStore interface. Any data which could cause stateful
changes to execution that lives outside the store poses the biggest risk for indeterminism. We recommend in general
avoiding the use of cached data, in memory stores, or persisting any state outside the scope of a Store.
App Integration
Integration of parallel execution is abstracted into two interfaces: DeliverTxFunc and TxRunner.
// DeliverTxFunc is the function called for each transaction in order to produce
// a single ExecTxResult. `memTx` is an optional in-memory representation of
// the transaction, which can be used to avoid decoding the transaction.
type DeliverTxFunc func(
tx []byte,
memTx Tx,
ms storetypes.MultiStore,
txIndex int,
incarnationCache map[string]any,
) *abci.ExecTxResult
// TxRunner defines an interface for types which can be used to execute the
// DeliverTxFunc. It should return an array of *abci.ExecTxResult corresponding
// to the result of executing each transaction provided to the Run function.
type TxRunner interface {
Run(
ctx context.Context,
ms storetypes.MultiStore,
txs [][]byte,
deliverTx DeliverTxFunc,
) ([]*abci.ExecTxResult, error)
}
The TxRunner is the primary interface that developers wire into their application. The baseapp package provides an option to set it up:
func (app *BaseApp) SetBlockSTMTxRunner(txRunner sdk.TxRunner) {
app.txRunner = txRunner
}
Runner Implementations
Two implementations of TxRunner are provided in the baseapp/txnrunner package:
NewDefaultRunner(txDecoder sdk.TxDecoder) *DefaultRunner
NewSTMRunner(
txDecoder sdk.TxDecoder,
stores []storetypes.StoreKey,
workers int,
estimate bool,
coinDenom func(storetypes.MultiStore) string,
) *STMRunner
NewDefaultRunner is used by BaseApp by default and provides serial execution without using the Block-STM code paths. You do not need to wire this in explicitly.
NewSTMRunner constructs a runner which uses parallel execution. Its parameters are:
-
txDecoder — A standard sdk.TxDecoder, readily available in any SDK application.
-
stores — A list of every store key used in your application. Since Block-STM needs to track store usage across transactions, it must be passed all module-level store keys. Here is an example taken from the Cosmos EVM:
keys := storetypes.NewKVStoreKeys(
authtypes.StoreKey, banktypes.StoreKey, stakingtypes.StoreKey,
minttypes.StoreKey, distrtypes.StoreKey, slashingtypes.StoreKey,
govtypes.StoreKey, consensusparamtypes.StoreKey,
upgradetypes.StoreKey, feegrant.StoreKey, evidencetypes.StoreKey,
authzkeeper.StoreKey,
// IBC keys
ibcexported.StoreKey, ibctransfertypes.StoreKey,
// Cosmos EVM store keys
evmtypes.StoreKey, feemarkettypes.StoreKey, erc20types.StoreKey,
)
oKeys := storetypes.NewObjectStoreKeys(
banktypes.ObjectStoreKey, evmtypes.ObjectKey,
)
var nonTransientKeys []storetypes.StoreKey
for _, k := range keys {
nonTransientKeys = append(nonTransientKeys, k)
}
for _, k := range oKeys {
nonTransientKeys = append(nonTransientKeys, k)
}
workers — The number of parallel workers. Experimentation has shown diminishing returns above your system’s hardware parallelism. The recommended value is:
import "runtime"
workers := min(runtime.GOMAXPROCS(0), runtime.NumCPU())
-
estimate — Controls whether the system should proactively determine transaction read/write conflicts before execution. Set this to true in all cases.
-
coinDenom — A function that returns the staking coin denom at runtime. This is used during estimation to reason about which keys in the bank module will be modified when fees are collected. A hard-coded value is acceptable; the value should be your chain’s bond denom.
Full Wiring Example
Here is a complete example taken from the Cosmos EVM’s evmd application:
bApp.SetBlockSTMTxRunner(txnrunner.NewSTMRunner(
encodingConfig.TxConfig.TxDecoder(),
nonTransientKeys,
min(goruntime.GOMAXPROCS(0), goruntime.NumCPU()),
true,
func(ms storetypes.MultiStore) string { return sdk.DefaultBondDenom },
))
Parallel Transaction Optimization
Once Block-STM is wired in, you may initially notice that most blocks execute slower than with serial execution. This is due to the overhead of re-executing transactions when any two have conflicting reads or writes. To realize performance gains, you need to reduce storage access conflicts between transactions.
An example of this can be seen in PR #26005 where new account creation involved assigning an ID whose value was retrieved and incremented via a single key in the x/auth module. The linked PR converts account ID generation to use deterministic UUID generation instead of relying on a conflicting storage location. The result is that multiple transactions in the same block which each create new accounts no longer access this key and can be run in parallel without re-executions.
Work has already been done within the SDK and Cosmos EVM for common transaction types such as bank sends and EVM gas sends. The following steps describe the additional configuration needed.
Enable Virtual Fee Collection (EVM-specific)
This alters how fee collection works for EVM transactions, accumulating fees to the fee collector module in the EndBlocker instead of using regular sends during transaction execution.
app.EVMKeeper.EnableVirtualFeeCollection()
Set Up the Object Store in the Bank Keeper
This enables the bank keeper to collect fees in the EndBlocker instead of requiring every transaction to send fees directly to the FeeCollector module account.
app.BankKeeper = app.BankKeeper.WithObjStoreKey(oKeys[banktypes.ObjectStoreKey])
Custom Modules
All other changes to parallelize common transactions were done in a way that does not require configuration.
For custom transaction types or custom modules, additional changes to KV store access patterns may be required. There is no generalized approach for this yet. The common pattern for functionality that requires access to the same storage key is to write intermediate values to transient or object storage and use an EndBlocker to collect the values after all transaction execution completes.
Benchmarks
Environment
- Machine: Apple M3 Pro, 11 cores
- OS: macOS (Darwin 25.3.0)
- Package:
github.com/cosmos/cosmos-sdk/internal/blockstm
Random Workload (10k txs, 100 keys)
| Workers | ns/op | B/op | allocs/op | Speedup |
|---|
| sequential | 1,169M | 9.9M | 220K | 1.0x |
| 1 | 1,208M | 36.3M | 544K | 0.97x |
| 5 | 324M | 37.1M | 552K | 3.6x |
| 10 | 218M | 43.7M | 621K | 5.4x |
| 15 | 211M | 77.4M | 975K | 5.5x |
| 20 | 226M | 78.0M | 982K | 5.2x |
No-Conflict Workload (10k txs)
| Workers | ns/op | B/op | allocs/op | Speedup |
|---|
| sequential | 1,381M | 11.7M | 221K | 1.0x |
| 1 | 1,358M | 80.3M | 1,095K | 1.0x |
| 5 | 291M | 81.1M | 1,103K | 4.8x |
| 10 | 209M | 83.5M | 1,131K | 6.6x |
| 15 | 200M | 83.7M | 1,135K | 6.9x |
| 20 | 204M | 83.8M | 1,136K | 6.8x |
Worst-Case Workload (full conflict, 10k txs)
| Workers | ns/op | B/op | allocs/op | Speedup |
|---|
| sequential | 1,239M | 9.6M | 220K | 1.0x |
| 1 | 1,280M | 34.5M | 520K | 0.97x |
| 5 | 295M | 42.8M | 607K | 4.2x |
| 10 | 224M | 70.1M | 899K | 5.5x |
| 15 | 246M | 77.3M | 980K | 5.0x |
| 20 | 262M | 77.4M | 980K | 4.7x |
Iterate Workload (10k txs, 100 keys)
| Workers | ns/op | B/op | allocs/op | Speedup |
|---|
| sequential | 1,286M | 16.5M | 290K | 1.0x |
| 1 | 1,332M | 75.2M | 843K | 0.97x |
| 5 | 288M | 76.5M | 855K | 4.5x |
| 10 | 252M | 123.1M | 1,280K | 5.1x |
| 15 | 317M | 359.5M | 3,405K | 4.1x |
| 20 | 319M | 363.5M | 3,419K | 4.0x |
Key Takeaways
- Peak speedup is 6.9x at 15 workers on the no-conflict workload
- Diminishing returns beyond 10–15 workers, with memory usage increasing significantly
- Even the worst-case (full conflict) scenario achieves ~5x speedup at 10 workers
- The iterate workload shows performance degradation beyond 10 workers, likely due to increased contention on range reads (memory usage jumps ~3x at 15+ workers)