变更记录
- 2025-03-10:初始草案 (@aaronc)
状态
提议中:尚未实现摘要
从历史上看,若干与编码和签名模式相关的问题导致 Cosmos SDK 交易存在这样一种可能: 在不使签名失效的前提下,交易可以被重新编码,从而改变其哈希 (在极少数情况下,甚至改变其含义)。 本文详细说明这些情况、它们的潜在风险、目前已被处理到什么程度, 并为后续改进提供建议。评审
关于 Cosmos SDK 交易,一个天真的假设是:对已提交交易的原始字节做哈希,就能为该交易生成一个安全且唯一的标识符。实际上,交易存在多种可被操纵的方式,使其产生不同的交易字节(以及由此产生的不同哈希),但仍然能够通过签名验证。 本文尝试枚举我们已识别出的各类潜在交易“可塑性”风险,以及它们在不同签名模式下已经或尚未被处理到什么程度。我们还指出了这样一些漏洞:如果开发者未来在没有认真考虑交易编码、签名模式和签名所涉及复杂性的情况下做出修改,就可能引入这些漏洞。与可塑性相关的风险
交易的可塑性会给终端用户带来以下潜在风险:- 未签名的数据可能被添加到交易中,并被状态机处理
- 客户端通常依赖交易哈希检查交易状态,但已提交交易的哈希是否与已处理交易的哈希一致,主要取决于网络参与者是否行为良好,而不是协议的根本性保证
- 交易可能被执行多次(重放保护失效)
可塑性的来源
非确定性的 Protobuf 编码
Cosmos SDK 交易在提交到网络时使用 protobuf 二进制编码。Protobuf 二进制本身并不是一种天然确定性的编码,这意味着同一个逻辑载荷可能存在多种有效的字节表示。从最基本的层面上说,这意味着一般情况下,protobuf 可以被解码后重新编码,从而在不改变字节逻辑含义的前提下生成不同的字节流(因此也会得到不同的哈希)。ADR 027:确定性的 Protobuf 序列化 详细说明了需要做哪些事情,才能生成我们认为“规范化”的、确定性的 protobuf 序列化。简而言之,已经识别出以下编码层面的可塑性来源,并由该规范进行处理:- 字段可以按任意顺序输出
- 默认字段值可以被包含或省略;除非使用了
optional,否则这不会改变含义 - 标量类型的
repeated字段可以使用 packed 或“常规”编码 varint可以包含额外的、会被忽略的位- 可以添加额外字段,而解码器通常只是简单忽略它们。ADR 020 规定,一般来说,这类额外字段应导致消息和交易被拒绝)
SIGN_MODE_DIRECT 时,上述任何可塑性都不会被容忍,因为:
- 消息和扩展的签名必须基于这些字段原始编码后的字节
- 外层交易封装(
TxRaw)必须遵循 ADR 027 规则,否则应被拒绝
SIGN_MODE_LEGACY_AMINO_JSON 签名的交易无法防御上述可塑性,因为被签名的是交易逻辑内容的 JSON 表示。这些逻辑内容可能对应任意数量的有效 protobuf 二进制编码,因此通常来说,Amino JSON 签名无法对交易哈希提供任何保证。
除了需要了解 protobuf 二进制本身的一般非确定性之外,开发者在开发与 protobuf 交易相关的新能力时,还需要特别注意确保未知的 protobuf 字段会被拒绝。Protobuf 序列化格式的设计假设是:编码器已知而解码器未知的数据,可以被解码器安全地忽略。这个假设在 Google 的集中式基础设施这个封闭环境中或许相对安全。然而,在分布式区块链系统中,这一假设通常是不安全的。如果较新的客户端对一个 protobuf 消息进行编码,其中包含面向较新服务端的数据,那么较旧的服务端简单地忽略并丢弃它不理解的指令并不安全。这些指令可能包含交易签名者所依赖的关键信息,仅仅假设它们不重要是不安全的。
ADR 020 为“非关键”字段规定了一些条款,允许旧服务端安全地忽略它们。就实践而言,我还没有见过任何有效的用法。这是维护者在设计中应当了解的一点,但它可能并非必要,甚至也未必 100% 安全。
非确定性的值编码
除了 protobuf 二进制本身存在的非确定性之外,一些 protobuf 字段数据还会使用一种自身也可能不具备确定性的微格式进行编码。以整数或小数编码为例。一些解码器可能允许前导零或尾随零的存在,而不改变其逻辑含义,例如00100 与 100,或 100.00 与 100。因此,如果某个签名模式以确定性方式编码数字,但解码器接受多种表示形式,
那么用户可能是对值 100 进行签名,而实际被编码的是 0100。在整数解码器接受前导零的情况下,Amino JSON 就可能出现这种问题。我认为当前的 Int 实现会拒绝这种情况,不过,
仍然很可能在交易中编码八进制或十六进制表示,而用户签名的却是十进制整数。
签名编码
签名本身使用与具体签名算法相关的微格式进行编码,而这些 微格式有时可能允许非确定性(同一个签名对应多种有效字节表示)。 SDK 当前支持的大多数签名算法实现都应当拒绝非规范字节。 然而,Multisignature protobuf 类型使用的是普通 protobuf 编码,并且不会检查
解码后的字节是否遵循了规范的 ADR 027 规则。因此,多签交易的签名
存在可塑性。
任何新的或自定义的签名算法都必须确保拒绝所有非规范字节,否则即使
使用 SIGN_MODE_DIRECT,也仍然可能因为用非规范表示重新编码签名而导致
交易哈希出现可塑性。
Amino JSON 未覆盖的字段
另一个需要谨慎处理的领域,是用于SIGN_MODE_LEGACY_AMINO_JSON 的 AminoSignDoc(见 aminojson.proto)与 TxBody 和 AuthInfo 的实际内容(见 tx.proto)之间的不一致。
如果向 TxBody 或 AuthInfo 添加了字段,那么这些字段必须要么在 AminoSignDoc 中有对应表示,要么在这些新字段被设置时拒绝 Amino JSON 签名。确保这一点是一个
高度依赖人工的过程,开发者很容易犯这样的错误:更新了 TxBody 或 AuthInfo,
却完全没有关注 Amino JSON 的 GetSignBytes 实现。这会形成一个关键
漏洞,使未签名内容可以进入交易,而签名验证却仍然会
通过。
签名模式总结与建议
SDK 官方支持的签名模式包括SIGN_MODE_DIRECT、SIGN_MODE_TEXTUAL、SIGN_MODE_DIRECT_AUX,
以及 SIGN_MODE_LEGACY_AMINO_JSON。
SIGN_MODE_LEGACY_AMINO_JSON 被钱包广泛使用,并且目前是 Nano Ledger 硬件设备上唯一受支持的签名模式
(尽管 SIGN_MODE_TEXTUAL 的设计目标也包括支持硬件设备)。
SIGN_MODE_DIRECT 是最简单的签名模式,其使用也相当普遍。
SIGN_MODE_DIRECT_AUX 是 SIGN_MODE_DIRECT 的一种变体,可在多签名者交易中由辅助签名者使用,
适用于那些不支付 gas 的签名者。
SIGN_MODE_TEXTUAL 原本旨在替代 SIGN_MODE_LEGACY_AMINO_JSON,但据我们所知,
目前尚未被任何客户端采用,因此并未被实际使用。
目前实现中的 SIGN_MODE_DIRECT 已处理所有已知的可塑性问题。
使用 SIGN_MODE_DIRECT 签名的交易,唯一已知仍可能出现的可塑性
只能来自签名字节本身。
由于签名本身不会被再次签名,任何签名模式都不可能直接解决这个问题,
因此签名算法必须谨慎地拒绝任何非规范编码的签名字节,
以防止可塑性。
对于 Multisignature 类型已知的可塑性,我们应确保在做签名验证时,任何有效签名
都遵循规范的 ADR 027 规则进行编码。
SIGN_MODE_DIRECT_AUX 提供与 SIGN_MODE_DIRECT 相同级别的安全性,因为
SignDocDirectAux会对原始编码后的TxBody字节进行签名,并且- 使用
SIGN_MODE_DIRECT_AUX的交易仍然要求主签名者使用SIGN_MODE_DIRECT对交易进行签名
SIGN_MODE_TEXTUAL 也提供与 SIGN_MODE_DIRECT 相同级别的安全性,因为它签名的是原始编码后的
TxBody 和 AuthInfo 字节的哈希。
遗憾的是,绝大多数尚未解决的可塑性风险都影响 SIGN_MODE_LEGACY_AMINO_JSON,而这一
签名模式仍然被广泛使用。
建议对 Amino JSON 签名进行以下改进:
- 应将
TxBody和AuthInfo的哈希添加到AminoSignDoc中,以解决编码层面的可塑性 - 在构造
AminoSignDoc时,应使用 protoreflect API 来确保TxBody或AuthInfo中不存在那些在AminoSignDoc中没有映射但却已被设置的字段 - 对于存在于
TxBody或AuthInfo中但不存在于AminoSignDoc中的字段(例如扩展选项),如果可能,应该将其 添加到AminoSignDoc中
测试
为了测试交易是否能够抵抗可塑性, 我们可以开发一套测试用例,并针对所有签名模式运行, 尝试以下方式来操纵交易字节:- 通过更改 protobuf 编码方式
- 重排字段顺序
- 设置默认值
- 为 varint 添加额外比特,或
- 设置新的未知字段
- 修改以字符串编码的整数和小数值,在其前后添加前导零或尾随零
TxBody 或 AuthInfo
中不受 Amino 的 AminoSignDoc 支持的字段,则签名会失败。
在交易解码这一更一般的场景下,我们应编写单元测试以确保:
- 任何不符合 ADR 027 规范编码的
TxRaw字节都会导致解码失败,以及 - 任何顶层交易元素(包括
TxBody、AuthInfo、公钥和消息)如果 设置了未知字段,都会导致该交易被拒绝 (这可确保 ADR 020 的未知字段过滤被正确应用)
参考
Changelog
- 2025-03-10: Initial draft (@aaronc)
Status
PROPOSED: Not ImplementedAbstract
Several encoding and sign mode related issues have historically resulted in the possibility that Cosmos SDK transactions may be re-encoded in such a way as to change their hash (and in rare cases, their meaning) without invalidating the signature. This document details these cases, their potential risks, the extent to which they have been addressed, and provides recommendations for future improvements.Review
One naive assumption about Cosmos SDK transactions is that hashing the raw bytes of a submitted transaction creates a safe unique identifier for the transaction. In reality, there are multiple ways in which transactions could be manipulated to create different transaction bytes (and as a result different hashes) that still pass signature verification. This document attempts to enumerate the various potential transaction “malleability” risks that we have identified and the extent to which they have or have not been addressed in various sign modes. We also identify vulnerabilities that could be introduced if developers make changes in the future without careful consideration of the complexities involved with transaction encoding, sign modes and signatures.Risks Associated with Malleability
The malleability of transactions poses the following potential risks to end users:- unsigned data could get added to transactions and be processed by state machines
- clients often rely on transaction hashes for checking transaction status, but whether or not submitted transaction hashes match processed transaction hashes depends primarily on good network actors rather than fundamental protocol guarantees
- transactions could potentially get executed more than once (faulty replay protection)
Sources of Malleability
Non-deterministic Protobuf Encoding
Cosmos SDK transactions are encoded using protobuf binary encoding when they are submitted to the network. Protobuf binary is not inherently a deterministic encoding meaning that the same logical payload could have several valid bytes representations. In a basic sense, this means that protobuf in general can be decoded and re-encoded to produce a different byte stream (and thus different hash) without changing the logical meaning of the bytes. ADR 027: Deterministic Protobuf Serialization describes in detail what needs to be done to produce what we consider to be a “canonical”, deterministic protobuf serialization. Briefly, the following sources of malleability at the encoding level have been identified and are addressed by this specification:- fields can be emitted in any order
- default field values can be included or omitted, and this doesn’t change meaning unless
optionalis used repeatedfields of scalars may use packed or “regular” encodingvarints can include extra ignored bits- extra fields may be added and are usually simply ignored by decoders. ADR 020 specifies that in general such extra fields should cause messages and transactions to be rejected)
SIGN_MODE_DIRECT none of the above malleabilities will be tolerated because:
- signatures of messages and extensions must be done over the raw encoded bytes of those fields
- the outer tx envelope (
TxRaw) must follow ADR 027 rules or be rejected
SIGN_MODE_LEGACY_AMINO_JSON, however, have no way of protecting against the above malleabilities because what is signed is a JSON representation of the logical contents of the transaction. These logical contents could have any number of valid protobuf binary encodings, so in general there are no guarantees regarding transaction hash with Amino JSON signing.
In addition to being aware of the general non-determinism of protobuf binary, developers need to pay special attention to make sure that unknown protobuf fields get rejected when developing new capabilities related to protobuf transactions. The protobuf serialization format was designed with the assumption that unknown data known to encoders could safely be ignored by decoders. This assumption may have been fairly safe within the walled garden of Google’s centralized infrastructure. However, in distributed blockchain systems, this assumption is generally unsafe. If a newer client encodes a protobuf message with data intended for a newer server, it is not safe for an older server to simply ignore and discard instructions that it does not understand. These instructions could include critical information that the transaction signer is relying upon and just assuming that it is unimportant is not safe.
ADR 020 specifies some provisions for “non-critical” fields which can safely be ignored by older servers. In practice, I have not seen any valid usages of this. It is something in the design that maintainers should be aware of, but it may not be necessary or even 100% safe.
Non-deterministic Value Encoding
In addition to the non-determinism present in protobuf binary itself, some protobuf field data is encoded using a micro-format which itself may not be deterministic. Consider for instance integer or decimal encoding. Some decoders may allow for the presence of leading or trailing zeros without changing the logical meaning, ex.00100 vs 100 or 100.00 vs 100. So if a sign mode encodes numbers deterministically, but decoders accept multiple representations,
a user may sign over the value 100 while 0100 gets encoded. This would be possible with Amino JSON to the extent that the integer decoder accepts leading zeros. I believe the current Int implementation will reject this, however, it is
probably possible to encode a octal or hexadecimal representation in the transaction whereas the user signs over a decimal integer.
Signature Encoding
Signatures themselves are encoded using a micro-format specific to the signature algorithm being used and sometimes these micro-formats can allow for non-determinism (multiple valid bytes for the same signature). Most of the signature algorithms supported by the SDK should reject non-canonical bytes in their current implementation. However, theMultisignature protobuf type uses normal protobuf encoding and there is no check as to whether the
decoded bytes followed canonical ADR 027 rules or not. Therefore, multisig transactions can have malleability in
their signatures.
Any new or custom signature algorithms must make sure that they reject any non-canonical bytes, otherwise even
with SIGN_MODE_DIRECT there can be transaction hash malleability by re-encoding signatures with a non-canonical
representation.
Fields not covered by Amino JSON
Another area that needs to be addressed carefully is the discrepancy betweenAminoSignDoc(see aminojson.proto) used for SIGN_MODE_LEGACY_AMINO_JSON and the actual contents of TxBody and AuthInfo (see tx.proto).
If fields get added to TxBody or AuthInfo, they must either have a corresponding representing in AminoSignDoc or Amino JSON signatures must be rejected when those new fields are set. Making sure that this is done is a
highly manual process, and developers could easily make the mistake of updating TxBody or AuthInfo
without paying any attention to the implementation of GetSignBytes for Amino JSON. This is a critical
vulnerability in which unsigned content can now get into the transaction and signature verification will
pass.
Sign Mode Summary and Recommendations
The sign modes officially supported by the SDK areSIGN_MODE_DIRECT, SIGN_MODE_TEXTUAL, SIGN_MODE_DIRECT_AUX,
and SIGN_MODE_LEGACY_AMINO_JSON.
SIGN_MODE_LEGACY_AMINO_JSON is used commonly by wallets and is currently the only sign mode supported on Nano Ledger hardware devices
(although SIGN_MODE_TEXTUAL was designed to also support hardware devices).
SIGN_MODE_DIRECT is the simplest sign mode and its usage is also fairly common.
SIGN_MODE_DIRECT_AUX is a variant of SIGN_MODE_DIRECT that can be used by auxiliary signers in a multi-signer
transaction by those signers who are not paying gas.
SIGN_MODE_TEXTUAL was intended as a replacement for SIGN_MODE_LEGACY_AMINO_JSON, but as far as we know it
has not been adopted by any clients yet and thus is not in active use.
All known malleability concerns have been addressed in the current implementation of SIGN_MODE_DIRECT.
The only known malleability that could occur with a transaction signed with SIGN_MODE_DIRECT would
need to be in the signature bytes themselves.
Since signatures are not signed over, it is impossible for any sign mode to address this directly
and instead signature algorithms need to take care to reject any non-canonically encoded signature bytes
to prevent malleability.
For the known malleability of the Multisignature type, we should make sure that any valid signatures
were encoded following canonical ADR 027 rules when doing signature verification.
SIGN_MODE_DIRECT_AUX provides the same level of safety as SIGN_MODE_DIRECT because
- the raw encoded
TxBodybytes are signed over inSignDocDirectAux, and - a transaction using
SIGN_MODE_DIRECT_AUXstill requires the primary signer to sign the transaction withSIGN_MODE_DIRECT
SIGN_MODE_TEXTUAL also provides the same level of safety as SIGN_MODE_DIRECT because the hash of the raw encoded
TxBody and AuthInfo bytes are signed over.
Unfortunately, the vast majority of unaddressed malleability risks affect SIGN_MODE_LEGACY_AMINO_JSON and this
sign mode is still commonly used.
It is recommended that the following improvements be made to Amino JSON signing:
- hashes of
TxBodyandAuthInfoshould be added toAminoSignDocso that encoding-level malleablity is addressed - when constructing
AminoSignDoc, protoreflect API should be used to ensure that there no fields inTxBodyorAuthInfowhich do not have a mapping inAminoSignDochave been set - fields present in
TxBodyorAuthInfothat are not present inAminoSignDoc(such as extension options) should be added toAminoSignDocif possible
Testing
To test that transactions are resistant to malleability, we can develop a test suite to run against all sign modes that attempts to manipulate transaction bytes in the following ways:- changing protobuf encoding by
- reordering fields
- setting default values
- adding extra bits to varints, or
- setting new unknown fields
- modifying integer and decimal values encoded as strings with leading or trailing zeros
TxBody or AuthInfo
field not supported by Amino’s AminoSignDoc is set that signing fails.
In the general case of transaction decoding, we should have unit tests to ensure that
- any
TxRawbytes which do not follow ADR 027 canonical encoding cause decoding to fail, and - any top-level transaction elements including
TxBody,AuthInfo, public keys, and messages which have unknown fields set cause the transaction to be rejected (this ensures that ADR 020 unknown field filtering is properly applied)