变更记录
- 2021年12月06日:初始草案。
- 2022年02月07日:Ledger 团队已阅读草案并在概念上确认。
- 2022年05月16日:状态变更为 Accepted。
- 2022年08月11日:要求对交易原始字节进行签名。
- 2022年09月07日:添加自定义
Msg渲染器。 - 2022年09月18日:使用结构化格式替代文本行。
- 2022年11月23日:指定 CBOR 编码。
- 2022年12月01日:链接到单独 JSON 文件中的示例。
- 2022年12月06日:重新排序信封屏幕。
- 2022年12月14日:提及可逆性的例外情况。
- 2023年01月23日:将 Screen.Text 切换为 Title+Content。
- 2023年03月07日:将 SignDoc 从数组改为包含数组的结构体。
- 2023年03月20日:引入规范版本,初始值为 0。
状态
已接受。实现已开始。小数值渲染器的细节仍需打磨。 规范版本:0。摘要
本 ADR 指定了 SIGN_MODE_TEXTUAL,这是一种新的基于字符串的签名模式,面向硬件设备签名场景。背景
基于 Protobuf 的 SIGN_MODE_DIRECT 在 ADR-020 中引入,旨在在大多数场景下替代 SIGN_MODE_LEGACY_AMINO_JSON,例如移动钱包和 CLI 密钥环。然而,Ledger 硬件钱包仍在使用 SIGN_MODE_LEGACY_AMINO_JSON 向用户展示签名字节。硬件钱包无法迁移到 SIGN_MODE_DIRECT,原因如下:- SIGN_MODE_DIRECT 基于二进制,因此不适合向终端用户显示。从技术上讲,硬件钱包可以直接向用户显示签名字节,但这会被视为盲签,并带来安全风险。
- 由于内存限制,硬件无法解码 protobuf 签名字节,因为需要将 Protobuf 定义嵌入硬件设备中。
决策
在 SIGN_MODE_TEXTUAL 中,交易会被渲染为文本表示形式, 然后发送到安全设备或子系统,供用户审阅并签名。 与SIGN_MODE_DIRECT 不同,传输的数据即使在处理能力和显示能力受限的设备上,
也可以被简单解码为可读文本。
文本表示由一系列屏幕组成。
每个屏幕都应尽可能完整地显示出来,即使是在 Ledger 这类小型设备上也是如此。
一个屏幕大致相当于一行简短文本。
较大的屏幕可以分成多段显示,
类似长文本行的换行方式,
因此这里不做严格限制,不过以 40 个字符为较好目标。
屏幕用于显示单个标量值的键值对
(或采用紧凑表示法的复合值,例如 Coins),
或者用于引出或结束一个更大的分组。
文本可以包含完整的 Unicode 码点范围,包括控制字符和 nul。
设备负责决定如何显示其无法原生渲染的字符。
相关指导参见 附录 2。
屏幕具有非负缩进级别,用于表示复合结构或嵌套结构。
缩进级别 0 表示顶层。
缩进通过某种设备特定机制来显示。
消息引用表示法是一个合适的模型,例如
在功能较强的显示设备上使用前导 > 字符或竖线。
某些屏幕会被标记为专家屏幕,
仅当查看者选择启用更多细节时才显示。
专家屏幕用于展示很少有用的信息,
或者仅为了保证签名完整性而必须存在的信息(见下文)。
可逆渲染
我们要求交易的渲染必须是可逆的: 必须存在一个解析函数,使得对任意交易, 当其被渲染为文本表示后, 再解析该表示时,得到的 proto 消息在 proto 相等意义下 与原始消息等价。 请注意,这个逆函数并不需要对整个文本数据域 都执行正确解析或正确报告错误。 只要求有效交易的值域在 渲染与解析的组合下是可逆的。 还要注意,逆函数的存在保证了 渲染后的文本包含原始交易的全部信息, 而不仅仅是哈希或某个子集。 对于那些过大而无法有意义展示的数据, 例如长度超过 32 字节的字节串,我们对可逆性作出例外。 此时可以选择性地使用强加密哈希来渲染它们。在这些情况下, 要找到一个渲染结果相同但交易内容不同的交易, 在计算上仍然是不可行的。然而,我们必须确保哈希计算 足够简单,能够被可靠地独立执行, 这样在无法直接验证原始字节串时,至少哈希本身仍具有合理的可验证性。链状态
渲染函数(以及解析函数)可以依赖当前链状态。 这对于读取参数很有用,例如代币展示元数据, 或者读取用户特定偏好,例如语言或地址别名。 请注意,如果签名生成时观察到的状态与 交易被打包进区块时的状态之间发生变化, 那么交付时的渲染结果可能会不同。若确实如此, 签名将无效,交易也会被拒绝。签名与安全性
出于安全考虑,交易签名应具备三个属性:- 给定交易、签名和链状态,必须能够验证签名与交易匹配, 以证明签名者必然知晓各自的私钥。
- 在相同链状态下,计算上应当无法找到一个与原交易有显著差异、但给定签名依然有效的交易。
- 用户应能够通过一个简单、安全且显示能力受限的设备,对被签名的数据作出知情同意。
SIGN_MODE_TEXTUAL 的正确性与安全性通过证明从渲染结果到交易 proto 的逆函数来保证。
这意味着不可能存在另一个不同的 protocol buffer 消息渲染成同样的文本。
交易哈希可塑性
当客户端软件构造交易时,“原始”交易(TxRaw)会被序列化为 proto,
并对生成的字节序列计算哈希值。
这就是 TxHash,各种服务会使用它来在交易生命周期中跟踪已提交的交易。
如果有人能够生成一个修改后的交易,使其拥有不同的 TxHash,
但签名仍然可以通过校验,就可能出现各种不当行为。
SIGN_MODE_TEXTUAL 通过在渲染中将 TxHash 作为专家屏幕包含进去,
防止了这种交易可塑性。
SignDoc
SIGN_MODE_TEXTUAL 的 SignDoc 由如下数据结构构成:
细节
在下面的示例中,屏幕将显示为文本行, 缩进使用前导> 表示,
专家屏幕使用前导 * 标记。
交易信封的编码
我们将“交易信封”定义为交易中不属于TxBody.Messages 字段的所有数据。交易信封包括手续费、签名者信息和备注,但不包括 Msg。// 表示注释,在 Ledger 设备上不会显示。
交易体的编码
交易体是Tx.TxBody.Messages 字段,它是一个 Any 数组,其中每个 Any 都封装一个 sdk.Msg。由于 sdk.Msg 被广泛使用,它的编码方式与附录 1 中描述的常规 Any 数组(Protobuf: repeated google.protobuf.Any)略有不同。
示例
给定如下 Protobuf 消息:sdk.Msg 的交易,我们会得到如下编码:
自定义 Msg 渲染器
应用开发者可以选择不遵循其自定义 Msg 的默认渲染值输出。在这种情况下,他们可以实现自己的自定义 Msg 渲染器。这类似于 EIP4430,其中智能合约开发者决定向最终用户显示的描述字符串。
这是通过将 cosmos.msg.textual.v1.expert_custom_renderer Protobuf 选项设置为非空字符串来完成的。该选项只能设置在表示交易消息对象的 Protobuf 消息上(即实现 sdk.Msg 接口的消息)。
Msg 上时,已注册的函数会将该 Msg 转换为一个包含一个或多个字符串的数组,这些字符串可以使用键值格式(见第 3 点)、专家字段前缀(见第 5 点)以及任意缩进(见第 6 点)。这些字符串既可以通过默认值渲染器从某个 Msg 字段渲染而来,也可以通过自定义逻辑由多个字段生成。
<unique algorithm identifier> 是应用开发者选定的字符串约定,用于标识这个自定义 Msg 渲染器。例如,该自定义算法的文档或规范可以引用这个标识符。该标识符可以带有版本后缀(例如 _v1),以适应未来的变更(这类变更会破坏共识)。我们还建议添加 Protobuf 注释,以人类可读的方式说明所使用的自定义逻辑。
此外,渲染器必须提供 2 个函数:一个用于从 Protobuf 格式化为字符串,另一个用于将字符串解析为 Protobuf。这 2 个函数由应用开发者提供。为了满足第 1 点,解析函数必须是格式化函数的逆函数。SDK 不会在运行时检查这一性质。不过,我们强烈建议应用开发者在其应用仓库中加入完整的测试集来验证这种可逆性,以避免引入安全漏洞。
要求对 TxBody 和 AuthInfo 的原始字节进行签名
回顾一下,上链后进行默克尔化的交易字节是 TxRaw 的 Protobuf 二进制序列化结果,其中包含 body_bytes 和 auth_info_bytes。此外,交易哈希被定义为 TxRaw 字节的 SHA256 哈希。我们要求用户在 SIGN_MODE_TEXTUAL 中对这些字节进行签名,更具体地说,是对以下字符串进行签名:
++表示拼接,HEX是字节的十六进制表示形式,全部使用大写字母,不带0x前缀,len()编码为 Big-Endian uint64。
body 和 auth_info 的值不可被可塑修改,但如果只有第 1 点,交易哈希仍然可能具有可塑性,因为 SIGN_MODE_TEXTUAL 的字符串并不遵循 body_bytes 和 auth_info_bytes 中定义的字节顺序。如果没有这个哈希,恶意验证者或交易所可能会拦截一笔交易,在用户使用 SIGN_MODE_TEXTUAL 完成签名之后,通过调整 body_bytes 或 auth_info_bytes 内部的字节顺序来修改其交易哈希,然后再将其提交给 Tendermint。
通过将该哈希纳入 SIGN_MODE_TEXTUAL 的签名载荷中,我们可以保持与 SIGN_MODE_DIRECT 相同级别的保证。
这些字节只会在专家模式中显示,因此前面带有 *。
当前规范的更新
当前规范并非一成不变,未来预计还会继续迭代。我们将对此规范的更新分为两类:- 需要修改硬件设备内嵌应用的更新。
- 仅修改信封和数值渲染器的更新。
Screen 结构体或其对应 CBOR 编码的变更。这类更新需要修改硬件签名器应用,才能解码和解析新类型。同时还必须保证向后兼容,以便新的硬件应用能够与现有版本的 SDK 配合使用。这类更新需要多方协调:SDK 开发者、硬件应用开发者(当前为 Zondax)以及客户端开发者(例如 CosmJS)。此外,可能还需要重新提交硬件设备应用,而这取决于供应商,可能需要一些时间。因此,我们建议尽可能避免这类更新。
第 2 类更新包括对任意数值渲染器或交易信封的修改。例如,可以调整信封中字段的顺序,或者修改时间戳的格式化方式。由于 SIGN_MODE_TEXTUAL 向硬件设备发送的是 Screen,因此这类变更不需要更新硬件钱包应用。不过,它们会破坏状态机,必须据此进行文档说明。这类更新需要 SDK 开发者与客户端开发者(例如 CosmJS)协调,以便双方在时间上尽量接近地发布这些更新。
我们定义了一个规范版本,它是一个整数,每次发生任意一类更新时都必须递增。该规范版本将由 SDK 的实现暴露出来,并可传达给客户端。例如,SDK v0.50 可能使用规范版本 1,而 SDK v0.51 可能使用 2;借助这种版本控制,客户端可以根据目标 SDK 版本知道如何构造 SIGN_MODE_TEXTUAL 交易。
当前规范版本定义在本文顶部的“Status”部分。其初始值设为 0,以便在定义未来版本时保留灵活性,因为这样可以在 SignDoc Go 结构体或 Protobuf 中以向后兼容的方式新增字段。
硬件设备执行的附加格式化
请参见 附录 2。示例
以下示例存储在一个 JSON 文件中,包含以下字段:proto:交易的 ProtoJSON 表示,screens:渲染为 SIGN_MODE_TEXTUAL 屏幕的交易,cbor:交易的签名字节,即这些屏幕的 CBOR 编码。
影响
向后兼容性
SIGN_MODE_TEXTUAL 是纯增量式的,不会破坏与其他签名模式的任何向后兼容性。正面影响
- 在硬件设备上提供对人类友好的签名方式。
- 一旦 SIGN_MODE_TEXTUAL 发布,SIGN_MODE_LEGACY_AMINO_JSON 就可以被弃用并移除。从更长期来看,一旦整个生态系统完成迁移,Amino 也可以被彻底移除。
负面影响
- 某些字段仍然以非人类可读的方式编码,例如十六进制形式的公钥。
- 需要发布新的 ledger 应用,相关情况仍不明确。
中性影响
- 如果交易较为复杂,字符串数组可能会任意增长,一些用户可能会跳过部分屏幕并进行盲签。
进一步讨论
- 数值渲染器的一些细节仍需完善,参见 附录 1。
- ledger 应用是否能够同时支持 SIGN_MODE_LEGACY_AMINO_JSON 和 SIGN_MODE_TEXTUAL?
- 一个开放问题:我们是否应该增加一个 Protobuf 字段选项,以便应用开发者覆盖某些 Protobuf 字段和消息的文本表示?这将类似于以太坊的 EIP4430,由合约开发者决定文本表示方式。
- 国际化。
参考资料
Changelog
- Dec 06, 2021: Initial Draft.
- Feb 07, 2022: Draft read and concept-ACKed by the Ledger team.
- May 16, 2022: Change status to Accepted.
- Aug 11, 2022: Require signing over tx raw bytes.
- Sep 07, 2022: Add custom
Msg-renderers. - Sep 18, 2022: Structured format instead of lines of text
- Nov 23, 2022: Specify CBOR encoding.
- Dec 01, 2022: Link to examples in separate JSON file.
- Dec 06, 2022: Re-ordering of envelope screens.
- Dec 14, 2022: Mention exceptions for invertability.
- Jan 23, 2023: Switch Screen.Text to Title+Content.
- Mar 07, 2023: Change SignDoc from array to struct containing array.
- Mar 20, 2023: Introduce a spec version initialized to 0.
Status
Accepted. Implementation started. Small value renderers details still need to be polished. Spec version: 0.Abstract
This ADR specifies SIGN_MODE_TEXTUAL, a new string-based sign mode that is targetted at signing with hardware devices.Context
Protobuf-based SIGN_MODE_DIRECT was introduced in ADR-020 and is intended to replace SIGN_MODE_LEGACY_AMINO_JSON in most situations, such as mobile wallets and CLI keyrings. However, the Ledger hardware wallet is still using SIGN_MODE_LEGACY_AMINO_JSON for displaying the sign bytes to the user. Hardware wallets cannot transition to SIGN_MODE_DIRECT as:- SIGN_MODE_DIRECT is binary-based and thus not suitable for display to end-users. Technically, hardware wallets could simply display the sign bytes to the user. But this would be considered as blind signing, and is a security concern.
- hardware cannot decode the protobuf sign bytes due to memory constraints, as the Protobuf definitions would need to be embedded on the hardware device.
Decision
In SIGN_MODE_TEXTUAL, a transaction is rendered into a textual representation, which is then sent to a secure device or subsystem for the user to review and sign. UnlikeSIGN_MODE_DIRECT, the transmitted data can be simply decoded into legible text
even on devices with limited processing and display.
The textual representation is a sequence of screens.
Each screen is meant to be displayed in its entirety (if possible) even on a small device like a Ledger.
A screen is roughly equivalent to a short line of text.
Large screens can be displayed in several pieces,
much as long lines of text are wrapped,
so no hard guidance is given, though 40 characters is a good target.
A screen is used to display a single key/value pair for scalar values
(or composite values with a compact notation, such as Coins)
or to introduce or conclude a larger grouping.
The text can contain the full range of Unicode code points, including control characters and nul.
The device is responsible for deciding how to display characters it cannot render natively.
See annex 2 for guidance.
Screens have a non-negative indentation level to signal composite or nested structures.
Indentation level zero is the top level.
Indentation is displayed via some device-specific mechanism.
Message quotation notation is an appropriate model, such as
leading > characters or vertical bars on more capable displays.
Some screens are marked as expert screens,
meant to be displayed only if the viewer chooses to opt in for the extra detail.
Expert screens are meant for information that is rarely useful,
or needs to be present only for signature integrity (see below).
Invertible Rendering
We require that the rendering of the transaction be invertible: there must be a parsing function such that for every transaction, when rendered to the textual representation, parsing that representation yeilds a proto message equivalent to the original under proto equality. Note that this inverse function does not need to perform correct parsing or error signaling for the whole domain of textual data. Merely that the range of valid transactions be invertible under the composition of rendering and parsing. Note that the existence of an inverse function ensures that the rendered text contains the full information of the original transaction, not a hash or subset. We make an exception for invertibility for data which are too large to meaningfully display, such as byte strings longer than 32 bytes. We may then selectively render them with a cryptographically-strong hash. In these cases, it is still computationally infeasible to find a different transaction which has the same rendering. However, we must ensure that the hash computation is simple enough to be reliably executed independently, so at least the hash is itself reasonably verifiable when the raw byte string is not.Chain State
The rendering function (and parsing function) may depend on the current chain state. This is useful for reading parameters, such as coin display metadata, or for reading user-specific preferences such as language or address aliases. Note that if the observed state changes between signature generation and the transaction’s inclusion in a block, the delivery-time rendering might differ. If so, the signature will be invalid and the transaction will be rejected.Signature and Security
For security, transaction signatures should have three properties:- Given the transaction, signatures, and chain state, it must be possible to validate that the signatures matches the transaction, to verify that the signers must have known their respective secret keys.
- It must be computationally infeasible to find a substantially different transaction for which the given signatures are valid, given the same chain state.
- The user should be able to give informed consent to the signed data via a simple, secure device with limited display capabilities.
SIGN_MODE_TEXTUAL is guaranteed by demonstrating an inverse function from the rendering to transaction protos.
This means that it is impossible for a different protocol buffer message to render to the same text.
Transaction Hash Malleability
When client software forms a transaction, the “raw” transaction (TxRaw) is serialized as a proto
and a hash of the resulting byte sequence is computed.
This is the TxHash, and is used by various services to track the submitted transaction through its lifecycle.
Various misbehavior is possible if one can generate a modified transaction with a different TxHash
but for which the signature still checks out.
SIGN_MODE_TEXTUAL prevents this transaction malleability by including the TxHash as an expert screen
in the rendering.
SignDoc
The SignDoc forSIGN_MODE_TEXTUAL is formed from a data structure like:
Details
In the examples that follow, screens will be shown as lines of text, indentation is indicated with a leading ’>’, and expert screens are marked with a leading*.
Encoding of the Transaction Envelope
We define “transaction envelope” as all data in a transaction that is not in theTxBody.Messages field. Transaction envelope includes fee, signer infos and memo, but don’t include Msgs. // denotes comments and are not shown on the Ledger device.
Encoding of the Transaction Body
Transaction Body is theTx.TxBody.Messages field, which is an array of Anys, where each Any packs a sdk.Msg. Since sdk.Msgs are widely used, they have a slightly different encoding than usual array of Anys (Protobuf: repeated google.protobuf.Any) described in Annex 1.
Example
Given the following Protobuf message:sdk.Msg, we get the following encoding:
Custom Msg Renderers
Application developers may choose to not follow default renderer value output for their own Msgs. In this case, they can implement their own custom Msg renderer. This is similar to EIP4430, where the smart contract developer chooses the description string to be shown to the end user.
This is done by setting the cosmos.msg.textual.v1.expert_custom_renderer Protobuf option to a non-empty string. This option CAN ONLY be set on a Protobuf message representing transaction message object (implementing sdk.Msg interface).
Msg, a registered function will transform the Msg into an array of one or more strings, which MAY use the key/value format (described in point #3) with the expert field prefix (described in point #5) and arbitrary indentation (point #6). These strings MAY be rendered from a Msg field using a default value renderer, or they may be generated from several fields using custom logic.
The <unique algorithm identifier> is a string convention chosen by the application developer and is used to identify the custom Msg renderer. For example, the documentation or specification of this custom algorithm can reference this identifier. This identifier CAN have a versioned suffix (e.g. _v1) to adapt for future changes (which would be consensus-breaking). We also recommend adding Protobuf comments to describe in human language the custom logic used.
Moreover, the renderer must provide 2 functions: one for formatting from Protobuf to string, and one for parsing string to Protobuf. These 2 functions are provided by the application developer. To satisfy point #1, the parse function MUST be the inverse of the formatting function. This property will not be checked by the SDK at runtime. However, we strongly recommend the application developer to include a comprehensive suite in their app repo to test invertibility, as to not introduce security bugs.
Require signing over the TxBody and AuthInfo raw bytes
Recall that the transaction bytes merklelized on chain are the Protobuf binary serialization of TxRaw, which contains the body_bytes and auth_info_bytes. Moreover, the transaction hash is defined as the SHA256 hash of the TxRaw bytes. We require that the user signs over these bytes in SIGN_MODE_TEXTUAL, more specifically over the following string:
++denotes concatenation,HEXis the hexadecimal representation of the bytes, all in capital letters, no0xprefix,- and
len()is encoded as a Big-Endian uint64.
body and auth_info values are not malleable, but the transaction hash still might be malleable with point #1 only, because the SIGN_MODE_TEXTUAL strings don’t follow the byte ordering defined in body_bytes and auth_info_bytes. Without this hash, a malicious validator or exchange could intercept a transaction, modify its transaction hash after the user signed it using SIGN_MODE_TEXTUAL (by tweaking the byte ordering inside body_bytes or auth_info_bytes), and then submit it to Tendermint.
By including this hash in the SIGN_MODE_TEXTUAL signing payload, we keep the same level of guarantees as SIGN_MODE_DIRECT.
These bytes are only shown in expert mode, hence the leading *.
Updates to the current specification
The current specification is not set in stone, and future iterations are to be expected. We distinguish two categories of updates to this specification:- Updates that require changes of the hardware device embedded application.
- Updates that only modify the envelope and the value renderers.
Screen struct or its corresponding CBOR encoding. This type of updates require a modification of the hardware signer application, to be able to decode and parse the new types. Backwards-compatibility must also be guaranteed, so that the new hardware application works with existing versions of the SDK. These updates require the coordination of multiple parties: SDK developers, hardware application developers (currently: Zondax), and client-side developers (e.g. CosmJS). Furthermore, a new submission of the hardware device application may be necessary, which, dependending on the vendor, can take some time. As such, we recommend to avoid this type of updates as much as possible.
Updates in the 2nd category include changes to any of the value renderers or to the transaction envelope. For example, the ordering of fields in the envelope can be swapped, or the timestamp formatting can be modified. Since SIGN_MODE_TEXTUAL sends Screens to the hardware device, this type of change do not need a hardware wallet application update. They are however state-machine-breaking, and must be documented as such. They require the coordination of SDK developers with client-side developers (e.g. CosmJS), so that the updates are released on both sides close to each other in time.
We define a spec version, which is an integer that must be incremented on each update of either category. This spec version will be exposed by the SDK’s implementation, and can be communicated to clients. For example, SDK v0.50 might use the spec version 1, and SDK v0.51 might use 2; thanks to this versioning, clients can know how to craft SIGN_MODE_TEXTUAL transactions based on the target SDK version.
The current spec version is defined in the “Status” section, on the top of this document. It is initialized to 0 to allow flexibility in choosing how to define future versions, as it would allow adding a field either in the SignDoc Go struct or in Protobuf in a backwards-compatible way.
Additional Formatting by the Hardware Device
See annex 2.Examples
- A minimal MsgSend: see transaction.
- A transaction with a bit of everything: see transaction.
proto: the representation of the transaction in ProtoJSON,screens: the transaction rendered into SIGN_MODE_TEXTUAL screens,cbor: the sign bytes of the transaction, which is the CBOR encoding of the screens.
Consequences
Backwards Compatibility
SIGN_MODE_TEXTUAL is purely additive, and doesn’t break any backwards compatibility with other sign modes.Positive
- Human-friendly way of signing in hardware devices.
- Once SIGN_MODE_TEXTUAL is shipped, SIGN_MODE_LEGACY_AMINO_JSON can be deprecated and removed. On the longer term, once the ecosystem has totally migrated, Amino can be totally removed.
Negative
- Some fields are still encoded in non-human-readable ways, such as public keys in hexadecimal.
- New ledger app needs to be released, still unclear
Neutral
- If the transaction is complex, the string array can be arbitrarily long, and some users might just skip some screens and blind sign.
Further Discussions
- Some details on value renderers need to be polished, see Annex 1.
- Are ledger apps able to support both SIGN_MODE_LEGACY_AMINO_JSON and SIGN_MODE_TEXTUAL at the same time?
- Open question: should we add a Protobuf field option to allow app developers to overwrite the textual representation of certain Protobuf fields and message? This would be similar to Ethereum’s EIP4430, where the contract developer decides on the textual representation.
- Internationalization.