变更记录
- 2020-08-07:初始草案
- 2020-09-01:进一步澄清规则
状态
提议中摘要
在对消息进行签名时,需要一种完全确定性的结构序列化方式,并且该方式能够在多种语言和客户端之间一致工作。我们需要确保,无论使用哪种受支持的语言来序列化某个数据结构,得到的原始字节都保持一致。 Protobuf 序列化不是双射的(也就是说,对于给定的 protobuf 文档,实际上存在数量几乎不受限制的有效二进制表示)1。 本文档描述了一种适用于 protobuf 文档子集的确定性序列化方案,它覆盖了这一使用场景,同时也可以复用于其他场景。背景
在 Cosmos SDK 中进行签名验证时,签名方和验证方需要在不传输序列化结果的前提下,就 ADR-020 中定义的SignDoc 的同一序列化结果达成一致。
当前,对于区块签名,我们使用了一种变通方案:在客户端侧将 Tx
的所有字段转换为字节后,创建一个新的 TxRaw
实例(定义见 adr-020-protobuf-transaction-encoding)。这会在发送和签名交易时增加一个额外的手动步骤。
决策
以下编码方案将供其他 ADR 使用,尤其用于SignDoc 序列化。
规范
范围
本 ADR 定义了一个 protobuf3 序列化器。其输出是合法的 protobuf 序列化结果,因此任何 protobuf 解析器都可以对其进行解析。 由于定义确定性序列化的复杂性,版本 1 不支持 map。未来这可能会改变。实现必须将包含 map 的文档视为无效输入并拒绝处理。背景 - Protobuf3 编码
protobuf3 中的大多数数值类型都编码为 varints。 Varint 最多为 10 个字节,并且由于每个 varint 字节携带 7 位数据, varint 实际上是uint70(70 位无符号整数)的一种表示方式。在编码时,数值会从其基础类型转换为 uint70;在解码时,解析得到的 uint70 会再转换为相应的数值类型。
符合 protobuf3 的 varint 最大合法值为
FF FF FF FF FF FF FF FF FF 7F(即 2**70 -1)。如果字段类型是
{,u,s}int64,解码时 70 位中的最高 6 位会被丢弃,从而引入 6 位的可塑性。如果字段类型是 {,u,s}int32,解码时 70 位中的最高 38 位会被丢弃,从而引入 38 位的可塑性。
除其他非确定性来源外,本 ADR 消除了编码可塑性的可能性。
序列化规则
该序列化基于 protobuf3 encoding ,并增加以下约束:- 字段必须且只能按升序序列化一次
- 不得添加额外字段或任何额外数据
- 必须省略默认值
- 标量数值类型的
repeated字段必须使用 packed encoding - Varint 编码不得长于必要长度:
- 不得有尾随的零字节(按小端表示,即按大端表示时不得有前导零)。根据上面的规则 3,默认值
0必须省略,因此该规则在这种情况下不适用。 - Varint 的最大值必须为
FF FF FF FF FF FF FF FF FF 01。 换句话说,解码后,70 位无符号整数的最高 6 位必须为0。(10 字节 varint 由 10 组 7 位组成,即 70 位,其中只有最低的 70-6=64 位是有效的。) - 32 位数值在 varint 编码中的最大值必须为
FF FF FF FF 0F, 但有一个例外(如下)。换句话说,解码后,70 位无符号整数的最高 38 位必须为0。- 上述规则的唯一例外是负数
int32,它必须使用完整的 10 个字节进行编码,以完成符号扩展2。
- 上述规则的唯一例外是负数
- 布尔值在 varint 编码中的最大值必须为
01(也就是它只能是0或1)。根据上面的规则 3,默认值0必须省略,因此如果包含某个布尔值,它的值必须为1。
- 不得有尾随的零字节(按小端表示,即按大端表示时不得有前导零)。根据上面的规则 3,默认值
""、0)、null 或未定义,从而形成 3 种不同的文档。
省略被设置为默认值的字段是合法的,因为解析器必须为序列化中缺失的字段赋予默认值4。对于标量类型,省略默认值是规范要求5。对于 repeated
字段,不对其进行序列化是表达空列表的唯一方式。枚举必须将数值为 0 的元素作为第一个元素,而这就是默认值6。消息字段的默认值则是未设置7。
省略默认值还提供了一定程度的前向兼容性:只要新增加的字段未被使用(即保持其默认值),新版 protobuf schema 的使用者与旧版 schema 的使用者就会生成相同的序列化结果。
实现
主要有三种实现策略,按从最少定制开发到最多定制开发排序:-
使用默认遵循上述规则的 protobuf 序列化器。 例如,
gogoproto 已知在大多数情况下是兼容的,但在使用某些注解(如
nullable = false)时并非如此。另一种选择也可能是对现有序列化器进行相应配置。 -
在编码前规范化默认值。 如果你的序列化器遵循规则 1 和 2,并允许你在序列化时显式取消设置字段,你可以将默认值规范化为未设置状态。使用
protobuf.js 时可以这样做:
-
为所需类型手写序列化器。 如果上述方法都不适用,你可以自行编写序列化器。对于 SignDoc,在 Go 中基于现有 protobuf 工具,大致会是下面这样:
测试向量
给定 protobuf 定义Article.proto
影响
提供这样一种编码方式后,我们就可以对 Cosmos SDK 签名场景中所需的所有 protobuf 文档实现确定性序列化。正面影响
- 规则定义明确,可独立于参考实现进行验证
- 足够简单,能够保持交易签名实现门槛较低
- 它使我们能够继续在 SignDoc 中使用 0 和其他空值,从而避免为 0 sequence 设计变通方案。这并不意味着不应合并 链接 中的更改,只是其重要性已经没有那么高。
负面影响
- 在实现交易签名时,必须理解并实现上述编码规则。
- 规则 3 的存在为实现增加了一些复杂度。
- 某些数据结构可能需要编写自定义序列化代码。因此,这些代码的可移植性不强,每个实现序列化的客户端都需要额外工作,才能正确处理自定义数据结构。
中性影响
在 Cosmos SDK 中的使用
基于上述原因(见“负面影响”部分),我们更倾向于对共享数据结构保留变通方案。例如,前面提到的TxRaw 就是使用原始字节作为一种变通方式。这样它们就可以使用任何合法的 Protobuf 库,而无需实现符合本标准的自定义序列化器(以及由此带来的相关 bug 风险)。
参考资料
- 1 当一条消息被序列化时,其已知字段或未知字段应按何种顺序写出,并没有保证。序列化顺序属于实现细节,任何特定实现的细节在未来都可能发生变化。因此,protocol buffer 解析器必须能够按任意顺序解析字段。 摘自 Link
- 2 Link
- 3 请注意,对于标量消息字段,一旦消息被解析,就无法判断某个字段是被显式设置为了默认值(例如某个布尔值是否被设置为 false),还是根本没有被设置:在定义消息类型时应牢记这一点。例如,如果你不希望某种行为默认也会发生,就不要使用一个在设置为 false 时会开启该行为的布尔值。 摘自 Link
- 4 当一条消息被解析时,如果编码后的消息不包含某个特定的单值元素,则解析后对象中的对应字段会被设置为该字段的默认值。 摘自 Link
- 5 另请注意,如果某个标量消息字段被设置为其默认值,该值将不会在 wire 上被序列化。 摘自 Link
- 6 对于枚举类型,默认值是第一个定义的枚举值,并且该值必须为 0。 摘自 Link
- 7 对于消息字段,该字段不会被设置。其确切值取决于具体语言。 摘自 Link
- 编码规则以及部分推理内容取自 canonical-proto3 Aaron Craelius
Changelog
- 2020-08-07: Initial Draft
- 2020-09-01: Further clarify rules
Status
ProposedAbstract
Fully deterministic structure serialization, which works across many languages and clients, is needed when signing messages. We need to be sure that whenever we serialize a data structure, no matter in which supported language, the raw bytes will stay the same. Protobuf serialization is not bijective (i.e. there exist a practically unlimited number of valid binary representations for a given protobuf document)1. This document describes a deterministic serialization scheme for a subset of protobuf documents, that covers this use case but can be reused in other cases as well.Context
For signature verification in Cosmos SDK, the signer and verifier need to agree on the same serialization of aSignDoc as defined in
ADR-020 without transmitting the
serialization.
Currently, for block signatures we are using a workaround: we create a new TxRaw
instance (as defined in adr-020-protobuf-transaction-encoding)
by converting all Tx
fields to bytes on the client side. This adds an additional manual
step when sending and signing transactions.
Decision
The following encoding scheme is to be used by other ADRs, and in particular forSignDoc serialization.
Specification
Scope
This ADR defines a protobuf3 serializer. The output is a valid protobuf serialization, such that every protobuf parser can parse it. No maps are supported in version 1 due to the complexity of defining a deterministic serialization. This might change in future. Implementations must reject documents containing maps as invalid input.Background - Protobuf3 Encoding
Most numeric types in protobuf3 are encoded as varints. Varints are at most 10 bytes, and since each varint byte has 7 bits of data, varints are a representation ofuint70 (70-bit unsigned integer). When
encoding, numeric values are casted from their base type to uint70, and when
decoding, the parsed uint70 is casted to the appropriate numeric type.
The maximum valid value for a varint that complies with protobuf3 is
FF FF FF FF FF FF FF FF FF 7F (i.e. 2**70 -1). If the field type is
{,u,s}int64, the highest 6 bits of the 70 are dropped during decoding,
introducing 6 bits of malleability. If the field type is {,u,s}int32, the
highest 38 bits of the 70 are dropped during decoding, introducing 38 bits of
malleability.
Among other sources of non-determinism, this ADR eliminates the possibility of
encoding malleability.
Serialization rules
The serialization is based on the protobuf3 encoding with the following additions:- Fields must be serialized only once in ascending order
- Extra fields or any extra data must not be added
- Default values must be omitted
repeatedfields of scalar numeric types must use packed encoding- Varint encoding must not be longer than needed:
- No trailing zero bytes (in little endian, i.e. no leading zeroes in big
endian). Per rule 3 above, the default value of
0must be omitted, so this rule does not apply in such cases. - The maximum value for a varint must be
FF FF FF FF FF FF FF FF FF 01. In other words, when decoded, the highest 6 bits of the 70-bit unsigned integer must be0. (10-byte varints are 10 groups of 7 bits, i.e. 70 bits, of which only the lowest 70-6=64 are useful.) - The maximum value for 32-bit values in varint encoding must be
FF FF FF FF 0Fwith one exception (below). In other words, when decoded, the highest 38 bits of the 70-bit unsigned integer must be0.- The one exception to the above is negative
int32, which must be encoded using the full 10 bytes for sign extension2.
- The one exception to the above is negative
- The maximum value for Boolean values in varint encoding must be
01(i.e. it must be0or1). Per rule 3 above, the default value of0must be omitted, so if a Boolean is included it must have a value of1.
- No trailing zero bytes (in little endian, i.e. no leading zeroes in big
endian). Per rule 3 above, the default value of
"", 0), null or undefined, leading to 3
different documents.
Omitting fields set to default values is valid because the parser must assign
the default value to fields missing in the serialization4. For scalar
types, omitting defaults is required by the spec5. For repeated
fields, not serializing them is the only way to express empty lists. Enums must
have a first element of numeric value 0, which is the default6. And
message fields default to unset7.
Omitting defaults allows for some amount of forward compatibility: users of
newer versions of a protobuf schema produce the same serialization as users of
older versions as long as newly added fields are not used (i.e. set to their
default value).
Implementation
There are three main implementation strategies, ordered from the least to the most custom development:-
Use a protobuf serializer that follows the above rules by default. E.g.
gogoproto is known to
be compliant by in most cases, but not when certain annotations such as
nullable = falseare used. It might also be an option to configure an existing serializer accordingly. -
Normalize default values before encoding them. If your serializer follows
rule 1. and 2. and allows you to explicitly unset fields for serialization,
you can normalize default values to unset. This can be done when working with
protobuf.js:
-
Use a hand-written serializer for the types you need. If none of the above
ways works for you, you can write a serializer yourself. For SignDoc this
would look something like this in Go, building on existing protobuf utilities:
Test vectors
Given the protobuf definitionArticle.proto
Consequences
Having such an encoding available allows us to get deterministic serialization for all protobuf documents we need in the context of Cosmos SDK signing.Positive
- Well defined rules that can be verified independent of a reference implementation
- Simple enough to keep the barrier to implement transaction signing low
- It allows us to continue to use 0 and other empty values in SignDoc, avoiding the need to work around 0 sequences. This does not imply the change from Link should not be merged, but not too important anymore.
Negative
- When implementing transaction signing, the encoding rules above must be understood and implemented.
- The need for rule number 3. adds some complexity to implementations.
- Some data structures may require custom code for serialization. Thus the code is not very portable - it will require additional work for each client implementing serialization to properly handle custom data structures.
Neutral
Usage in Cosmos SDK
For the reasons mentioned above (“Negative” section) we prefer to keep workarounds for shared data structure. Example: the aforementionedTxRaw is using raw bytes
as a workaround. This allows them to use any valid Protobuf library without
the need of implementing a custom serializer that adheres to this standard (and related risks of bugs).
References
- 1 When a message is serialized, there is no guaranteed order for how its known or unknown fields should be written. Serialization order is an implementation detail and the details of any particular implementation may change in the future. Therefore, protocol buffer parsers must be able to parse fields in any order. from Link
- 2 Link
- 3 Note that for scalar message fields, once a message is parsed there’s no way of telling whether a field was explicitly set to the default value (for example whether a boolean was set to false) or just not set at all: you should bear this in mind when defining your message types. For example, don’t have a boolean that switches on some behavior when set to false if you don’t want that behavior to also happen by default. from Link
- 4 When a message is parsed, if the encoded message does not contain a particular singular element, the corresponding field in the parsed object is set to the default value for that field. from Link
- 5 Also note that if a scalar message field is set to its default, the value will not be serialized on the wire. from Link
- 6 For enums, the default value is the first defined enum value, which must be 0. from Link
- 7 For message fields, the field is not set. Its exact value is language-dependent. from Link
- Encoding rules and parts of the reasoning taken from canonical-proto3 Aaron Craelius