递归长度前缀(Recursive Length Prefix,RLP)是一种在以太坊执行客户端中被广泛使用的序列化格式。它的用途是对任意嵌套的二进制数据数组进行编码,也是以太坊中用于序列化对象的主要编码方式。RLP 只编码结构,而将字符串、整数、浮点数等特定原子数据类型的编码交由更高层协议处理。 在以太坊中,整数必须表示为无前导零的大端二进制形式,因此整数值 0 等价于空字节数组。RLP 编码函数接收一个项(item)作为输入,该项可以是取值位于 [0x00, 0x7f] 范围内的单个字节,也可以是长度为 0 到 55 字节的字符串。如果字符串长度超过 55 字节,则其 RLP 编码由以下部分组成:一个单字节,其值为 0xb7(十进制 183)加上该字符串长度的长度(以二进制表示)所占的字节数;随后是字符串长度;最后是字符串本身。RLP 用于哈希校验,其中交易通过对交易数据的 RLP 哈希进行签名,而区块则通过其区块头的 RLP 哈希来标识。RLP 也用于网络传输中的数据编码,以及某些需要高效编码默克尔树数据结构的场景。以太坊执行层使用 RLP 作为序列化对象的主要编码方式,但在以太坊 2.0 的新共识层中,更新的简单序列化(Simple Serialize,SSZ)取代了 RLP。 Cosmos 的 Stargate 版本引入 protobuf 作为客户端序列化和状态序列化的主要编码格式。所有用于状态和客户端的 EVM 模块类型,例如交易消息、创世配置、查询服务等,都将实现为协议缓冲区消息。Cosmos SDK 也支持旧版的 Amino 编码。Protocol Buffers(protobuf)是一种与编程语言无关的二进制序列化格式,体积比 JSON 更小、速度也更快。它用于序列化消息等结构化数据,并被设计为兼具高效率与可扩展性。该编码格式由一种与语言无关的语言 Protocol Buffers Language(proto3)定义,编码后的消息可用于为多种编程语言生成代码。protobuf 的主要优势在于其高效率,这带来了更小的消息体积以及更快的序列化和反序列化速度。RLP 的解码过程如下:根据输入数据的第一个字节(即前缀)判断并解码数据类型、实际数据长度以及偏移量;再根据数据的类型和偏移量,对数据进行相应解码。

前置阅读

Cosmos SDK 编码

了解 Cosmos SDK 中的 protobuf 和 Amino 编码

以太坊 RLP

理解递归长度前缀编码

编码格式

Cosmos EVM 的主要编码方式Cosmos 的 Stargate 版本引入了 protobuf 作为客户端序列化和状态序列化的主要编码格式。
所有 EVM 模块类型(交易消息、创世配置、查询服务)都实现为协议缓冲区消息。
优势:
  • 与编程语言无关的二进制序列化
  • 比 JSON 更小的消息体积
  • 更快的序列化/反序列化
  • 具备模式校验的强类型支持

交易编码实现

x/vm 模块通过转换为 go-ethereum 的 Transaction 格式并使用 RLP 来处理 MsgEthereumTx 编码:
// TxEncoder overwrites sdk.TxEncoder to support MsgEthereumTx
func (g txConfig) TxEncoder() sdk.TxEncoder {
  return func(tx sdk.Tx) ([]byte, error) {
    msg, ok := tx.(*evmtypes.MsgEthereumTx)
    if ok {
      return msg.AsTransaction().MarshalBinary()
    }
    return g.TxConfig.TxEncoder()(tx)
  }
}

// TxDecoder overwrites sdk.TxDecoder to support MsgEthereumTx
func (g txConfig) TxDecoder() sdk.TxDecoder {
  return func(txBytes []byte) (sdk.Tx, error) {
    tx := &ethtypes.Transaction{}
    err := tx.UnmarshalBinary(txBytes)
    if err == nil {
      msg := &evmtypes.MsgEthereumTx{}
      msg.FromEthereumTx(tx)
      return msg, nil
    }
    return g.TxConfig.TxDecoder()(txBytes)
  }
}

The Recursive Length Prefix (RLP) is a serialization format used extensively in Ethereum’s execution clients. Its purpose is to encode arbitrarily nested arrays of binary data, and it is the main encoding method used to serialize objects in Ethereum. RLP only encodes structure and leaves encoding specific atomic data types, such as strings, integers, and floats, to higher-order protocols. In Ethereum, integers must be represented in big-endian binary form with no leading zeroes, making the integer value zero equivalent to the empty byte array. The RLP encoding function takes in an item, which is defined as a single byte whose value is in the [0x00, 0x7f] range or a string of 0-55 bytes long. If the string is more than 55 bytes long, the RLP encoding consists of a single byte with value 0xb7 (dec. 183) plus the length in bytes of the length of the string in binary form, followed by the length of the string, followed by the string. RLP is used for hash verification, where a transaction is signed by signing the RLP hash of the transaction data, and blocks are identified by the RLP hash of their header. RLP is also used for encoding data over the wire and for some cases where there should be support for efficient encoding of the merkle tree data structure. The Ethereum execution layer uses RLP as the primary encoding method to serialize objects, but the newer Simple Serialize (SSZ) replaces RLP as the encoding for the new consensus layer in Ethereum 2.0. The Cosmos Stargate release introduces protobuf as the main encoding format for both client and state serialization. All the EVM module types that are used for state and clients, such as transaction messages, genesis, query services, etc., will be implemented as protocol buffer messages. The Cosmos SDK also supports the legacy Amino encoding. Protocol Buffers (protobuf) is a language-agnostic binary serialization format that is smaller and faster than JSON. It is used to serialize structured data, such as messages, and is designed to be highly efficient and extensible. The encoding format is defined in a language-agnostic language called Protocol Buffers Language (proto3), and the encoded messages can be used to generate code for a variety of programming languages. The main advantage of protobuf is its efficiency, which results in smaller message sizes and faster serialization and deserialization times. The RLP decoding process is as follows: according to the first byte (i.e., prefix) of input data and decoding the data type, the length of the actual data and offset; according to the type and offset of data, decode the data correspondingly.

Prerequisite Readings

Cosmos SDK Encoding

Learn about protobuf and Amino encoding in Cosmos SDK

Ethereum RLP

Understand Recursive Length Prefix encoding

Encoding Formats

Primary encoding for Cosmos EVMThe Cosmos Stargate release introduces protobuf as the main encoding format for both client and state serialization.
All EVM module types (transaction messages, genesis, query services) are implemented as protocol buffer messages.
Advantages:
  • Language-agnostic binary serialization
  • Smaller message sizes than JSON
  • Faster serialization/deserialization
  • Strongly typed with schema validation

Transaction Encoding Implementation

The x/vm module handles MsgEthereumTx encoding by converting to go-ethereum’s Transaction format and using RLP:
// TxEncoder overwrites sdk.TxEncoder to support MsgEthereumTx
func (g txConfig) TxEncoder() sdk.TxEncoder {
  return func(tx sdk.Tx) ([]byte, error) {
    msg, ok := tx.(*evmtypes.MsgEthereumTx)
    if ok {
      return msg.AsTransaction().MarshalBinary()
    }
    return g.TxConfig.TxEncoder()(tx)
  }
}

// TxDecoder overwrites sdk.TxDecoder to support MsgEthereumTx
func (g txConfig) TxDecoder() sdk.TxDecoder {
  return func(txBytes []byte) (sdk.Tx, error) {
    tx := &ethtypes.Transaction{}
    err := tx.UnmarshalBinary(txBytes)
    if err == nil {
      msg := &evmtypes.MsgEthereumTx{}
      msg.FromEthereumTx(tx)
      return msg, nil
    }
    return g.TxConfig.TxDecoder()(txBytes)
  }
}