变更记录
- 2020/08/18:初始版本
- 2021/01/15:分析与算法更新
状态
提议中摘要
本 ADR 为所有可寻址的 Cosmos SDK 账户定义了一种地址格式。其中包括:新的公钥算法、多重签名公钥以及模块账户。背景
问题 #3685 指出,当前公钥地址空间存在重叠。我们确认这会显著降低 Cosmos SDK 的安全性。问题
攻击者可以控制地址生成函数的输入。这会导致生日攻击,从而显著压缩安全空间。 为了解决这一点,我们需要将不同类型账户的输入彼此隔离: 某一种账户类型的安全性被攻破,不应影响其他账户类型的安全性。初始提案
最初的一个提案是扩展地址长度, 并为不同类型的地址添加前缀。 @ethanfrey 解释了一种最初在 Link 中使用的替代方案:我在构建 weave 时花了不少时间思考这个问题……另一个 cosmos Sdk。 基本思路是,我将一个条件定义为一种类型,并将其格式化为人类可读字符串,再附加一些二进制数据。这个条件会被哈希为一个 Address(同样是 20 字节)。使用这个前缀后,就不可能针对带有不同条件的给定地址找到原像(例如 ed25519 与 secp256k1)。 这里有更深入的说明 Link 代码在这里,主要看顶部处理条件的部分。Link并解释了为什么这种方法应当具备足够的抗碰撞能力:
是的,据我所知,只要原像是唯一且不可塑的,20 字节应当具备抗碰撞性。2^160 的空间中,大约在 2^80 个元素附近才会开始较大概率出现碰撞(生日悖论)。如果你想为数据库中某个已有元素找到碰撞,复杂度仍然是 2^160。只有当所有这些元素都写入状态时,才会出现 2^80 的情况。 你举的一个好例子是:某段公钥字节同时是 codec 支持的两种算法下的有效公钥。这意味着如果其中任意一种被攻破,即使账户原本由更安全的变体保护,你仍然可以攻破这些账户。只有在原像中没有携带可区分的类型信息(也就是哈希为地址之前)时,才会出现这个问题。 如果 20 字节空间在安全性上确实存在问题,我希望听到相关论证,因为我也愿意在 weave 中增加地址长度。我只是参考了 cosmos、ethereum 和 bitcoin 都使用 20 字节,觉得应该足够了。再加上上面的论证,让我倾向于认为它是安全的。不过我还没有做更深入的分析。这引出了第一个提案(后来我们证明它还不够好): 我们将密钥类型与公钥拼接,进行哈希,并取该哈希的前 20 个字节,可总结为
sha256(keyTypePrefix || keybytes)[:20]。
评审与讨论
在 #5694 中,我们讨论了多种解决方案。 我们一致认为,20 字节不具备长期适应性,而扩展地址长度是支持不同类型地址、各种签名类型等场景的唯一方式。 这也否定了最初的提案。 在该 issue 中,我们讨论了多种修改方案:- 哈希函数的选择。
- 将前缀移到哈希函数之外:
keyTypePrefix + sha256(keybytes)[:20][post-hash-prefix-proposal]。 - 使用双重哈希:
sha256(keyTypePrefix + sha256(keybytes)[:20])。 - 将
keybytes哈希切片从 20 字节增加到 32 或 40 字节。我们的结论是,由优良哈希函数生成的 32 字节在未来仍然是安全的。
需求
- 支持当前正在使用的工具,我们不希望破坏现有生态,也不希望引入过长的适配周期。参考:Link
- 尽量保持地址长度较小,因为地址在状态中被广泛使用,既可能是键的一部分,也可能是对象值的一部分。
范围
本 ADR 只定义地址字节的生成过程。对于终端用户与地址的交互(例如通过 API、CLI 等),我们仍然使用 bech32 将这些地址格式化为字符串。本 ADR 不会改变这一点。 使用 Bech32 进行字符串编码,可以让我们支持校验和错误码,并处理用户输入错误。决策
我们定义以下账户类型,并为其定义地址函数:- 简单账户:由常规公钥表示(例如:secp256k1、sr25519)
- 朴素多签:由其他可寻址对象组成的账户(例如:朴素多签)
- 带有原生地址键的组合账户(例如:bls、group 模块账户)
- 模块账户:基本上指任何不能签名交易、且由模块内部管理的账户
旧版公钥地址保持不变
当前(2021 年 1 月),Cosmos SDK 唯一官方支持的用户账户是secp256k1 基础账户和旧版 amino 多签。
它们已被现有的 Cosmos SDK zone 使用。它们使用以下地址格式:
- secp256k1:
ripemd160(sha256(pk_bytes))[:20] - 旧版 amino 多签:
sha256(aminoCdc.Marshal(pk))[:20]
哈希函数选择
与 Cosmos SDK 的其他部分一致,我们将使用sha256。
基础地址
我们首先定义一个用于生成地址的基础算法,称为Hash。需要特别说明的是,它用于由单个密钥对表示的账户。对于每一种公钥模式,我们都必须有一个关联的 typ 字符串,下一节会解释。hash 是上一节定义的密码学哈希函数。
+ 表示字节拼接,不使用任何分隔符。
该算法是与专业密码学家咨询后的结果。
动机在于:该算法可以让地址保持相对较小(typ 的长度不会影响最终地址的长度),
并且比 [post-hash-prefix-proposal] 更安全(后者使用公钥哈希的前 20 个字节,会显著缩小地址空间)。
此外,这位密码学家还说明,将 typ 放入哈希中,是为了防御 switch table attack。
address.Hash 是一个底层函数,用于为新密钥类型生成基础地址。例如:
- BLS:
address.Hash("bls", pubkey)
组合地址
对于简单的组合账户(例如一种新的朴素多签),我们对address.Hash 做了泛化。地址通过递归地为子账户创建地址、对这些地址排序,再将其组合为一个单一地址来构造。这样可以确保密钥的顺序不会影响最终地址。
typ 参数应当是一个模式描述符,包含所有重要属性,并且具有确定性的序列化方式(例如:utf8 字符串)。
LengthPrefix 是一个在地址前附加 1 个字节的函数。该字节的值是附加前地址位长度的长度。地址长度最多不能超过 255 位。
我们使用 LengthPrefix 来消除冲突。它保证:对于两组地址列表 as = {a1, a2, ..., an} 和 bs = {b1, b2, ..., bm},只要每个 bi 和 ai 的长度都不超过 255,则只有当 as = bs 时,concatenate(map(as, (a) => LengthPrefix(a))) = map(bs, (b) => LengthPrefix(b))。
实现提示:账户实现应当缓存地址。
多签地址
对于新的多签公钥,我们定义typ 参数时不依赖任何编码方案(amino 或 protobuf)。这样可以避免编码方案非确定性带来的问题。
示例:
派生地址
我们必须能够从一个地址以密码学方式派生出另一个地址。派生过程必须保证哈希属性,因此我们使用前面定义的Hash 函数:
模块账户地址
模块账户将具有"module" 类型。模块账户可以有子账户。子模块账户将基于模块名和派生键序列来创建。通常,第一个派生键应当表示派生账户的某个类别。派生过程有明确的顺序:模块名、子模块键、子子模块键…… 一个模块账户的创建示例如下:
address.Module 函数使用 address.Hash,其中类型参数为 "module",并将模块名的字节表示与子模块键拼接。最后两个组成部分必须能被唯一分隔,以避免潜在冲突(例如:modulename=“ab” 且 submodulekey=“bc”,其派生键将与 modulename=“a” 且 submodulekey=“bbc” 相同)。
我们使用空字节('\x00')来分隔模块名和子模块键。这样可行,因为空字节不是合法模块名的一部分。最后,子子模块账户通过递归应用 Derive 函数来创建。
我们也可以在第一步使用 Derive 函数(而不是将模块名、零字节和子模块键拼接)。我们决定使用拼接方式,以避免多一层派生并加快计算速度。
为了向后兼容现有的 authtypes.NewModuleAddress,我们在 Module 函数中添加了一个特殊分支:当未提供派生键时,回退到“legacy”实现。
模式类型
Hash 函数中使用的 typ 参数对于每种账户类型都应当唯一。
由于所有 Cosmos SDK 账户类型都会序列化到状态中,我们建议使用 protobuf 消息名字符串。
例如:所有公钥类型都有唯一的 protobuf 消息类型,类似于:
cosmos.crypto.sr25519.PubKey。
这些名称以标准化方式直接从 .proto 文件派生出来,并且也用于其他场景,例如 Any 的 type URL。我们可以通过
proto.MessageName(msg) 轻松获取该名称。
影响
向后兼容性
本 ADR 与已提交并由 Cosmos SDK 仓库直接支持的内容兼容。正面影响
- 为新的公钥、复杂账户和模块生成地址的简单算法
- 该算法推广了原生组合键
- 提高了地址的安全性和抗冲突能力
- 该方法可扩展到未来的使用场景,只要使用的其他地址类型不与此处规定的地址长度(20 或 32 字节)冲突即可。
- 支持新的账户类型。
负面影响
- 地址本身不传达密钥类型,而带前缀的方法本可以做到这一点
- 地址长度增加了 60%,会占用更多存储空间
- 需要重构 KVStore 的 store key,以支持变长地址
中性影响
- 使用 protobuf 消息名作为密钥类型前缀
进一步讨论
有些账户可能有固定名称,或者可能以其他方式构造(例如:模块)。我们曾讨论过一种带预定义名称的账户(例如:me.regen)的想法,这种账户可供机构使用。
不展开细节地说,只要这些地址的长度不同于这里描述的基于哈希的地址,它们就是兼容的。
更具体地说,任何特殊账户地址的长度都不能等于 20 或 32 字节。
附录:咨询会议
2020 年 12 月底,我们与 Alan Szepieniec 进行了一次会议,就上述方案进行咨询。 Alan 的总体观察:- 我们不需要第二原像抗性
- 为了抗冲突,我们需要 32 字节的地址空间
- 当攻击者能够控制某个带地址对象的输入时,就会面临生日攻击问题
- 对于哈希而言,智能合约存在一个问题
- 可以利用 sha2 挖矿来破坏地址原像
- 任何能攻破 blake3 的攻击也会攻破 blake2
- Alan 对当前 blake 哈希算法的安全性分析相当有信心。它曾是决赛入围方案,作者在安全分析领域也很知名。
- Alan 建议对前缀进行哈希:
address(pub_key) = hash(hash(key_type) + pub_key)[:32],主要优点是:- 我们可以自由使用任意长度的前缀名称
- 仍然不会有冲突风险
- switch tables
- 关于 penalization 的讨论 -> 关于在哈希后添加前缀
- Aaron 询问了哈希后前缀(
address(pub_key) = key_type + hash(pub_key))及其差异。Alan 指出,这种方法的地址空间更长,强度也更高。
- 使用相同算法合并树状地址是可行的
- 我们需要为模块地址设置一个原像前缀,以将其保持在 32 字节空间内:
hash(hash('module') + module_key) - Aaron 的观察:我们本来就需要处理变长问题(以避免破坏 secp256k1 密钥)。
- Posseidon / Rescue
- 问题:风险要大得多,因为我们对算术结构的密码分析技术和历史了解不多。这仍然是一个新的方向,也是活跃研究领域。
- Alan 的建议:Falcon,速度 / 大小比非常好。
- Aaron:我们应该考虑这个吗? Alan:根据早期推测,这种技术将在 2050 年具备攻破 EC 密码学的能力。但这其中存在很大的不确定性。不过在递归 / 链接 / 模拟方面正在出现一些突破,可能会加快这一进展。
- 假设我们在两种不同用例中,对同一个密钥使用两种不同的地址算法,这样仍然安全吗?Alan:如果我们想隐藏公钥(这不是我们的用例),那安全性会更低,但也有对应的修复方法。
参考资料
Changelog
- 2020/08/18: Initial version
- 2021/01/15: Analysis and algorithm update
Status
ProposedAbstract
This ADR defines an address format for all addressable Cosmos SDK accounts. That includes: new public key algorithms, multisig public keys, and module accounts.Context
Issue #3685 identified that public key address spaces are currently overlapping. We confirmed that it significantly decreases security of Cosmos SDK.Problem
An attacker can control an input for an address generation function. This leads to a birthday attack, which significantly decreases the security space. To overcome this, we need to separate the inputs for different kind of account types: a security break of one account type shouldn’t impact the security of other account types.Initial proposals
One initial proposal was extending the address length and adding prefixes for different types of addresses. @ethanfrey explained an alternate approach originally used in Link:I spent quite a bit of time thinking about this issue while building weave… The other cosmos Sdk. Basically I define a condition to be a type and format as human readable string with some binary data appended. This condition is hashed into an Address (again at 20 bytes). The use of this prefix makes it impossible to find a preimage for a given address with a different condition (eg ed25519 vs secp256k1). This is explained in depth here Link And the code is here, look mainly at the top where we process conditions. LinkAnd explained how this approach should be sufficiently collision resistant:
Yeah, AFAIK, 20 bytes should be collision resistance when the preimages are unique and not malleable. A space of 2^160 would expect some collision to be likely around 2^80 elements (birthday paradox). And if you want to find a collision for some existing element in the database, it is still 2^160. 2^80 only is if all these elements are written to state. The good example you brought up was eg. a public key bytes being a valid public key on two algorithms supported by the codec. Meaning if either was broken, you would break accounts even if they were secured with the safer variant. This is only as the issue when no differentiating type info is present in the preimage (before hashing into an address). I would like to hear an argument if the 20 bytes space is an actual issue for security, as I would be happy to increase my address sizes in weave. I just figured cosmos and ethereum and bitcoin all use 20 bytes, it should be good enough. And the arguments above which made me feel it was secure. But I have not done a deeper analysis.This led to the first proposal (which we proved to be not good enough): we concatenate a key type with a public key, hash it and take the first 20 bytes of that hash, summarized as
sha256(keyTypePrefix || keybytes)[:20].
Review and Discussions
In #5694 we discussed various solutions. We agreed that 20 bytes it’s not future proof, and extending the address length is the only way to allow addresses of different types, various signature types, etc. This disqualifies the initial proposal. In the issue we discussed various modifications:- Choice of the hash function.
- Move the prefix out of the hash function:
keyTypePrefix + sha256(keybytes)[:20][post-hash-prefix-proposal]. - Use double hashing:
sha256(keyTypePrefix + sha256(keybytes)[:20]). - Increase to keybytes hash slice from 20 byte to 32 or 40 bytes. We concluded that 32 bytes, produced by a good hash functions is future secure.
Requirements
- Support currently used tools - we don’t want to break an ecosystem, or add a long adaptation period. Ref: Link
- Try to keep the address length small - addresses are widely used in state, both as part of a key and object value.
Scope
This ADR only defines a process for the generation of address bytes. For end-user interactions with addresses (through the API, or CLI, etc.), we still use bech32 to format these addresses as strings. This ADR doesn’t change that. Using Bech32 for string encoding gives us support for checksum error codes and handling of user typos.Decision
We define the following account types, for which we define the address function:- simple accounts: represented by a regular public key (ie: secp256k1, sr25519)
- naive multisig: accounts composed by other addressable objects (ie: naive multisig)
- composed accounts with a native address key (ie: bls, group module accounts)
- module accounts: basically any accounts which cannot sign transactions and which are managed internally by modules
Legacy Public Key Addresses Don’t Change
Currently (Jan 2021), the only officially supported Cosmos SDK user accounts aresecp256k1 basic accounts and legacy amino multisig.
They are used in existing Cosmos SDK zones. They use the following address formats:
- secp256k1:
ripemd160(sha256(pk_bytes))[:20] - legacy amino multisig:
sha256(aminoCdc.Marshal(pk))[:20]
Hash Function Choice
As in other parts of the Cosmos SDK, we will usesha256.
Basic Address
We start with defining a base algorithm for generating addresses which we will callHash. Notably, it’s used for accounts represented by a single key pair. For each public key schema we have to have an associated typ string, explained in the next section. hash is the cryptographic hash function defined in the previous section.
+ is bytes concatenation, which doesn’t use any separator.
This algorithm is the outcome of a consultation session with a professional cryptographer.
Motivation: this algorithm keeps the address relatively small (length of the typ doesn’t impact the length of the final address)
and it’s more secure than [post-hash-prefix-proposal] (which uses the first 20 bytes of a pubkey hash, significantly reducing the address space).
Moreover the cryptographer motivated the choice of adding typ in the hash to protect against a switch table attack.
address.Hash is a low level function to generate base addresses for new key types. Example:
- BLS:
address.Hash("bls", pubkey)
Composed Addresses
For simple composed accounts (like a new naive multisig) we generalize theaddress.Hash. The address is constructed by recursively creating addresses for the sub accounts, sorting the addresses and composing them into a single address. It ensures that the ordering of keys doesn’t impact the resulting address.
typ parameter should be a schema descriptor, containing all significant attributes with deterministic serialization (eg: utf8 string).
LengthPrefix is a function which prepends 1 byte to the address. The value of that byte is the length of the address bits before prepending. The address must be at most 255 bits long.
We are using LengthPrefix to eliminate conflicts - it assures, that for 2 lists of addresses: as = {a1, a2, ..., an} and bs = {b1, b2, ..., bm} such that every bi and ai is at most 255 long, concatenate(map(as, (a) => LengthPrefix(a))) = map(bs, (b) => LengthPrefix(b)) if as = bs.
Implementation Tip: account implementations should cache addresses.
Multisig Addresses
For a new multisig public keys, we define thetyp parameter not based on any encoding scheme (amino or protobuf). This avoids issues with non-determinism in the encoding scheme.
Example:
Derived Addresses
We must be able to cryptographically derive one address from another one. The derivation process must guarantee hash properties, hence we use the already definedHash function:
Module Account Addresses
A module account will have"module" type. Module accounts can have sub accounts. The submodule account will be created based on module name, and sequence of derivation keys. Typically, the first derivation key should be a class of the derived accounts. The derivation process has a defined order: module name, submodule key, subsubmodule key… An example module account is created using:
address.Module function is using address.Hash with "module" as the type argument, and byte representation of the module name concatenated with submodule key. The two last component must be uniquely separated to avoid potential clashes (example: modulename=“ab” & submodulekey=“bc” will have the same derivation key as modulename=“a” & submodulekey=“bbc”).
We use a null byte ('\x00') to separate module name from the submodule key. This works, because null byte is not a part of a valid module name. Finally, the sub-submodule accounts are created by applying the Derive function recursively.
We could use Derive function also in the first step (rather than concatenating module name with zero byte and the submodule key). We decided to do concatenation to avoid one level of derivation and speed up computation.
For backward compatibility with the existing authtypes.NewModuleAddress, we add a special case in Module function: when no derivation key is provided, we fallback to the “legacy” implementation.
Schema Types
Atyp parameter used in Hash function SHOULD be unique for each account type.
Since all Cosmos SDK account types are serialized in the state, we propose to use the protobuf message name string.
Example: all public key types have a unique protobuf message type similar to:
cosmos.crypto.sr25519.PubKey.
These names are derived directly from .proto files in a standardized way and used
in other places such as the type URL in Anys. We can easily obtain the name using
proto.MessageName(msg).
Consequences
Backwards Compatibility
This ADR is compatible with what was committed and directly supported in the Cosmos SDK repository.Positive
- a simple algorithm for generating addresses for new public keys, complex accounts and modules
- the algorithm generalizes native composed keys
- increased security and collision resistance of addresses
- the approach is extensible for future use-cases - one can use other address types, as long as they don’t conflict with the address length specified here (20 or 32 bytes).
- support new account types.
Negative
- addresses do not communicate key type, a prefixed approach would have done this
- addresses are 60% longer and will consume more storage space
- requires a refactor of KVStore store keys to handle variable length addresses
Neutral
- protobuf message names are used as key type prefixes
Further Discussions
Some accounts can have a fixed name or may be constructed in other way (eg: modules). We were discussing an idea of an account with a predefined name (eg:me.regen), which could be used by institutions.
Without going into details, these kinds of addresses are compatible with the hash based addresses described here as long as they don’t have the same length.
More specifically, any special account address must not have a length equal to 20 or 32 bytes.
Appendix: Consulting session
End of Dec 2020 we had a session with Alan Szepieniec to consult the approach presented above. Alan general observations:- we don’t need 2-preimage resistance
- we need 32bytes address space for collision resistance
- when an attacker can control an input for object with an address then we have a problem with birthday attack
- there is an issue with smart-contracts for hashing
- sha2 mining can be use to breaking address pre-image
- any attack breaking blake3 will break blake2
- Alan is pretty confident about the current security analysis of the blake hash algorithm. It was a finalist, and the author is well known in security analysis.
- Alan recommends to hash the prefix:
address(pub_key) = hash(hash(key_type) + pub_key)[:32], main benefits:- we are free to user arbitrary long prefix names
- we still don’t risk collisions
- switch tables
- discussion about penalization -> about adding prefix post hash
- Aaron asked about post hash prefixes (
address(pub_key) = key_type + hash(pub_key)) and differences. Alan noted that this approach has longer address space and it’s stronger.
- merging tree like addresses with same algorithm are fine
- we will need to set a pre-image prefix for module addresse to keept them in 32-byte space:
hash(hash('module') + module_key) - Aaron observation: we already need to deal with variable length (to not break secp256k1 keys).
- Posseidon / Rescue
- Problem: much bigger risk because we don’t know much techniques and history of crypto-analysis of arithmetic constructions. It’s still a new ground and area of active research.
- Alan suggestion: Falcon: speed / size ration - very good.
- Aaron - should we think about it? Alan: based on early extrapolation this thing will get able to break EC cryptography in 2050 . But that’s a lot of uncertainty. But there is magic happening with recurions / linking / simulation and that can speedup the progress.
- Let’s say we use same key and two different address algorithms for 2 different use cases. Is it still safe to use it? Alan: if we want to hide the public key (which is not our use case), then it’s less secure but there are fixes.