变更记录

  • 2019-10-15:初始草案
  • 2020-05-25:移除相关性根惩罚
  • 2020-07-01:更新为使用 S 曲线函数而非线性函数

背景

在基于权益证明的链中,共识权力集中在少数验证者手中,会因审查风险增加、活性失效、分叉攻击等问题对网络造成损害。然而,尽管这种中心化会给网络带来负外部性,这种影响并不会被那些继续委托给已足够大的验证者的委托人直接感知。我们希望有一种方式,能够将中心化带来的负外部性成本转移给这些大型验证者及其委托人。

决策

设计

为了解决这个问题,我们将实现一种称为“按比例惩罚”的机制。目标是:验证者规模越大,其应承担的惩罚越重。最直接的初步方案,是让验证者的惩罚比例与其在共识投票权中的占比成正比。
slash_amount = k * power // power is the faulting validator's voting power and k is some on-chain constant
然而,这会激励持有大量质押的验证者将自己的投票权拆分到多个账户中(女巫攻击),这样一旦发生故障,每个账户都会按更低的比例被惩罚。解决办法是,不仅要考虑某个验证者自身的投票权占比,还要考虑在指定时间范围内所有被惩罚验证者的投票权占比总和。
slash_amount = k * (power_1 + power_2 + ... + power_n) // where power_i is the voting power of the ith validator faulting in the specified time frame and k is some on-chain constant
这样一来,如果有人将一个占比 10% 的验证者拆分为两个各占 5% 的验证者,而这两个验证者都发生故障,那么由于它们在同一时间范围内一起出错,最终仍会按合计 10% 的占比受到惩罚。 但在实际中,我们很可能并不希望“故障质押量”与“惩罚比例”之间是线性关系。尤其是,仅有 5% 的质押发生双签,实际上并不足以对安全性构成重大威胁;而当 30% 的质押出现故障时,由于已经非常接近 Tendermint 安全性受到威胁的临界点,就显然应该适用更高的惩罚系数。若采用线性关系,这两者之间只会有 6 倍差距,而它们对网络带来的风险差异实际上远大于此。我们提议使用 S 曲线(形式上即逻辑函数)来解决这个问题。S 曲线能够很好地满足这一目标:在数值较小时,惩罚因子可以保持在较低水平;而在接近某个风险开始显著上升的阈值点时,惩罚因子会快速增长。

参数化

这要求对逻辑函数进行参数化。其参数化方式已经非常成熟,共包含四个参数:
  1. 最小惩罚因子
  2. 最大惩罚因子
  3. S 曲线的拐点位置(本质上就是你希望 S 的中心落在哪里)
  4. S 曲线的增长速率(S 被拉长到什么程度)

非女巫验证者之间的相关性

可以注意到,这个模型并不会区分多个验证者究竟是由同一批运营者运行,还是由不同运营者运行。从某种意义上说,这实际上还是一个额外的优点。它会激励验证者尽量让自己的部署方案与其他验证者不同,以避免与其他验证者发生相关性故障,否则就要承担更高的惩罚。例如,运营者应避免使用相同的热门云托管平台,或使用相同的 Staking as a Service 提供商。这将使网络更加稳健,也更具去中心化特性。

恶意连带

恶意连带指的是攻击者故意让自己被惩罚,从而加重他人的惩罚。在这里,这种情况可能会成为一种担忧。不过按照本文描述的协议,攻击者自身也会和受害者一样受到同等影响,因此这种行为对恶意实施者并没有太大收益。

实现

在 slashing 模块中,我们将新增两个队列,用于追踪近期所有的惩罚事件。对于双签故障,我们将“近期惩罚”定义为过去 unbonding period 内发生的惩罚;对于活性故障,我们将“近期惩罚”定义为过去 jail period 内发生的惩罚。
type SlashEvent struct {
    Address                     sdk.ValAddress
    ValidatorVotingPercent      sdk.Dec
    SlashedSoFar                sdk.Dec
}
一旦这些惩罚事件超过各自对应的“近期惩罚周期”,就会从队列中被剪枝移除。 每当发生一次新的惩罚时,都会创建一个 SlashEvent 结构体,其中记录出错验证者的投票权占比,并将 SlashedSoFar 初始化为 0。由于近期惩罚事件会在 unbonding period 和 unjail period 到期前被剪枝,因此同一个验证者理论上不可能在同一时间于同一个队列中存在多个 SlashEvent。 随后,我们会遍历队列中的所有 SlashEvent,累加它们的 ValidatorVotingPercent,以此计算队列中所有验证者新的统一惩罚比例,计算方式采用上文引入的“根之和的平方”公式。 得到 NewSlashPercent 之后,我们会再次遍历队列中的所有 SlashEvent。如果某个 SlashEvent 满足 NewSlashPercent > SlashedSoFar,则调用 staking.Slash(slashEvent.Address, slashEvent.Power, Math.Min(Math.Max(minSlashPercent, NewSlashPercent - SlashedSoFar), maxSlashPercent)(这里传入的是验证者在任何惩罚发生之前的 power,以确保扣减的代币数量正确)。随后,将该 SlashEvent.SlashedSoFar 更新为 NewSlashPercent。

状态

提议中

后果

正面影响

  • 通过抑制向大型验证者委托,提升去中心化程度
  • 激励验证者之间降低相关性
  • 对攻击行为的惩罚比对意外故障更严厉
  • 惩罚比例的参数化更灵活

负面影响

  • 比当前实现具有更高的计算开销,并且需要在链上存储更多关于“近期惩罚事件”的数据。

Changelog

  • 2019-10-15: Initial draft
  • 2020-05-25: Removed correlation root slashing
  • 2020-07-01: Updated to include S-curve function instead of linear

Context

In Proof of Stake-based chains, centralization of consensus power amongst a small set of validators can cause harm to the network due to increased risk of censorship, liveness failure, fork attacks, etc. However, while this centralization causes a negative externality to the network, it is not directly felt by the delegators contributing towards delegating towards already large validators. We would like a way to pass on the negative externality cost of centralization onto those large validators and their delegators.

Decision

Design

To solve this problem, we will implement a procedure called Proportional Slashing. The desire is that the larger a validator is, the more they should be slashed. The first naive attempt is to make a validator’s slash percent proportional to their share of consensus voting power.
slash_amount = k * power // power is the faulting validator's voting power and k is some on-chain constant
However, this will incentivize validators with large amounts of stake to split up their voting power amongst accounts (sybil attack), so that if they fault, they all get slashed at a lower percent. The solution to this is to take into account not just a validator’s own voting percentage, but also the voting percentage of all the other validators who get slashed in a specified time frame.
slash_amount = k * (power_1 + power_2 + ... + power_n) // where power_i is the voting power of the ith validator faulting in the specified time frame and k is some on-chain constant
Now, if someone splits a validator of 10% into two validators of 5% each which both fault, then they both fault in the same time frame, they both will get slashed at the sum 10% amount. However in practice, we likely don’t want a linear relation between amount of stake at fault, and the percentage of stake to slash. In particular, solely 5% of stake double signing effectively did nothing to majorly threaten security, whereas 30% of stake being at fault clearly merits a large slashing factor, due to being very close to the point at which Tendermint security is threatened. A linear relation would require a factor of 6 gap between these two, whereas the difference in risk posed to the network is much larger. We propose using S-curves (formally logistic functions to solve this). S-Curves capture the desired criterion quite well. They allow the slashing factor to be minimal for small values, and then grow very rapidly near some threshold point where the risk posed becomes notable.

Parameterization

This requires parameterizing a logistic function. It is very well understood how to parameterize this. It has four parameters:
  1. A minimum slashing factor
  2. A maximum slashing factor
  3. The inflection point of the S-curve (essentially where do you want to center the S)
  4. The rate of growth of the S-curve (How elongated is the S)

Correlation across non-sybil validators

One will note, that this model doesn’t differentiate between multiple validators run by the same operators vs validators run by different operators. This can be seen as an additional benefit in fact. It incentivizes validators to differentiate their setups from other validators, to avoid having correlated faults with them or else they risk a higher slash. So for example, operators should avoid using the same popular cloud hosting platforms or using the same Staking as a Service providers. This will lead to a more resilient and decentralized network.

Griefing

Griefing, the act of intentionally getting oneself slashed in order to make another’s slash worse, could be a concern here. However, using the protocol described here, the attacker also gets equally impacted by the grief as the victim, so it would not provide much benefit to the griefer.

Implementation

In the slashing module, we will add two queues that will track all of the recent slash events. For double sign faults, we will define “recent slashes” as ones that have occurred within the last unbonding period. For liveness faults, we will define “recent slashes” as ones that have occurred withing the last jail period.
type SlashEvent struct {
    Address                     sdk.ValAddress
    ValidatorVotingPercent      sdk.Dec
    SlashedSoFar                sdk.Dec
}
These slash events will be pruned from the queue once they are older than their respective “recent slash period”. Whenever a new slash occurs, a SlashEvent struct is created with the faulting validator’s voting percent and a SlashedSoFar of 0. Because recent slash events are pruned before the unbonding period and unjail period expires, it should not be possible for the same validator to have multiple SlashEvents in the same Queue at the same time. We then will iterate over all the SlashEvents in the queue, adding their ValidatorVotingPercent to calculate the new percent to slash all the validators in the queue at, using the “Square of Sum of Roots” formula introduced above. Once we have the NewSlashPercent, we then iterate over all the SlashEvents in the queue once again, and if NewSlashPercent > SlashedSoFar for that SlashEvent, we call the staking.Slash(slashEvent.Address, slashEvent.Power, Math.Min(Math.Max(minSlashPercent, NewSlashPercent - SlashedSoFar), maxSlashPercent) (we pass in the power of the validator before any slashes occurred, so that we slash the right amount of tokens). We then set SlashEvent.SlashedSoFar amount to NewSlashPercent.

Status

Proposed

Consequences

Positive

  • Increases decentralization by disincentivizing delegating to large validators
  • Incentivizes Decorrelation of Validators
  • More severely punishes attacks than accidental faults
  • More flexibility in slashing rates parameterization

Negative

  • More computationally expensive than current implementation. Will require more data about “recent slashing events” to be stored on chain.