变更记录

  • 20-01-2020:初稿

状态

提案中

背景

遥测对于调试以及理解应用程序正在做什么、其运行表现如何至关重要。我们的目标是从模块以及 Cosmos SDK 的其他核心部分暴露指标。 此外,我们还应支持多种可配置的 sink,供运维人员自行选择。默认情况下,当启用遥测时,应用程序应跟踪并暴露存储在内存中的指标。运维人员还可以选择启用额外的 sink,目前我们只支持 Prometheus,因为它经过充分验证、易于设置、开源,并且拥有丰富的生态工具。 我们还必须以尽可能无缝的方式将指标集成到 Cosmos SDK 中,从而可以按需添加或移除指标,而不会带来太多阻力。为此,我们将使用 go-metrics 库。 最后,运维人员可以结合特定配置选项启用遥测。启用后,指标将通过 API 服务器的 /metrics?format={text|prometheus} 对外暴露。

决策

我们将在 app.toml 中新增一个配置块,用于定义遥测设置:
###############################################################################
###                         Telemetry Configuration                         ###
###############################################################################

[telemetry]

# Prefixed with keys to separate services
service-name = {{ .Telemetry.ServiceName }}

# Enabled enables the application telemetry functionality. When enabled,
# an in-memory sink is also enabled by default. Operators may also enabled
# other sinks such as Prometheus.
enabled = {{ .Telemetry.Enabled }}

# Enable prefixing gauge values with hostname
enable-hostname = {{ .Telemetry.EnableHostname }}

# Enable adding hostname to labels
enable-hostname-label = {{ .Telemetry.EnableHostnameLabel }}

# Enable adding service to labels
enable-service-label = {{ .Telemetry.EnableServiceLabel }}

# PrometheusRetentionTime, when positive, enables a Prometheus metrics sink.
prometheus-retention-time = {{ .Telemetry.PrometheusRetentionTime }}
该配置支持两种 sink:内存和 Prometheus。我们创建了一个 Metrics 类型,为运维人员完成所有初始化工作,从而使指标采集过程变得无缝。
// Metrics defines a wrapper around application telemetry functionality. It allows
// metrics to be gathered at any point in time. When creating a Metrics object,
// internally, a global metrics is registered with a set of sinks as configured
// by the operator. In addition to the sinks, when a process gets a SIGUSR1, a
// dump of formatted recent metrics will be sent to STDERR.
type Metrics struct {
    memSink           *metrics.InmemSink
  prometheusEnabled bool
}

// Gather collects all registered metrics and returns a GatherResponse where the
// metrics are encoded depending on the type. Metrics are either encoded via
// Prometheus or JSON if in-memory.
func (m *Metrics)

Gather(format string) (GatherResponse, error) {
    switch format {
    case FormatPrometheus:
    return m.gatherPrometheus()
    case FormatText:
    return m.gatherGeneric()
    case FormatDefault:
    return m.gatherGeneric()

default:
    return GatherResponse{
}, fmt.Errorf("unsupported metrics format: %s", format)
}
}
此外,Metrics 还允许我们在任意时刻收集当前这组指标。运维人员也可以选择发送 SIGUSR1 信号,将格式化后的指标转储并打印到 STDERR。 在应用程序的引导和构建阶段,如果 Telemetry.Enabled 为 true,API 服务器将创建一个 Metrics 对象引用的实例,并据此注册指标处理器。
func (s *Server)

Start(cfg config.Config)

error {
  // ...
    if cfg.Telemetry.Enabled {
    m, err := telemetry.New(cfg.Telemetry)
    if err != nil {
    return err
}

s.metrics = m
    s.registerMetrics()
}

  // ...
}

func (s *Server)

registerMetrics() {
    metricsHandler := func(w http.ResponseWriter, r *http.Request) {
    format := strings.TrimSpace(r.FormValue("format"))

gr, err := s.metrics.Gather(format)
    if err != nil {
    rest.WriteErrorResponse(w, http.StatusBadRequest, fmt.Sprintf("failed to gather metrics: %s", err))

return
}

w.Header().Set("Content-Type", gr.ContentType)
    _, _ = w.Write(gr.Metrics)
}

s.Router.HandleFunc("/metrics", metricsHandler).Methods("GET")
}
应用开发者可以跟踪计数器、gauge、摘要以及键值指标。模块无需额外工作即可利用性能分析指标。实现方式非常简单:
func (k BaseKeeper)

MintCoins(ctx sdk.Context, moduleName string, amt sdk.Coins)

error {
    defer metrics.MeasureSince(time.Now(), "MintCoins")
  // ...
}

影响

正面

  • 提升对应用程序性能和行为的可观测性

负面

中性

参考资料


Changelog

  • 20-01-2020: Initial Draft

Status

Proposed

Context

Telemetry is paramount into debugging and understanding what the application is doing and how it is performing. We aim to expose metrics from modules and other core parts of the Cosmos SDK. In addition, we should aim to support multiple configurable sinks that an operator may choose from. By default, when telemetry is enabled, the application should track and expose metrics that are stored in-memory. The operator may choose to enable additional sinks, where we support only Prometheus for now, as it’s battle-tested, simple to setup, open source, and is rich with ecosystem tooling. We must also aim to integrate metrics into the Cosmos SDK in the most seamless way possible such that metrics may be added or removed at will and without much friction. To do this, we will use the go-metrics library. Finally, operators may enable telemetry along with specific configuration options. If enabled, metrics will be exposed via /metrics?format={text|prometheus} via the API server.

Decision

We will add an additional configuration block to app.toml that defines telemetry settings:
###############################################################################
###                         Telemetry Configuration                         ###
###############################################################################

[telemetry]

# Prefixed with keys to separate services
service-name = {{ .Telemetry.ServiceName }}

# Enabled enables the application telemetry functionality. When enabled,
# an in-memory sink is also enabled by default. Operators may also enabled
# other sinks such as Prometheus.
enabled = {{ .Telemetry.Enabled }}

# Enable prefixing gauge values with hostname
enable-hostname = {{ .Telemetry.EnableHostname }}

# Enable adding hostname to labels
enable-hostname-label = {{ .Telemetry.EnableHostnameLabel }}

# Enable adding service to labels
enable-service-label = {{ .Telemetry.EnableServiceLabel }}

# PrometheusRetentionTime, when positive, enables a Prometheus metrics sink.
prometheus-retention-time = {{ .Telemetry.PrometheusRetentionTime }}
The given configuration allows for two sinks — in-memory and Prometheus. We create a Metrics type that performs all the bootstrapping for the operator, so capturing metrics becomes seamless.
// Metrics defines a wrapper around application telemetry functionality. It allows
// metrics to be gathered at any point in time. When creating a Metrics object,
// internally, a global metrics is registered with a set of sinks as configured
// by the operator. In addition to the sinks, when a process gets a SIGUSR1, a
// dump of formatted recent metrics will be sent to STDERR.
type Metrics struct {
    memSink           *metrics.InmemSink
  prometheusEnabled bool
}

// Gather collects all registered metrics and returns a GatherResponse where the
// metrics are encoded depending on the type. Metrics are either encoded via
// Prometheus or JSON if in-memory.
func (m *Metrics)

Gather(format string) (GatherResponse, error) {
    switch format {
    case FormatPrometheus:
    return m.gatherPrometheus()
    case FormatText:
    return m.gatherGeneric()
    case FormatDefault:
    return m.gatherGeneric()

default:
    return GatherResponse{
}, fmt.Errorf("unsupported metrics format: %s", format)
}
}
In addition, Metrics allows us to gather the current set of metrics at any given point in time. An operator may also choose to send a signal, SIGUSR1, to dump and print formatted metrics to STDERR. During an application’s bootstrapping and construction phase, if Telemetry.Enabled is true, the API server will create an instance of a reference to Metrics object and will register a metrics handler accordingly.
func (s *Server)

Start(cfg config.Config)

error {
  // ...
    if cfg.Telemetry.Enabled {
    m, err := telemetry.New(cfg.Telemetry)
    if err != nil {
    return err
}

s.metrics = m
    s.registerMetrics()
}

  // ...
}

func (s *Server)

registerMetrics() {
    metricsHandler := func(w http.ResponseWriter, r *http.Request) {
    format := strings.TrimSpace(r.FormValue("format"))

gr, err := s.metrics.Gather(format)
    if err != nil {
    rest.WriteErrorResponse(w, http.StatusBadRequest, fmt.Sprintf("failed to gather metrics: %s", err))

return
}

w.Header().Set("Content-Type", gr.ContentType)
    _, _ = w.Write(gr.Metrics)
}

s.Router.HandleFunc("/metrics", metricsHandler).Methods("GET")
}
Application developers may track counters, gauges, summaries, and key/value metrics. There is no additional lifting required by modules to leverage profiling metrics. To do so, it’s as simple as:
func (k BaseKeeper)

MintCoins(ctx sdk.Context, moduleName string, amt sdk.Coins)

error {
    defer metrics.MeasureSince(time.Now(), "MintCoins")
  // ...
}

Consequences

Positive

  • Exposure into the performance and behavior of an application

Negative

Neutral

References