本文档详细说明了 QA 流程。 其目的是供工程师在未来进行 CometBFT 测试时复现实验环境。 releases 中描述的 QA 流程(第一版)已应用于 v0.34.x 版本,以获得一组可作为基准基线的结果。 随后会将该基线与后续版本得到的结果进行比较。 在 releases 中基于测试网的测试用例里,我们重点关注其中两个: 200 节点测试 和 轮换节点测试。

软件依赖

运行测试所需的基础设施

结果提取要求

  • 使用 Prometheus DB 从节点收集指标
  • 用于处理查询的 Prometheus DB(可以与前一个不是同一台节点)
  • 测试网中某个全节点的 blockstore DB

200 节点测试网

运行测试

本节说明这些测试是如何执行的,以便复现。
  1. [如果你之前还没做过] 按照 testnet 仓库顶层 README.md 中的步骤 1-4 配置 Terraform 和 doctl。
  2. 将文件 testnets/testnet200.toml 复制为 testnet.toml(不要提交此变更)。
  3. 将 Makefile 中的变量 VERSION_TAG 设置为要测试的 git hash。
    • 如果你运行的是基线测试,即同构网络(所有节点运行相同版本), 那么请确保 makefile 变量 VERSION2_WEIGHT 设置为 0。
    • 如果你运行的是混合网络,请将变量 VERSION2_TAG 设置为你希望在网络中部署的另一个版本。 然后调整权重变量 VERSION_WEIGHT 和 VERSION2_WEIGHT, 以配置两种已设置版本各自运行的节点占比。
  4. 按照 README.md 中的步骤 5-10 配置并启动 200 节点测试网。
    • 警告:测试完成后务必立刻运行 make terraform-destroy(见步骤 9)。
  5. 作为基本检查,连接到 Prometheus 节点的 Web 界面(9090 端口), 并查看 cometbft_consensus_height 指标的图表。所有节点 的高度都应持续增长。
    • 你可以在 ansible/hosts 的 [prometheus] 部分找到 Prometheus 节点的 IP 地址。
    • 以下 URL 会显示 cometbft_consensus_height 和 cometbft_mempool_size 指标:
      http://<PROMETHEUS-NODE-IP>:9090/classic/graph?g0.range_input=1h&g0.expr=cometbft_consensus_height&g0.tab=0&g1.range_input=1h&g1.expr=cometbft_mempool_size&g1.tab=0
      
  6. 现在你需要启动会产生交易负载的 load runner。
    • 如果你不知道当前测试版本的饱和负载,就需要先探测出来。
      • 运行 make loadrunners-init。这会把 loader 脚本复制到 testnet-load-runner 节点,并安装负载工具。
      • 在 ansible/hosts 的 [loadrunners] 部分中找到 testnet-load-runner 节点的 IP 地址。
      • 通过 ssh 登录 testnet-load-runner。
        • 编辑负载运行器节点上的脚本 /root/200-node-loadscript.sh, 填入一个全节点的 IP 地址(例如 validator000)。 所有来自负载运行器节点的交易都会发送到这个节点。
        • 在负载运行器节点上运行 /root/200-node-loadscript.sh。
          • 该脚本运行大约需要 40 分钟,因此建议先启动 tmux, 以防 ssh 会话中断。
          • 它会循环执行持续 90 秒、负载各不相同的实验。
    • 如果你已经知道饱和负载,那么可以直接在略低于饱和点的负载下运行测试(多次),每次持续 90 秒:
      • 将 makefile 变量 LOAD_CONNECTIONS、LOAD_TX_RATE 设置为能产生目标交易负载的值。
      • 将 LOAD_TOTAL_TIME 设置为 90(秒)。
      • 运行 make runload 并等待其完成。你可能需要多运行几次,以便比较不同轮次的数据。
  7. 运行 make retrieve-data,将测试网中的所有相关数据收集到编排机器上。
    • 或者,你也可以分别运行 make retrieve-prometheus-data 和 make retrieve-blockstore。 最终结果是一样的。
    • make retrieve-blockstore 支持在 makefile 变量 RETRIEVE_TARGET_HOST 中使用以下取值:
      • any:(默认值)选择一个全节点,仅从该节点拉取 blockstore。
      • all:从所有全节点拉取 blockstore;这会非常慢并消耗大量带宽, 因此请谨慎使用。
      • 某个具体全节点的名称(例如 validator01):仅从该节点拉取 blockstore。
  8. 验证数据是否已无误收集:
    • 至少一个 CometBFT 验证者的 blockstore DB
    • 来自 Prometheus 节点的 Prometheus 数据库
    • 为了更稳妥起见,你可以对 prometheus.zip 文件以及(其中一个)blockstore.db.zip 文件运行 zip -T
  9. 运行 make terraform-destroy
    • 别忘了输入 yes!否则会出问题。

结果提取

这里描述的结果提取方法目前仍然高度依赖手工操作(且具有探索性质)。 CometBFT 团队应在每次迭代中持续改进,以提高自动化程度。

步骤

  1. 将 blockstore 解压到某个目录中。
  2. 为了识别饱和点:
    1. 提取所有实验的延迟报告。
      • 在包含 blockstore.db 文件夹的目录中运行以下命令。
      • 建议将 go run 命令中的 hash 调整为尽可能新的版本。
      • mkdir results
        go run github.com/cometbft/cometbft/test/loadtime/cmd/report@3003ef7 --database-type goleveldb --data-dir ./ > results/report.txt
        
    2. 文件 report.txt 包含一个无序的实验列表,实验之间使用不同的并发连接数和交易速率。 你需要按实验拆分数据。
      • 创建文件 report01.txt、report02.txt、report04.txt,并针对 report.txt 中的每个实验, 将其相关行复制到与连接数匹配的文件名中,例如:
        for cnum in 1 2 4; do echo "$cnum"; grep "Connections: $cnum" results/report.txt -B 2 -A 10 > results/report$cnum.txt;  done
        
      • 将 report01.txt 中的实验按 tx rate 升序排序。report02.txt 和 report04.txt 也同样处理。
      • 否则,也可以直接保留 report.txt 并跳到下一步。
    3. 通过并排显示 report01.txt、report02.txt、report04.txt 的内容来生成文件 report_tabbed.txt。
      • 这实际上会创建一个表格,其中行表示某个特定 tx rate,列表示某个特定 websocket 连接数。
      • 将各列文件合并为一个表格文件:
        • 先把所有列文件中的制表符替换为空格。例如, sed -i.bak 's/\t/ /g' results/report1.txt。
      • 再将新的列文件合并为一个: paste results/report1.txt results/report2.txt results/report4.txt | column -s $'\t' -t > report_tabbed.txt
  3. 为了生成“延迟 vs 吞吐量”图,请将数据提取为 CSV:
    •  go run github.com/cometbft/cometbft/test/loadtime/cmd/report@3003ef7 --database-type goleveldb --data-dir ./ --csv results/raw.csv
      
    • 按照 latency_throughput.py 脚本的说明进行操作。 该图有助于可视化饱和点。
    • 或者,按照 latency_plotter.py 脚本的说明进行操作。 该脚本会针对每个实验和配置生成一系列图表,可能有助于 可视化延迟与吞吐量的变化。

提取 Prometheus 指标

  1. 如果 prometheus server 作为服务运行(例如 systemd 单元),请先停止它。
  2. 解压从测试网取回的 prometheus 数据库,并将其移动到本地,替换掉 本地 prometheus 数据库。
  3. 启动 prometheus server,并确保启动时没有出现错误日志。
  4. 确定你希望在图表中绘制的时间窗口。
  5. 针对该时间窗口执行 prometheus_plotter.py 脚本。

轮换节点测试网

运行测试

本节说明这些测试是如何执行的,以便复现。
  1. [如果你之前还没做过] 按照 testnet 仓库顶层 README.md 中的步骤 1-4 配置 Terraform 和 doctl。
  2. 将文件 testnet_rotating.toml 复制为 testnet.toml(不要提交此变更)。
  3. 将变量 VERSION_TAG 设置为要测试的 git hash。
  4. 运行 make terraform-apply EPHEMERAL_SIZE=25。
    • 警告:测试完成后务必立刻运行 make terraform-destroy。
  5. 按照 README.md 中的步骤 6-10,配置并启动轮换节点测试网中“稳定”的那部分。
  6. 作为基本检查,连接到 Prometheus 节点的 Web 界面,并查看 tendermint_consensus_height 指标的图表。 所有节点的高度都应持续增长。
  7. 在另一个 shell 中:
    • 运行 make runload LOAD_CONNECTIONS=X LOAD_TX_RATE=Y LOAD_TOTAL_TIME=Z。
    • X 和 Y 应反映低于饱和点的负载(更多信息可参见 这一段)。
    • Z(单位:秒)应足够大,使其在整个测试期间持续运行,直到我们在步骤 9 手动停止它。 原则上,Z 的一个合适取值是 7200(2 小时)。
  8. 运行 make rotate,启动脚本来创建临时节点,并在它们追上后将其终止。
    • 警告:如果你从笔记本电脑上运行此命令,那么在整个实验期间, 笔记本都必须保持开机并联网。
    • http://<PROMETHEUS-NODE-IP>:9090/classic/graph?g0.range_input=100m&g0.expr=cometbft_consensus_height%7Bjob%3D~%22ephemeral.*%22%7D%20or%20cometbft_blocksync_latest_block_height%7Bjob%3D~%22ephemeral.*%22%7D&g0.tab=0&g1.range_input=100m&g1.expr=cometbft_mempool_size%7Bjob!~%22ephemeral.*%22%7D&g1.tab=0&g2.range_input=100m&g2.expr=cometbft_consensus_num_txs%7Bjob!~%22ephemeral.*%22%7D&g2.tab=0 是一个可用于监控该测试用例进展的 Prometheus URL 示例。
  9. 当链高度达到 3000 时,停止 make runload 脚本。
  10. 在高度达到 3000 之后,当 rotate 脚本完成两轮迭代(即所有临时节点都已完成两次追赶)时,停止 make rotate。
  11. 运行 make stop-network。
  12. 运行 make retrieve-data,将测试网中的所有相关数据收集到编排机器上。
  13. 验证数据是否已无误收集:
    • 至少一个 CometBFT 验证者的 blockstore DB
    • 来自 Prometheus 节点的 Prometheus 数据库
    • 为了更稳妥起见,你可以对 prometheus.zip 文件以及(其中一个)blockstore.db.zip 文件运行 zip -T
  14. 运行 make terraform-destroy
步骤 8 到 10 目前高度依赖手工操作,后续迭代中会继续改进。

结果提取

为了获得延迟图,请按照上文 200 节点实验的说明操作, 但 results.txt 文件只包含一个实验。 至于 Prometheus,可以采用与 200 节点实验相同的方法。

Vote Extensions 测试网

运行测试

本节说明这些测试是如何执行的,以便复现。
  1. [如果你之前还没做过] 按照 testnet 仓库顶层 README.md 中的步骤 1-4 配置 Terraform 和 doctl。
  2. 将文件 varyVESize.toml 复制为 testnet.toml(不要提交此变更)。
  3. 将 Makefile 中的变量 VERSION_TAG 设置为要测试的 git hash。
  4. 按照 README.md 中的步骤 5-10 配置并启动测试网。
    • 警告:测试完成后务必立刻运行 make terraform-destroy。
  5. 配置 load runner 以产生所需的交易负载。
    • 将 makefile 变量 ROTATE_CONNECTIONS、ROTATE_TX_RATE 设置为能产生目标交易负载的值。
    • 将 ROTATE_TOTAL_TIME 设置为 150(秒)。
    • 将 ITERATIONS 设置为每种配置需要运行的迭代次数。
  6. 执行 testnet 仓库中 README.md 文件的步骤 5-10。
  7. 针对每个期望的 vote_extension_size,重复以下步骤:
    1. 更新配置(如果你没有修改 vote_extension_size,可以跳过此步骤)。
      • 将 testnet.toml 中的 vote_extensions_size 更新为目标值。
      • make configgen
      • ANSIBLE_SSH_RETRIES=10 ansible-playbook ./ansible/re-init-testapp.yaml -u root -i ./ansible/hosts --limit=validators -e "testnet_dir=testnet" -f 20
      • make restart
    2. 运行测试。
      • make runload 每次调用时,它都会将测试重复执行 ITERATIONS 次。
    3. 收集数据。
      • make retrieve-data 将测试网中的所有相关数据收集到编排机器上的 experiments 文件夹中。 会创建两个子文件夹:一个用于某个 CometBFT 验证者的 blockstore DB,另一个用于 Prometheus DB 数据。
      • 对 prometheus.zip 文件以及(其中一个)blockstore.db.zip 文件运行 zip -T,以验证数据是否已无误收集。
  8. 清理你的环境。
    • make terraform-destroy;别忘了必须输入 yes 才能完成。

结果提取

为了获得延迟图,请按照上文 200 节点实验的说明操作,但:
  • results.txt 文件只包含一个实验。
  • 因此,不需要任何 for 循环。
至于 Prometheus,可以采用与 200 节点实验相同的方法。
This document provides a detailed description of the QA process. It is intended to be used by engineers reproducing the experimental setup for future tests of CometBFT. The (first iteration of the) QA process as described in the RELEASES.md document was applied to version v0.34.x in order to have a set of results acting as a benchmarking baseline. This baseline is then compared with results obtained in later versions. Out of the testnet-based test cases described in the releases document, we focused on two of them: 200 Node Test and Rotating Nodes Test.

Software Dependencies

Infrastructure Requirements to Run the Tests

  • An account at Digital Ocean (DO), with a high droplet limit (>202)
  • The machine to orchestrate the tests should have the following installed:

Requirements for Result Extraction

  • Prometheus DB to collect metrics from nodes
  • Prometheus DB to process queries (may be a different node from the previous one)
  • blockstore DB of one of the full nodes in the testnet

200 Node Testnet

Running the test

This section explains how the tests were carried out for reproducibility purposes.
  1. [If you haven’t done it before] Follow steps 1-4 of the README.md at the top of the testnet repository to configure Terraform and doctl.
  2. Copy file testnets/testnet200.toml onto testnet.toml (do NOT commit this change).
  3. Set the variable VERSION_TAG in the Makefile to the git hash that is to be tested.
    • If you are running the base test, which implies a homogeneous network (all nodes are running the same version), then make sure makefile variable VERSION2_WEIGHT is set to 0.
    • If you are running a mixed network, set the variable VERSION2_TAG to the other version you want deployed in the network. Then adjust the weight variables VERSION_WEIGHT and VERSION2_WEIGHT to configure the desired proportion of nodes running each of the two configured versions.
  4. Follow steps 5-10 of the README.md to configure and start the 200 node testnet.
    • WARNING: Do NOT forget to run make terraform-destroy as soon as you are done with the tests (see step 9).
  5. As a sanity check, connect to the Prometheus node’s web interface (port 9090) and check the graph for the cometbft_consensus_height metric. All nodes should be increasing their heights.
    • You can find the Prometheus node’s IP address in ansible/hosts under section [prometheus].
    • The following URL will display the metrics cometbft_consensus_height and cometbft_mempool_size:
      http://<PROMETHEUS-NODE-IP>:9090/classic/graph?g0.range_input=1h&g0.expr=cometbft_consensus_height&g0.tab=0&g1.range_input=1h&g1.expr=cometbft_mempool_size&g1.tab=0
      
  6. You now need to start the load runner that will produce transaction load.
    • If you don’t know the saturation load of the version you are testing, you need to discover it.
      • Run make loadrunners-init. This will copy the loader scripts to the testnet-load-runner node and install the load tool.
      • Find the IP address of the testnet-load-runner node in ansible/hosts under section [loadrunners].
      • ssh into testnet-load-runner.
        • Edit the script /root/200-node-loadscript.sh in the load runner node to provide the IP address of a full node (for example, validator000). This node will receive all transactions from the load runner node.
        • Run /root/200-node-loadscript.sh from the load runner node.
          • This script will take about 40 minutes to run, so it is suggested to first run tmux in case the ssh session breaks.
          • It is running 90-second-long experiments in a loop with different loads.
    • If you already know the saturation load, you can simply run the test (several times) for 90 seconds with a load somewhat below saturation:
      • Set makefile variables LOAD_CONNECTIONS, LOAD_TX_RATE to values that will produce the desired transaction load.
      • Set LOAD_TOTAL_TIME to 90 (seconds).
      • Run make runload and wait for it to complete. You may want to run this several times so the data from different runs can be compared.
  7. Run make retrieve-data to gather all relevant data from the testnet into the orchestrating machine.
    • Alternatively, you may want to run make retrieve-prometheus-data and make retrieve-blockstore separately. The end result will be the same.
    • make retrieve-blockstore accepts the following values in makefile variable RETRIEVE_TARGET_HOST:
      • any: (which is the default) picks up a full node and retrieves the blockstore from that node only.
      • all: retrieves the blockstore from all full nodes; this is extremely slow and consumes plenty of bandwidth, so use it with care.
      • the name of a particular full node (e.g., validator01): retrieves the blockstore from that node only.
  8. Verify that the data was collected without errors:
    • at least one blockstore DB for a CometBFT validator
    • the Prometheus database from the Prometheus node
    • for extra care, you can run zip -T on the prometheus.zip file and (one of) the blockstore.db.zip file(s)
  9. Run make terraform-destroy
    • Don’t forget to type yes! Otherwise you’re in trouble.

Result Extraction

The method for extracting the results described here is highly manual (and exploratory) at this stage. The CometBFT team should improve it at every iteration to increase the amount of automation.

Steps

  1. Unzip the blockstore into a directory.
  2. To identify saturation points:
    1. Extract the latency report for all the experiments.
      • Run these commands from the directory containing the blockstore.db folder.
      • It is advisable to adjust the hash in the go run command to the latest possible.
      • mkdir results
        go run github.com/cometbft/cometbft/test/loadtime/cmd/report@3003ef7 --database-type goleveldb --data-dir ./ > results/report.txt
        
    2. File report.txt contains an unordered list of experiments with varying concurrent connections and transaction rate. You will need to separate data per experiment.
      • Create files report01.txt, report02.txt, report04.txt, and for each experiment in file report.txt, copy its related lines to the filename that matches the number of connections, for example:
        for cnum in 1 2 4; do echo "$cnum"; grep "Connections: $cnum" results/report.txt -B 2 -A 10 > results/report$cnum.txt;  done
        
      • Sort the experiments in report01.txt in ascending tx rate order. Likewise for report02.txt and report04.txt.
      • Otherwise just keep report.txt and skip to the next step.
    3. Generate file report_tabbed.txt by showing the contents of report01.txt, report02.txt, report04.txt side by side.
      • This effectively creates a table where rows are a particular tx rate and columns are a particular number of websocket connections.
      • Combine the column files into a single table file:
        • Replace tabs by spaces in all column files. For example, sed -i.bak 's/\t/ /g' results/report1.txt.
      • Merge the new column files into one: paste results/report1.txt results/report2.txt results/report4.txt | column -s $'\t' -t > report_tabbed.txt
  3. To generate a latency vs throughput plot, extract the data as a CSV:
    •  go run github.com/cometbft/cometbft/test/loadtime/cmd/report@3003ef7 --database-type goleveldb --data-dir ./ --csv results/raw.csv
      
    • Follow the instructions for the latency_throughput.py script. This plot is useful to visualize the saturation point.
    • Alternatively, follow the instructions for the latency_plotter.py script. This script generates a series of plots per experiment and configuration that may help with visualizing latency vs throughput variation.

Extracting Prometheus Metrics

  1. Stop the prometheus server if it is running as a service (e.g., a systemd unit).
  2. Unzip the prometheus database retrieved from the testnet, and move it to replace the local prometheus database.
  3. Start the prometheus server and make sure no error logs appear at startup.
  4. Identify the time window you want to plot in your graphs.
  5. Execute the prometheus_plotter.py script for the time window.

Rotating Node Testnet

Running the test

This section explains how the tests were carried out for reproducibility purposes.
  1. [If you haven’t done it before] Follow steps 1-4 of the README.md at the top of the testnet repository to configure Terraform and doctl.
  2. Copy file testnet_rotating.toml onto testnet.toml (do NOT commit this change).
  3. Set variable VERSION_TAG to the git hash that is to be tested.
  4. Run make terraform-apply EPHEMERAL_SIZE=25.
    • WARNING: Do NOT forget to run make terraform-destroy as soon as you are done with the tests.
  5. Follow steps 6-10 of the README.md to configure and start the “stable” part of the rotating node testnet.
  6. As a sanity check, connect to the Prometheus node’s web interface and check the graph for the tendermint_consensus_height metric. All nodes should be increasing their heights.
  7. On a different shell:
    • Run make runload LOAD_CONNECTIONS=X LOAD_TX_RATE=Y LOAD_TOTAL_TIME=Z.
    • X and Y should reflect a load below the saturation point (see, e.g., this paragraph for further info).
    • Z (in seconds) should be big enough to keep running throughout the test, until we manually stop it in step 9. In principle, a good value for Z is 7200 (2 hours).
  8. Run make rotate to start the script that creates the ephemeral nodes and kills them when they are caught up.
    • WARNING: If you run this command from your laptop, the laptop needs to be up and connected for the full length of the experiment.
    • http://<PROMETHEUS-NODE-IP>:9090/classic/graph?g0.range_input=100m&g0.expr=cometbft_consensus_height%7Bjob%3D~%22ephemeral.*%22%7D%20or%20cometbft_blocksync_latest_block_height%7Bjob%3D~%22ephemeral.*%22%7D&g0.tab=0&g1.range_input=100m&g1.expr=cometbft_mempool_size%7Bjob!~%22ephemeral.*%22%7D&g1.tab=0&g2.range_input=100m&g2.expr=cometbft_consensus_num_txs%7Bjob!~%22ephemeral.*%22%7D&g2.tab=0 is an example Prometheus URL you can use to monitor the test case’s progress.
  9. When the height of the chain reaches 3000, stop the make runload script.
  10. When the rotate script has made two iterations (i.e., all ephemeral nodes have caught up twice) after height 3000 was reached, stop make rotate.
  11. Run make stop-network.
  12. Run make retrieve-data to gather all relevant data from the testnet into the orchestrating machine.
  13. Verify that the data was collected without errors:
    • at least one blockstore DB for a CometBFT validator
    • the Prometheus database from the Prometheus node
    • for extra care, you can run zip -T on the prometheus.zip file and (one of) the blockstore.db.zip file(s)
  14. Run make terraform-destroy
Steps 8 to 10 are highly manual at the moment and will be improved in next iterations.

Result Extraction

In order to obtain a latency plot, follow the instructions above for the 200 node experiment, but the results.txt file contains only one experiment. As for Prometheus, the same method as for the 200 node experiment can be applied.

Vote Extensions Testnet

Running the test

This section explains how the tests were carried out for reproducibility purposes.
  1. [If you haven’t done it before] Follow steps 1-4 of the README.md at the top of the testnet repository to configure Terraform and doctl.
  2. Copy file varyVESize.toml onto testnet.toml (do NOT commit this change).
  3. Set variable VERSION_TAG in the Makefile to the git hash that is to be tested.
  4. Follow steps 5-10 of the README.md to configure and start the testnet.
    • WARNING: Do NOT forget to run make terraform-destroy as soon as you are done with the tests.
  5. Configure the load runner to produce the desired transaction load.
    • Set makefile variables ROTATE_CONNECTIONS, ROTATE_TX_RATE to values that will produce the desired transaction load.
    • Set ROTATE_TOTAL_TIME to 150 (seconds).
    • Set ITERATIONS to the number of iterations that each configuration should run for.
  6. Execute steps 5-10 of the README.md file at the testnet repository.
  7. Repeat the following steps for each desired vote_extension_size:
    1. Update the configuration (you can skip this step if you didn’t change the vote_extension_size).
      • Update the vote_extensions_size in the testnet.toml to the desired value.
      • make configgen
      • ANSIBLE_SSH_RETRIES=10 ansible-playbook ./ansible/re-init-testapp.yaml -u root -i ./ansible/hosts --limit=validators -e "testnet_dir=testnet" -f 20
      • make restart
    2. Run the test.
      • make runload This will repeat the tests ITERATIONS times every time it is invoked.
    3. Collect your data.
      • make retrieve-data Gathers all relevant data from the testnet into the orchestrating machine, inside folder experiments. Two subfolders are created: one blockstore DB for a CometBFT validator and one for the Prometheus DB data.
      • Verify that the data was collected without errors with zip -T on the prometheus.zip file and (one of) the blockstore.db.zip file(s).
  8. Clean up your setup.
    • make terraform-destroy; don’t forget that you need to type yes for it to complete.

Result Extraction

In order to obtain a latency plot, follow the instructions above for the 200 node experiment, but:
  • The results.txt file contains only one experiment.
  • Therefore, no need for any for loops.
As for Prometheus, the same method as for the 200 node experiment can be applied.