How to set real-time alarm notification of CPU and bandwidth by using Alibaba Cloud Cloud Monitoring

cloud 2026-10-02 阅读 3
3

In the technical fields of operations and system administration, many newcomers to cloud computing often make a critical mistake:

Once the server is purchased and the application is deployed, people often assume that everything is set and done.

.

Until one day in the middle of the night, the traffic suddenly soared, causing the bandwidth to be stuck, or a dead cycle process directly pushed the CPU to 100 percent, causing the system to go down, the customer's phone was blown up, and the operation and maintenance personnel hurriedly logged into the background to check the log. This kind of "passive fire" pain, all experienced engineers do not want to experience a second time.

In fact, Alibaba Cloud itself comes with a very powerful infrastructure monitoring tool—

CloudMonitor

. As long as you properly configure the real-time alarm notification of CPU and network bandwidth, you can receive notification within the first few minutes of changes in system indicators, thus easily handling the crisis before it spreads.

Next, I will combine real operation and maintenance experience to build an efficient and zero-miss real-time alarm system of Aliyun.

1. Why are the CPU and bandwidth the "lifeline" of alerts?

In the monitoring indicators of cloud servers (ECS), there are various alarm items,

CPU usage

and

Inbound/Outbound Bandwidth

It has always been the core of cores.

CPU usage: reflects the server's computational load. If the CPU is maintained above 90% for a long time, it will lead to a sharp increase in HTTP request delay, and it will trigger system OOM (memory overflow) or direct jam.

Network bandwidth (NetworkOut / NetworkIn): reflects the data transfer load. Once the export bandwidth is full (for example, you bought 5Mbps bandwidth and ran to 4.9Mbps), the server will have serious packet loss and network delay, and the external user will look like "the website cannot be opened" or "the interface timed out".

By keeping an eye on these two indicators, more than 90% of server infrastructure failures can be warned in advance.

2. Preparatory work: the "infrastructure" for alert notifications

Before configuring rules in the CloudMonitor console, we need to first set up the notification channels. If an alert is triggered but there is no designated recipient, even the most perfect rule is rendered useless.

1. Verify the status of the CloudMonitor agent.

Log in to the Alibaba Cloud console and go to

Cloud Monitoring -> Host Monitoring

. Ensure that on your ECS instance

ArgusAgent

The plugin status is displayed as

"Running"

. Only when the plugin is functioning properly can Cloud Monitoring collect more granular metrics from within the system.

2. Configure alert contacts and contact groups

Click the following in the left-hand navigation bar, in order:

"Alarm Service" -> "Alarm Contacts"

:

Create contact person: fill in the mobile phone number and e-mail address of the operation and maintenance personnel or developer, and bind the DingTalk robot (WebHook) or flying book/enterprise WeChat robot

.

Create a contact group: Pull related contacts into the same group (for example, "Core Operations Group" or "Duty Personnel Group").

A word of advice: We strongly recommend integrating a DingTalk/Feishu group robot. Compared with the traditional mail (easy to hang) and SMS (easy to be harassed and intercepted), the response speed of the group robot with @ everyone function in emergency is the fastest.

3. Hands-on Exercise: Step-by-step Configuration of CPU and Bandwidth Alert Rules

After the preparations are ready, we formally enter the process of creating alarm rules.

1. Access the alert rule creation entry: Console navigation.

Log on to the Alibaba Cloud console. In the top search bar, enter CloudMonitor. Expand the menu bar on the left

Alarm service

->

Alarm Rules

, click in the page

Create alarm rules

Button.

2. Select Associated Resources and Products: Locate Monitoring Targets.

Product Type: Select ECS (if it is EIP Bandwidth Plan or SLB, select the corresponding product).

Resource range: It is recommended to select an instance and select the core server to be monitored. If the number of servers is large, the following can be based on the "application group" for unified management.

3. Configure CPU utilization alert rules: Core computing metric.

In the Add Rule panel, add the first monitoring metric:

Monitoring Metrics: Select (ECS)CPU Usage (cpu_total).

Threshold and level configuration: Urgent (Critical): 3 consecutive cycles (default 1 minute),CPU usage $\ge 90\%$. Warn: 3 consecutive cycles, CPU usage $\ge 80\%$.

Channel selection: emergency level check phone + SMS + DingTalk; Warning level check email + DingTalk.

4. Configure network bandwidth alert rules: network throughput metric.

Continue to click Add Rule to configure network bandwidth metrics:

Monitoring Metrics: Select (ECS) public network outflow bandwidth (IntranetOut or InternetOut, depending on whether the business goes through the public network or intranet).

Threshold setting technique: Bandwidth alarm cannot be set blindly, but should be combined with your ECS actual configuration. For example, if the maximum public network bandwidth you purchased is 10 Mbps, the warning line can be set to: warning level: outflow rate $\ge 8\text{ Mbps}$(that is, $80\%$that reaches the maximum). Emergency level: Outflow rate $\ge 9.5\text{ Mbps}$(to be capped).

5. Configure notification delivery and the effective time: Complete rule creation.

Alarm Contact Group: Select the previously created Core O & M Group ".

Anti-Disturb Settings: Set the effective time of the rule (for example, 24 hours a day).

Advanced configuration: Set the alarm repetition frequency to "5 minutes/time" or "15 minutes/time" to avoid alarm storm impact.

Click OK after confirmation

Done. Creation completed.

4. Verification and troubleshooting logic after an alarm is triggered

After the rules are set, how to verify that this set of alarm processes is effective and smooth?

1. Verification method (How to trigger the test?)

You can use the tools that come with the Linux system to simulate stress testing:

CPU pressure test: Run stress -- cpu 2 -- timeout the 300s command line in the test environment to manually load up the CPU.

Bandwidth pressure measurement: Use iperf3 or download large files from the server to the external network to increase the bandwidth.

Verification criteria: Observe whether the DingTalk group or mobile phone text messages can accurately receive the alarm notification sent by Alibaba Cloud within 3-5 minutes.

Production environment test prevention warning

Stress testing directly in the production environment is strictly prohibited! Be sure to select a test server or create a new temporary instance to verify the alarm link.

2.3-step gold processing after the alarm is triggered.

When receiving a CPU or bandwidth alarm, do not restart the server blindly. It is recommended to take the following step-by-step troubleshooting route:

[Alert notification received]

│

├──> CPU alarm ──> login to the server ──> run 'top' / 'htop' ──> locate the high-occupancy PID ──> check the process log or kill the abnormal thread

│

these-> bandwidth alarm-> log on to the console-> view traffic monitoring curve-> run 'iftop'/'nethogs'-> identify abnormal connection IP-> configure Security group to block or expand bandwidth

Extended Thinking of 5. Enterprise Operation and Maintenance: Standardized Management of Account and Resources

In the process of building a complete set of monitoring alarm system, in addition to the technical configuration itself, many enterprises tend to ignore

Security of underlying infrastructure assets and account compliance

.

With the expansion of business scale, many teams will face the needs of multi-account management, independent project settlement or overseas business expansion. In this process, the aspects involved are

Alibaba Cloud account purchase

Therefore, account real-name verification and permission isolation become particularly crucial.

Permission Minimization Principle: Never directly assign AdministratorAccess global highest permissions to O & M engineers or monitoring services. It is recommended that you create a dedicated role through RAM (Access Control) and only grant read and write permissions (such as AliyunCloudMonitorFullAccess) to CloudMonitor (CMS).

Account architecture: For companies that need to isolate multiple environments (development, testing, and production), reasonable Alibaba Cloud account purchase and architecture planning can isolate risks from the source. The production environment's alarms are directly connected to the core operations team, while the development ring

Environmental alarms are distributed to specific research and development personnel, so that they do not interfere with each other and have clear rights and responsibilities.

Lifecycle management: Whether it is the renewal of the cloud server ECS or the evolution of monitoring alarm rules, it must be strongly bound to the resource lifecycle under the account. If the server is released, the corresponding alarm rules should be cleared synchronously to avoid invalid alarm garbage.

Conclusion

Real-time alarm is not to increase the workload of operation and maintenance, on the contrary, it is to "reduce the burden" for engineers ".

An Alibaba Cloud CPU and bandwidth alarm system with reasonable design and scientific thresholds is like installing "security guards" on servers that patrol 24 hours a day ". When everything is normal, you can sleep peacefully; when the risk is first revealed, it will send the most accurate information to you in the first place to help you kill the fault in the cradle.

Before leaving work tonight, you might as well take 10 minutes to log on to ariyun console and check if your server alarm rules are set correctly?

2
← 返回新闻中心