Monitoring and tuning
Server Monitoring and Performance Optimization
I set up server monitoring that tells the right person about real problems, then tune performance in a fixed order of baseline, bottleneck, one change, and measure again. It is for hosting owners and site operators whose servers feel slow, fill their disks or only get checked after a customer complains.
By Shahid Malla, WHMCS developer and hosting infrastructure engineer · Updated
What should I monitor on a hosting server?
Monitor what visitors see from outside, plus the resources that run out first on a hosting server: CPU and disk I/O, memory, disk space and inodes, the database, the mail queue and the services themselves. Thresholds come from your baseline, so the table describes what to watch and not fixed numbers.
| Signal | How I measure it | When it should alert |
|---|---|---|
| Uptime from outside | A keyword check on a real page, plus TLS expiry and DNS, from a network other than the server's | Repeated failures, or a certificate close to expiry |
| CPU load and IO wait | uptime, vmstat, iostat -x, with history in sar | Sustained load above the core count, or IO wait well above baseline |
| Memory and swap | free -m, the swap columns of vmstat, OOM killer lines in the kernel log | Sustained swapping, or any OOM kill |
| Disk space and inodes | df -h and df -i | A trend toward full. Inode exhaustion shows "No space left on device" while df -h still shows free space. |
| MySQL or MariaDB | Slow query log, thread and connection counts | Rising slow queries, or "Too many connections" errors |
| Mail queue | exim -bpc | A growing queue, or a sudden spike in outbound mail |
| Service health | cPanel's service monitor, chkservd, plus an independent check on httpd, exim, dovecot, mysqld, named and sshd | A service down, or restarting repeatedly |
| Scheduled jobs | A heartbeat ping at the end of each backup and WHMCS cron run | A heartbeat that fails to arrive |
Who gets the alert, and how?
Each alert goes to a named person through a channel that still works when the server is down, and its urgency decides whether it interrupts them.
- Urgent. Site down, disk nearly full, mail queue exploding. A phone push or SMS to whoever is on duty.
- Warning. A trend toward a limit. An email or chat message to read in working hours.
- Digest. Slow query counts and resource trends in a weekly summary.
Never route alerts through the server that may be failing, because Exim on a dead machine cannot report its own death. I test routing by firing each alert on purpose, for example by stopping a non-critical service. An alert nobody acts on gets removed, since constant noise teaches people to ignore the one that matters. For backup and cron jobs, the heartbeat checks in the table cover the failure that produces no error at all: the job that never ran.
In what order should I tune a slow server?
Tune in a fixed order: baseline, find the bottleneck, change one thing, measure again. Skipping a step usually moves the problem somewhere else.
- Baseline. Record load, IO wait, memory, slow query count and response times for a few key URLs, at a quiet time and at a busy one.
- Find the bottleneck. Ask which resource saturates first. These commands show it.
- Change one thing. One setting, with a rollback note.
- Measure again under comparable load, then keep or revert the change on the evidence.
uptime vmstat 5 5 iostat -x 5 3 ps -eo user,pcpu,pmem,comm --sort=-pcpu | head
How do I tune the web server and PHP-FPM?
Size PHP-FPM workers to the memory you actually have, not the traffic you hope for, then check the web server and OPcache against measurements.
- PHP-FPM workers. Set
pm.max_childrenfrom the RAM reserved for PHP divided by average memory per worker. Too high pushes the server into swap. Too low queues requests, and the FPM log then reports that the pool reachedpm.max_children. Usepm.max_requeststo recycle leaking workers andrequest_slowlog_timeoutto catch slow scripts. - Web server. On Apache, I use the event MPM with PHP-FPM and tune keep-alive and connection limits from the access logs. LiteSpeed has its own settings, which are tuned the same way, from measurements.
- OPcache. Make
opcache.memory_consumptionandopcache.max_accelerated_fileslarge enough that the cache never fills.opcache_get_status()shows the hit rate and whether it ran out, because scripts that do not fit are recompiled on every request.
How do I tune MySQL or MariaDB?
Find the slow queries first and size the buffer pool second, because most database slowness comes from queries and missing indexes, not server variables.
I enable the slow query log (slow_query_log and long_query_time) for a representative period, then run EXPLAIN on the worst queries and fix the query or add an index. For InnoDB, innodb_buffer_pool_size should fit the working set while leaving memory for PHP-FPM and the operating system. Raising max_connections to hide a connection leak only moves the failure. Scripts such as mysqltuner give hints that I treat as questions to test and not as rules. I try variables live with SET GLOBAL, measure, and only then write them to the config file.
When does object caching help, and how do I find a noisy neighbor?
Object caching helps dynamic sites that repeat the same database queries, and a noisy neighbor is one account using most of a shared server's CPU or disk. You find both by measuring.
Redis or Memcached keep repeated query results in memory. WordPress benefits through a persistent object cache plugin, while a static site gains nothing. On shared servers each account needs its own cache instance, or a separate prefix and authentication, so accounts cannot read each other's data. To find a noisy neighbor, I total CPU and memory per user from the process list, then read the busiest URLs from that account's logs (on EasyApache 4 the domain logs sit under /etc/apache2/logs/domlogs). Common culprits are constant wp-cron.php calls, xmlrpc.php floods and bots crawling search pages. The fixes are caching, rate limits, bot rules and per-account PHP-FPM limits. Where hard per-account CPU, memory and I/O caps are required, commercial isolation software such as CloudLinux is an option.
What I will not promise about performance
I do not promise speed-up percentages, uptime figures or benchmark scores, because they depend on your application, traffic and hardware. I report what I measured before and after on your server. If the baseline shows the machine is truly out of CPU, memory or disk, I recommend more capacity, and I do not work on hardware faults. Monitoring is typically in place within a day or two, with tuning time depending on the findings. I quote a fixed price or bill $55 to $65 per hour. Related reading covers the first-day cPanel settings, log review under hardening, backup tests and WHMCS support and maintenance.
Who this is for
- Hosting owners who hear about outages from their customers
- Operators of shared servers where one or two accounts use most of the resources
- Site owners on a VPS whose pages slow down at busy times
- Teams whose alerts are either silent or constant
What is included
- External uptime, certificate expiry and DNS checks from outside your network
- Metrics for CPU load, IO wait, memory, swap, disk space and inodes
- MySQL or MariaDB slow query logging and connection tracking
- Mail queue, service health and backup heartbeat checks
- Alert routing by urgency to named people, tested by firing each alert
- A recorded baseline before any tuning begins
- Tuning of web server, PHP-FPM, OPcache, database and caching, one change at a time
- Written notes on what changed and what was measured before and after
How the work runs
-
1
Baseline
I record normal load, key page response times, slow queries and resource use across a typical period and a busy one, so later changes are judged against facts.
-
2
Instrument
I install monitoring and alerting, then test the routing by triggering each alert on purpose and confirming the right person receives it.
-
3
Diagnose
I find the resource that saturates first, whether CPU, memory, disk I/O, the database or the network, using measurements and not assumptions.
-
4
Change one thing
I make a single change with a rollback note, measure again under comparable load and keep or revert it on the evidence before the next change.
-
5
Hand over
You receive the dashboards, the alert list and notes on every change with its before and after numbers. Support continues for two weeks.
Frequently asked questions
Why monitor from outside the server as well as inside?
A server that is down cannot report its own outage. An external check sees what visitors see: DNS resolution, the network path, the TLS certificate and the page itself. Internal metrics then explain why. I use both, and I send alerts through a channel that does not depend on the monitored server.
How do you avoid a flood of alerts?
Every alert must trigger a defined action, and thresholds come from the recorded baseline instead of generic defaults. Urgent alerts interrupt a person, warnings wait for working hours and trends go into a weekly digest. If nobody acts on an alert for a month, I remove it or downgrade it.
Do you tune before measuring anything?
No. Without a baseline you cannot tell whether a change helped, and a change made blind often moves the bottleneck instead of removing it. I record the current state first, change one setting at a time and compare results under similar load.
Can you tell me how much faster my sites will be?
No, and anyone who names a figure before measuring is guessing. The gain depends on what limits your server today, which might be memory, slow queries or a single heavy account. I report what I measured before and after on your server, including changes that made no difference.
Which monitoring tools do you use?
I choose by size. One or two servers are well served by a hosted uptime service plus a light agent. A larger fleet suits a self-hosted stack such as Prometheus with Alertmanager, Zabbix or Netdata. You own the accounts and the data, so you can change tools later.
Do I need a bigger server instead of tuning?
Sometimes. The baseline shows whether memory, CPU, disk or software configuration is the limit. Tuning often finds easy gains, but if the machine is genuinely out of a resource at normal load, the honest answer is more capacity, and I will say that instead of tuning around it.
Related services
-
cPanel and WHM Setup Service
I install and configure cPanel and WHM on a new VPS or dedicated server, from license and hostname to mail aut...
-
Server Hardening Service for Linux and cPanel Servers
I harden Linux and cPanel servers with SSH keys, firewall review, patching, PHP isolation and protected backup...
-
cPanel Server Migration and Backup Setup
I move cPanel accounts, mail and DNS to a new server with a tested cutover, then set up local and off-site bac...
-
WHMCS Support and Maintenance
Ongoing WHMCS care, hourly or monthly: staged upgrades, PHP changes, cron and backup checks, module updates an...
Ready to talk about your project?
Send the details and I reply within one business day with questions, an estimate and a plan.