Know before your users do
Built-in health checks watch replication lag, disk, memory and failed queries — each with tunable warning and critical thresholds and a full fire-and-recovery history.
The checks that matter, pre-built
Replication lag, disk pressure, memory, failed queries, detached parts and more — every check ships with sensible defaults and editable warning/critical thresholds.
- Editable thresholds per check
- Fire and recovery history per check
- Per-host status rollup
One webhook, every channel
Point chmonitor at a single webhook URL and alerts land in Slack, Discord, PagerDuty or Opsgenie — fired on threshold breach, resolved on recovery.
- Breach and recovery notifications
- No per-channel integrations to maintain
- Works on every deploy target
From alert to diagnosis in one click
Any failing check carries an AI audit action: it opens the agent with the check context pre-loaded, so the diagnosis starts from the failure, not from a blank prompt.
- Check context handed to the agent automatically
- Schema-aware root-cause suggestions
- Read-only — the agent recommends, you apply
Everything in the box
Replication lag
Warn before replicas fall behind.
Disk & memory
Capacity pressure with critical thresholds.
Failed queries
Error-rate spikes surfaced early.
Webhooks
Slack, Discord, PagerDuty, Opsgenie — one URL.
History
Every fire and recovery, timestamped.
AI audit
One-click diagnosis for any failing check.
Common questions
Do I need an external alerting stack?
No. Checks are evaluated by chmonitor itself; you only supply a webhook URL for delivery.
Can I change what counts as critical?
Yes — every check has editable warning and critical thresholds, per check.
Will I get a notification when things recover?
Yes, recovery fires its own notification and is recorded in the check history.
More chmonitor features
Start monitoring — advisor and alerts built in
Open the hosted dashboard and connect a cluster, or self-host with Docker, Kubernetes, or from source. Self-hosting is always free.