Alerting

Know before your users do

Built-in health checks watch replication lag, disk, memory and failed queries — each with tunable warning and critical thresholds and a full fire-and-recovery history.

Native adaptersSlack, Discord, PagerDuty, Opsgenie, and more
Tunablewarning and critical thresholds per check
Historyfull fire and recovery timeline
AI auditone-click prompt for failing checks
Health checks

The checks that matter, pre-built

Replication lag, disk pressure, memory, failed queries, detached parts and more — every check ships with sensible defaults and editable warning/critical thresholds.

  • Editable thresholds per check
  • Fire and recovery history per check
  • Per-host status rollup
Notifications

Native adapters, per-rule routing

Alerts go to Slack, Discord, Teams, Google Chat, Telegram, ntfy, Pushover, Twilio SMS, Opsgenie, PagerDuty, and healthchecks.io — fired on threshold breach, resolved on recovery. Route per rule or host; a generic webhook is optional, not the only path.

  • Breach and recovery notifications
  • Per-rule and per-host routing
  • Works on every deploy target
AI audit

From alert to diagnosis in one click

Any failing check carries an AI audit action: it opens the agent with the check context pre-loaded, so the diagnosis starts from the failure, not from a blank prompt.

  • Check context handed to the agent automatically
  • Schema-aware root-cause suggestions
  • Read-only — the agent recommends, you apply
Capabilities

Everything in the box

Replication lag

Warn before replicas fall behind.

Disk & memory

Capacity pressure with critical thresholds.

Failed queries

Error-rate spikes surfaced early.

Channels

Slack, Discord, PagerDuty, Opsgenie, and more — per-rule routing.

History

Every fire and recovery, timestamped.

AI audit

One-click diagnosis for any failing check.

FAQ

Common questions

Do I need an external alerting stack?

No. Checks are evaluated by chmonitor itself. Configure the channel you use — native adapters or an optional generic webhook.

Can I change what counts as critical?

Yes — every check has editable warning and critical thresholds, per check.

Will I get a notification when things recover?

Yes, recovery fires its own notification and is recorded in the check history.

Start monitoring — advisor and alerts built in

Open the hosted dashboard and connect a cluster, or self-host with Docker, Kubernetes, or from source. Self-hosting is always free.