Deploy alerting and route notifications
This guide brings up the two components that turn the metrics store into notifications, and says where a notification goes. At the end vmalert evaluates the repository's rules against VictoriaMetrics, Alertmanager groups what fires and hands it to a receiver the deployment names, and both UIs answer on the hostnames the overlay patches. What the rules ask and why the pair is split this way is in alerting design.
Before you start
- A cluster with VictoriaMetrics running, which every rule expression is evaluated against.
- An overlay of your own over
deploy/kubernetes/base, since a deployment patches the external URLs and replaces the Alertmanager config rather than editing the base. - Docker on the host, which is all
make check-alertingneeds. - A shell at the repository root for
make upandmake -s ca, andkubectlagainst the cluster for the silence. - The alert rules reference page, which states every rule with its expression, its
forand its labels.
Publish both components
Patch the two external URLs in the overlay, one per component. The base sets
VMALERT_EXTERNAL_URLto the placeholderhttps://vmalert.tally.example.comandALERTMANAGER_EXTERNAL_URLtohttps://alertmanager.tally.example.com, and an unpatched value sends a reader who follows a Source link to a host that answers nothing:yamlpatches: - patch: | apiVersion: apps/v1 kind: Deployment metadata: name: vmalert spec: template: spec: containers: - name: vmalert env: - name: VMALERT_EXTERNAL_URL value: https://vmalert.tally.127-0-0-1.nip.io:8443 # The same placeholder on the other side of the pair. - patch: | apiVersion: apps/v1 kind: StatefulSet metadata: name: alertmanager spec: template: spec: containers: - name: alertmanager env: - name: ALERTMANAGER_EXTERNAL_URL value: https://alertmanager.tally.127-0-0-1.nip.io:8443The pair above is the one in
deploy/kubernetes/overlays/dev/kustomization.yaml. Each overlay patches the variable rather than the argument, because-external.urlreads its value from the argument and Kubernetes expands$(VAR)in an argument from the container's own environment.Patch the route hostname of each component together with its variable. The base leaves both hostnames at
vmalert.tally.example.comandalertmanager.tally.example.com, and both external URLs at the matching placeholder, which is the pairing the dev overlay keeps.Neither route carries a credential. Whatever reaches the Alertmanager hostname reads every firing alert with its labels, which by the rules of
rules.yamlrender the cloud names and the exporter and target addresses of the OpenStack control plane, and the receivers the deployment delivers to; whatever reaches the vmalert hostname reads every rule expression and every active alert with the same labels. What each route withholds — the write paths,/api/v2/statusand/metricson the one, the root prefix with/flags,/metricsand/debug/pprofon the other — is in what the dev route publishes; the reads above are not among them. On the dev overlay both hostnames resolve to127.0.0.1and the Gateway is bound to the developer's own machine. A deployment that puts them anywhere else puts an authenticating proxy in front of them, or publishes neither route and reaches both from inside the cluster.
Name a receiver
Add the deployment's own delivery by replacing the generated ConfigMap from the overlay rather than by patching the file in the base:
yamlconfigMapGenerator: - name: alertmanager-config behavior: replace files: - config.yamlCarry the
webhook_configsor theemail_configsunder the receiver of the overlay's ownconfig.yaml. That is the mechanism replace the scrape targets records for the scrape config, used here for the same reason: a file the base owns stays the base's, and the deployment's copy is a file of its own. Which route a delivery is reached through is in where an alert is delivered.
Edit the rules
Edit
rules.yaml,scrape-rules.yamlorconfig.yamlin its component's directory. None of them is applied as a file: each is the source of aconfigMapGeneratorin its component'skustomization.yaml, generatingvmalert-rules,vmalert-scrape-rulesandalertmanager-config, and kustomize appends a content hash to each name. Editing a file changes the generated name, which changes the pod spec that mounts it, which rolls the pod. Neither component is left evaluating rules or routing by a config that no longer matches the tree.scrape-rules.yamlnames the discovered scrape jobs, so an overlay that replaces the scrape config replacesvmalert-scrape-ruleswith it, as replace the scrape targets describes.Validate the files before the commit:
shmake check-alertingIt loads
rules.yamlandscrape-rules.yamlinto vmalert with-dryRunand runsamtool check-configoverconfig.yaml, each in the image the cluster runs, so an expression or a routing field the pinned version refuses fails here rather than in the cluster.
Reach both on a dev cluster
Bring the cluster up.
make uppublishes both on the dev overlay's hostnames:texthttps://vmalert.tally.127-0-0-1.nip.io:8443/vmalert/ https://alertmanager.tally.127-0-0-1.nip.io:8443Reach either with the dev CA that signed the Gateway's certificate:
shmake -s ca > tally-ca.crt curl --cacert tally-ca.crt \ https://vmalert.tally.127-0-0-1.nip.io:8443/api/v1/rules
Silence an alert
Create the silence from inside the cluster, with the
amtoolthe image ships:shkubectl --context kind-tally -n tally exec statefulset/alertmanager -- \ amtool --alertmanager.url=http://127.0.0.1:9093 silence add \ --author=<who> --comment='<why>' --duration=2h alertname=<name>Which write paths the dev route refuses, and how a refused one is answered, is in what the dev route publishes.
Check the result
About five minutes after
make up,TallyScrapeTargetDownfires for the jobsopenstack-db-exporter,ceilometerandopenstack-collector. All three are static targets of the dev overlay:ceilometernames an exporter that runs beside an OpenStack control plane rather than in this cluster, and the other two read the simulator and the collector of the compose stack, which run only whilemake simulator-uppublishes a month. That is the designed dev state described under replace the scrape targets, and not a fault to chase. What to do when it fires against a deployment is in its runbook,TallyScrapeTargetDown.No other rule fires on a cluster nothing has reported to.
TallyScrapeJobMissingstays silent because the jobs the dev overlay'sscrape-rules.yamlnames,reporting-apiandotel-collector, resolve to targets, andTallyExporterServiceSilentneeds an exporter target that answers a scrape, which is the target that is down. The remaining seven readtally_series the store does not carry yet, and an expression over nothing returns nothing.TallyRecordedSeriesMissingis quiet for a different reason: itsabsent()clause is true here, and thecount(tally_current_resources) > 0it is paired with is what keeps it from reporting a cluster that has not reconciled yet as a stalled write path.When an alert you expect does not appear, read
/api/v1/ruleson the vmalert hostname and find the rule. A non-emptylastErroron it means the datasource query failed rather than that the condition was false, and the message names what VictoriaMetrics answered.