Grafana dashboards
Grafana provisions four dashboard files from deploy/kubernetes/base/grafana/dashboards, which reach the pod through a generated ConfigMap. The provider in provisioning/dashboards/tally.yaml loads them into the folder Tally with disableDeletion: true and allowUiUpdates: false, so an edit or a delete in the UI is refused and the files in the repository are what a cluster serves.
Every panel target names the datasource uid victoriametrics, which is fixed in provisioning/datasources/vm.yaml rather than generated. That datasource points at the vmauth container in the Grafana pod, on this pod's loopback, which carries the Prometheus read paths on to the store and refuses everything else.
dashboards_test.go pins the file set the ConfigMap ships, that every file parses, the uids and the titles, the datasource uid on every query target, the variables the expressions read, the refresh of the project_id list, the unlimited reading of the quota gauges, and the drift note of the reconciliation dashboard.
Dashboards
A panel with no query target is listed with none as its expression.
fleet-overview.json
Title Tally / Fleet Overview, uid tally-fleet-overview.
| Variable | Multi | Query |
|---|---|---|
platform | yes | label_values(tally_current_resources, platform) |
cloud | yes | label_values(tally_current_resources{platform=~"$platform"}, cloud) |
| Panel | Type | Expression |
|---|---|---|
| Resources by type and state | barchart | sum by (resource_type, state) (tally_current_resources{platform=~"$platform", cloud=~"$cloud"}) |
| Resource count trend | timeseries | sum by (resource_type, state) (tally_current_resources{platform=~"$platform", cloud=~"$cloud"}) |
| Clouds reporting | stat | count(count by (cloud) (tally_current_resources)) |
| Projects (OpenStack) | stat | sum(openstack_identity_projects{cloud=~"$cloud"}) |
| Top 10 projects by instance count | table | topk(10, sum by (tenant_id) (openstack_nova_limits_instances_used{cloud=~"$cloud"})) |
ingestion-health.json
Title Tally / Ingestion Health, uid tally-ingestion-health.
| Variable | Multi | Query |
|---|---|---|
platform | yes | label_values(tally_current_resources, platform) |
cloud | yes | label_values(tally_current_resources{platform=~"$platform"}, cloud) |
| Panel | Type | Expression |
|---|---|---|
| Event ingest rate | timeseries | sum by (cloud, source) (rate(tally_events_ingested_total{platform=~"$platform", cloud=~"$cloud"}[5m])) |
| Dedup rate | timeseries | sum by (cloud) (rate(tally_events_deduplicated_total{cloud=~"$cloud"}[5m])) |
| Rejected events | timeseries | sum by (cloud, reason) (increase(tally_events_rejected_total{cloud=~"$cloud"}[1h])) |
| Collector buffer depth | timeseries | tally_collector_buffer_depth{platform=~"$platform", cloud=~"$cloud"} |
| Oldest buffered event age | stat | tally_collector_oldest_buffered_seconds{platform=~"$platform", cloud=~"$cloud"} |
| Projection replays | timeseries | sum by (cloud) (rate(tally_projection_replays_total{cloud=~"$cloud"}[15m])) |
| Scrape health | stat | up |
project-drilldown.json
Title Tally / Project Drilldown, uid tally-project-drilldown.
| Variable | Multi | Query |
|---|---|---|
platform | yes | label_values(tally_current_resources, platform) |
cloud | yes | label_values(tally_current_resources{platform=~"$platform"}, cloud) |
project_id | no | label_values(openstack_nova_limits_instances_used, tenant_id) |
api_base | no | https://api.tally.example.com |
| Panel | Type | Expression |
|---|---|---|
| Resources by type | timeseries | count(openstack_nova_server_status{cloud=~"$cloud", tenant_id=~"$project_id"}) |
| Resources by type | timeseries | count(openstack_cinder_volume_status{cloud=~"$cloud", tenant_id=~"$project_id"}) |
| Resources by type | timeseries | count(openstack_neutron_floating_ip{cloud=~"$cloud", project_id=~"$project_id"}) |
| Resources by type | timeseries | count(openstack_neutron_router{cloud=~"$cloud", project_id=~"$project_id"}) |
| Resources by type | timeseries | count(openstack_glance_image_bytes{cloud=~"$cloud", tenant_id=~"$project_id"}) |
| Resources by type | timeseries | count(openstack_loadbalancer_loadbalancer_status{cloud=~"$cloud", project_id=~"$project_id"}) |
| Volume capacity | timeseries | sum(openstack_cinder_volume_gb{cloud=~"$cloud", tenant_id=~"$project_id"}) |
| Quota usage | gauge | sum(openstack_nova_limits_instances_used{cloud=~"$cloud", tenant_id=~"$project_id"} and (openstack_nova_limits_instances_max{cloud=~"$cloud", tenant_id=~"$project_id"} >= 0)) / (sum(openstack_nova_limits_instances_max{cloud=~"$cloud", tenant_id=~"$project_id"} >= 0) > 0) or (max(openstack_nova_limits_instances_max{cloud=~"$cloud", tenant_id=~"$project_id"}) == -1) |
| Quota usage | gauge | sum(openstack_nova_limits_vcpus_used{cloud=~"$cloud", tenant_id=~"$project_id"} and (openstack_nova_limits_vcpus_max{cloud=~"$cloud", tenant_id=~"$project_id"} >= 0)) / (sum(openstack_nova_limits_vcpus_max{cloud=~"$cloud", tenant_id=~"$project_id"} >= 0) > 0) or (max(openstack_nova_limits_vcpus_max{cloud=~"$cloud", tenant_id=~"$project_id"}) == -1) |
| Quota usage | gauge | sum(openstack_nova_limits_memory_used{cloud=~"$cloud", tenant_id=~"$project_id"} and (openstack_nova_limits_memory_max{cloud=~"$cloud", tenant_id=~"$project_id"} >= 0)) / (sum(openstack_nova_limits_memory_max{cloud=~"$cloud", tenant_id=~"$project_id"} >= 0) > 0) or (max(openstack_nova_limits_memory_max{cloud=~"$cloud", tenant_id=~"$project_id"}) == -1) |
| Recent lifecycle activity | timeseries | sum by (event_type) (rate(tally_events_ingested_total{cloud=~"$cloud"}[5m])) |
reconciliation-drift.json
Title Tally / Reconciliation Drift, uid tally-reconciliation-drift.
| Variable | Multi | Query |
|---|---|---|
platform | yes | label_values(tally_current_resources, platform) |
cloud | yes | label_values(tally_current_resources{platform=~"$platform"}, cloud) |
| Panel | Type | Expression |
|---|---|---|
| Reconciled resources by action | timeseries | sum by (cloud, action) (increase(tally_sync_resources_reconciled_total{cloud=~"$cloud"}[1h])) |
| Sync errors | timeseries | sum by (cloud) (increase(tally_sync_errors_total{cloud=~"$cloud"}[1h])) |
| Sync runs by status | stat | sum by (cloud, status) (increase(tally_sync_runs_total{cloud=~"$cloud"}[6h])) |
| Drift interpretation | text | none |
What each panel needs
Each panel reads the series of one source, and a source reaches the store through one scrape job or push.
| Source | Scrape job or push | Panels |
|---|---|---|
The Reporting API's tally_ series | reporting-api | Fleet Overview: Resources by type and state, Resource count trend, Clouds reporting. Ingestion Health: Event ingest rate, Dedup rate, Rejected events, Projection replays. Project Drilldown: Recent lifecycle activity. Reconciliation Drift: Reconciled resources by action, Sync errors, Sync runs by status, and Drift interpretation, a text panel that reads no series. The platform and cloud variables of every dashboard. |
The OpenStack collector's tally_collector_* series | openstack-collector | Ingestion Health: Collector buffer depth, Oldest buffered event age. |
The database exporter's openstack_* series | openstack-db-exporter, or the simulator on a dev cluster | Fleet Overview: Projects (OpenStack), Top 10 projects by instance count. Project Drilldown: Resources by type, Volume capacity, Quota usage. The project_id variable. |
The store's own up series | every scrape job | Ingestion Health: Scrape health. |
A deployment that runs notifications and reconciliation without the database exporter has every panel of the third row empty and an empty project_id list. On a dev cluster, fill the dashboards pushes those series.
Variables
Every dashboard carries platform and cloud. Both are read off tally_current_resources, cloud narrowed by whatever platform is set to, and both are multi-select with an All option whose value is .*.
project-drilldown.json carries two more. project_id is read off the tenant_id label of openstack_nova_limits_instances_used, which the exporter reports once per project, and it is read again on every change of the time range (refresh: 2). api_base is a textbox rather than a query, and the panel link of Recent lifecycle activity is built from it: it opens GET /api/v1/events for the selected project. Its default is https://api.tally.example.com, the base's placeholder hostname. A reader sets it to the hostname the overlay publishes the Reporting API under, with the port where the Gateway is not on 443; on the dev overlay that is https://api.tally.127-0-0-1.nip.io:8443.
See also
The metrics page states every tally_ series these panels read. The alert rules page states what fires on them.