Connect the collector to an OpenStack cloud
This guide points one tally-openstack-collector at one OpenStack cloud. At the end the services publish the notifications the collector reads, its queue is bound to their exchanges, and the events of that cloud arrive at the Reporting API. What the collector guarantees between those two ends is in how the collector consumes a bus.
Before you start
- A broker the collector reaches over AMQP, and an administrator shell on it for
rabbitmqctl. The account the collector connects with is created in create the broker account. The default queue type needs RabbitMQ 4.0 or newer, and an older broker takesTALLY_OSC_QUEUE_TYPE=classic, which choose the queue type covers. - Access to the configuration files of nova, neutron, cinder, glance and octavia, and permission to restart them.
- A shell that may run
tally-reporting-adminagainst the reporting database, for the ingest credential this cloud reports under. Issue and revoke credentials has those steps. - The collector settings page, which lists every variable this guide sets with its default.
- A collector to point at the cloud. This guide exports the variables into a shell; install the collector from the Debian package puts the same settings into
/etc/default/tally-openstack-collectorand runs it as a systemd service instead. The third way is theopenstack-collectorkustomize component the prod overlay lists, which runs the collector in the cluster of the Reporting API; deploy the stack to a cluster sets it up.
Configure the OpenStack services
Set nova to the unversioned notification format and to notifying on
vm_statechanges, and set the notification driver on nova, neutron, cinder, glance and octavia. The first block belongs to nova alone, the second to all five:ini[DEFAULT] # nova only notification_format = unversioned notify_on_state_change = vm_state [oslo_messaging_notifications] # nova, neutron, cinder, glance, and octavia driver = messagingv2Octavia's
[controller_worker] event_notificationsdefaults toTrueand needs no change. Oslo's own default for the driver is the empty string, so octavia stays silent until the second block is set on it.Restart the services whose configuration changed, and check that they came back:
shsystemctl restart nova-api nova-compute neutron-server cinder-api glance-api octavia-api systemctl is-active nova-api nova-compute neutron-server cinder-api glance-api octavia-apitextactive active active active active active
Enable notifications on a Kolla deployment
Kolla-Ansible, and OSISM, which deploys through it, renders driver = noop into [oslo_messaging_notifications] of a service unless one of the service's notification topics is enabled, and it enables the notifications topic for Ceilometer alone. On such a deployment these steps take the place of the section above.
Enable the
notificationstopic of the five services in the Kolla configuration, which is/etc/kolla/globals.yml, orenvironments/kolla/configuration.ymlon OSISM:yamlnova_notification_topics: - name: notifications enabled: true neutron_notification_topics: - name: notifications enabled: true cinder_notification_topics: - name: notifications enabled: true glance_notification_topics: - name: notifications enabled: true octavia_notification_topics: - name: notifications enabled: trueEach variable replaces the role's whole list. A deployment that runs the Designate sink keeps the sink's topic as a second entry under nova and neutron:
yaml- name: "{{ designate_notifications_topic_name }}" enabled: "{{ designate_enable_notifications_sink | bool }}"Leave nova's notification settings alone where Ceilometer, Designate or the Infoblox IPAM agent is enabled.
notification_formatdefaults tounversioned, and Kolla setsnotify_on_state_change = vm_and_task_statefor those three. That value is a superset ofvm_state: it addscompute.instance.updatenotifications, which the collector counts as skipped. Where none of the three is enabled, set the option in a nova override,/etc/kolla/config/nova.conf, orenvironments/kolla/files/overlays/nova.confon OSISM:ini[notifications] notify_on_state_change = vm_stateReconfigure the five services:
shkolla-ansible reconfigure -i <inventory> --tags nova,cinder,neutron,glance,octaviaOn OSISM, once per service:
shosism apply -a reconfigure novaRead the rendered section back from one of the containers:
shdocker exec nova_api grep -A3 '^\[oslo_messaging_notifications\]' /etc/nova/nova.conf | grep -v '^transport_url'text[oslo_messaging_notifications] driver = messagingv2 topics = notificationsThe section also holds
transport_url, which carries the broker password and is filtered out for that reason.
Cap the broker's message size
Set RabbitMQ's
max_message_sizeinrabbitmq.confso thatTALLY_OSC_PREFETCHtimes that value fits the collector pod's memory limit. RabbitMQ's own default is 128 MiB. 4 MiB is far above any oslo notification and leaves 400 MiB resident at the default prefetch of 100:inimax_message_size = 4194304Restart the broker and read back what it runs with:
shrabbitmqctl eval 'application:get_env(rabbit, max_message_size).'text{ok,4194304}Where the deployment cannot change that bound, lower
TALLY_OSC_PREFETCHuntil the product fits the memory limit. At RabbitMQ's own default of 128 MiB, a prefetch of 3 holds 384 MiB:shexport TALLY_OSC_PREFETCH=3
Cap the default notification queues
oslo.messaging declares a queue of its own per notification topic and priority, and the collector reads none of them. The default queues oslo declares has the reasons. On a Kolla or OSISM deployment the broker runs in a container, and every rabbitmqctl call on this page runs as docker exec rabbitmq rabbitmqctl there.
List the default queues with their consumers:
shrabbitmqctl list_queues name consumers messages | grep -E '^notifications\.'textnotifications.info 0 18342 notifications.error 0 27A consumer count of 0 means nothing reads the queue at this moment. Before capping it, confirm that the cloud runs no service that consumes it, which Ceilometer's notification agent does: an agent that is stopped or restarting shows 0 as well, and the length cap drops the oldest messages of the backlog it would have read as soon as the policy is set. A queue a service consumes is that service's backlog and gets no cap.
Cap them with a policy that keeps 10000 messages per queue and drops a message after 600000 ms:
shrabbitmqctl set_policy -p / --apply-to queues notifications-cap '^notifications\.' '{"max-length":10000,"message-ttl":600000}'The pattern does not match
tally-notifications. Where a service consumes one of these queues, narrow the pattern to the others. RabbitMQ applies one policy per queue, so on a broker whose existing policy already matches these queues, add the two keys to that policy instead.Read the policy back:
shrabbitmqctl list_policies -p /textvhost name pattern apply-to definition priority / notifications-cap ^notifications\. queues {"max-length":10000,"message-ttl":600000} 0
Bind the exchanges and topics
Name the service exchanges in
TALLY_OSC_EXCHANGESand the notification topics inTALLY_OSC_TOPICS. An exchange is a service'scontrol_exchange, a topic one of itsnotification_topics. The defaultsnova,neutron,openstack,glanceandnotifications.infoare the stock settings: cinder sets nocontrol_exchangeand publishes on oslo's default,openstack. A deployment that runs octavia lists it as well, and one that renamed an exchange or publishes on a topic of its own lists its values instead:shexport TALLY_OSC_EXCHANGES=nova,neutron,openstack,glance,octavia export TALLY_OSC_TOPICS=notifications.infoCheck which of the exchanges you listed the broker carries. The collector creates none of them:
shrabbitmqctl list_exchanges name type | grep -E '^(nova|neutron|openstack|glance|octavia)[[:space:]]'textnova topic neutron topic openstack topic glance topic octavia topicAn exchange missing from that output is skipped with a warning and bound within a minute of appearing. On a fresh cloud glance's exchange appears with the first image notification. That notification, and any published before the collector binds, reaches no queue, so create and delete one image before the cloud goes into billing.
TALLY_OSC_REQUIRE_EXCHANGES=truemakes the collector wait for every exchange instead.
Create the broker account
The account holds three permission patterns and one topic permission per exchange. What the broker account may do has the reasons. The patterns are not enough on RabbitMQ 4.3.0, which broker permissions covers.
Create the account and grant it the three patterns, configure, write and read in that order, on the virtual host the services'
transport_urlnames, which is/on Kolla.add_userreads the password from standard input, which keeps it out of the shell history and the process list:shrabbitmqctl add_user tally < tally-broker-password rabbitmqctl set_permissions -p / tally \ '^(tally-notifications|amq\.gen-.*)$' \ '^(tally-notifications|amq\.gen-.*)$' \ '^(tally-notifications|amq\.gen-.*|nova|neutron|openstack|glance|octavia)$'The file holds the password on one line; remove it afterwards. In a container the first call is
docker exec -i rabbitmq rabbitmqctl add_user tally < tally-broker-password.The read pattern lists the exchanges in
TALLY_OSC_EXCHANGESand no others, the ones the broker does not carry yet included. A cloud without octavia leaves it out here and in step 2. Theamq\.gen-.*alternative is the server-named queue of--dump, and an account that never runs the dump can leave it out.Restrict what the account may bind, once per exchange in the read pattern. The topic read pattern has to match every topic in
TALLY_OSC_TOPICS:shfor exchange in nova neutron openstack glance octavia; do rabbitmqctl set_topic_permissions -p / tally "$exchange" '^$' '^notifications\.info$' doneAn exchange in the read pattern without a topic permission is readable under every routing key, RPC included. RabbitMQ 4.3.6 answers
Exchange glance does not existfor an exchange the broker does not carry, where 3.13.7 accepts the call. The account is unrestricted on that exchange from the moment it appears, so repeat the call as soon as it does, which for glance on a fresh cloud is after the first image notification. Then list what the account bound there, because a binding that predates the topic permission stays:shrabbitmqctl list_bindings source_name destination_name routing_key | grep -E '^glance[[:space:]]+(tally-notifications|amq\.gen-)'textglance tally-notifications notifications.infoA line with another routing key than a topic in
TALLY_OSC_TOPICSis a binding made while the exchange was unrestricted. It keeps copying messages until it is unbound or its queue is deleted.Read both back:
shrabbitmqctl list_user_permissions tallytextvhost configure write read / ^(tally-notifications|amq\.gen-.*)$ ^(tally-notifications|amq\.gen-.*)$ ^(tally-notifications|amq\.gen-.*|nova|neutron|openstack|glance|octavia)$shrabbitmqctl list_topic_permissions -p /textuser exchange write read tally glance ^$ ^notifications\.info$ tally neutron ^$ ^notifications\.info$ tally nova ^$ ^notifications\.info$ tally octavia ^$ ^notifications\.info$ tally openstack ^$ ^notifications\.info$Every exchange in the read pattern has a row. One without a row is unrestricted, and is the one step 2 has to be repeated for.
Choose the queue type
The collector declares tally-notifications as a quorum queue unless TALLY_OSC_QUEUE_TYPE is classic. Classic and quorum compares the two types.
Read the broker version:
shrabbitmqctl versiontext4.3.6On RabbitMQ 4.0 or newer, leave
TALLY_OSC_QUEUE_TYPEunset. On an older broker, or to keep the queue on one node, set it toclassic:shexport TALLY_OSC_QUEUE_TYPE=classicA collector at the default on an older broker logs
TALLY_OSC_QUEUE_TYPE=quorum needs RabbitMQ 4.0 or newer and the broker reports <version>, ending onset TALLY_OSC_QUEUE_TYPE=classic for this broker, and consumes nothing.Where a collector has run against this broker before, list the queue it left:
shrabbitmqctl list_queues name type arguments messages consumers | grep -E '^tally-notifications[[:space:]]'texttally-notifications classic [{"x-queue-type","classic"}] 0 1A collector up to v0.2.0, and one set to
classic, declared the queue without arguments. The broker refuses the default declare over that queue, and over a quorum queue whose arguments carry nox-delivery-limitof -1, which is what a virtual host withdefault_queue_typequorumcreated. The collector then logsthe queue exists with other arguments than TALLY_OSC_QUEUE_TYPE=quorum declaresand consumes nothing, while the queue stays bound and keeps filling.Either set
TALLY_OSC_QUEUE_TYPE=classicbefore the collector is upgraded, which keeps the queue, or move the queue with the steps of move the queue to another type.
Move the queue to another type
These steps cover the upgrade from a collector up to v0.2.0 to the default, and any later change of TALLY_OSC_QUEUE_TYPE.
Keep the collector that fits the existing queue running until the message count, the fourth column, is 0. That collector is the older release, or the new one with the old setting. Where the collector was already restarted with the new setting, set the value back and restart it first:
shrabbitmqctl list_queues name type arguments messages consumers | grep -E '^tally-notifications[[:space:]]'texttally-notifications classic [{"x-queue-type","classic"}] 0 1Stop the collector and delete the queue:
shrabbitmqctl delete_queue tally-notificationsStart the collector with the new setting. For the default that is the upgraded collector with
TALLY_OSC_QUEUE_TYPEunset.
Notifications published between the stop and the delete are discarded with the queue, and those published between the delete and the new binding reach no queue. Reconcile a cloud finds what those notifications reported.
The way back is the same sequence. A collector set to classic, and a downgrade to v0.2.0, declares the queue without arguments, and the broker refuses that declare over the quorum queue. Wait until the message count is 0 with the quorum collector still running, stop it, delete the queue, and start the other collector.
Configure the collector
Set
TALLY_OSC_CLOUDto the cloud the ingest credential was issued for, and hand the collector the token inTALLY_OSC_TOKENor in the fileTALLY_OSC_TOKEN_FILEnames, which is the path a Kubernetes Secret volume takes. The API refuses an event whose cloud lies outside the credential's scope with the reasonscope, and a refused item is never resent:shexport TALLY_OSC_CLOUD=os-prod-eu1 export TALLY_OSC_TOKEN_FILE=/run/secrets/tally/ingest-tokenPoint the sender at the Reporting API.
TALLY_OSC_REPORTING_URLhas to be absolute, carry a host and usehttps; a deployment whose link to Tally is trusted setsTALLY_OSC_REPORTING_INSECURE=truefor a plaintext one, and the collector refuses to start otherwise:shexport TALLY_OSC_REPORTING_URL=https://tally-reporting.internalPut
TALLY_OSC_BUFFER_PATHon a volume that outlives the pod, sized fromTALLY_OSC_BUFFER_MAX_EVENTS. A buffered event runs a few hundred bytes, so the default of a million events reaches roughly half a gigabyte:shexport TALLY_OSC_BUFFER_PATH=/var/lib/tally/outbox.db export TALLY_OSC_BUFFER_MAX_EVENTS=1000000Set
TALLY_OSC_HTTP_PORTto the port/healthz,/readyzand/metricsanswer on, then start the collector and keep its log:shexport TALLY_OSC_HTTP_PORT=8080 tally-openstack-collector 2>&1 | tee collector.logjson{"time":"2026-07-09T14:22:00.512Z","level":"INFO","msg":"listening","service":"tally-openstack-collector","port":8080} {"time":"2026-07-09T14:22:00.731Z","level":"INFO","msg":"the AMQP session is established, consuming","service":"tally-openstack-collector","queue":"tally-notifications","exchanges":["nova","neutron","openstack","glance"],"topics":["notifications.info"]}
Check the result
In a second shell, ask for readiness. It answers 200 while the consumer holds the broker connection and the outbox answers; the
tally-openstack-collectorreference page states all three routes:shcurl -sS -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8080/readyztext200Read the two counters twice, a minute apart, while the cloud is in use. Both rise:
shcurl -sS http://127.0.0.1:8080/metrics | grep -E '^tally_collector_(consumed|delivered)_total'texttally_collector_consumed_total{cloud="os-prod-eu1",event_type="compute.instance.create.end",platform="openstack"} 14 tally_collector_delivered_total{cloud="os-prod-eu1",platform="openstack"} 12Count the two failures that keep delivered events at zero. An
x509error is a Reporting API certificate the collector does not trust, and a 401 is an ingest token the API refused:shgrep -cE 'x509|answered 401' collector.logtext0Read the collector's queue back. The line below is what the default declares, and a queue under
classicprintsclassicin the type column:shrabbitmqctl list_queues name type arguments effective_policy_definition consumers | grep -E '^tally-notifications[[:space:]]'texttally-notifications quorum [{"x-queue-type","quorum"},{"x-delivery-limit",-1}] #{} 1A policy or operator policy that matches
tally-notificationsand setsdelivery-limitto a positive number overrides the argument. The effective policy definition, which is#{}where no policy matches, must carry nodelivery-limit; excludetally-notificationsfrom such a policy's pattern.Read the last summary line. The collector logs one every
TALLY_OSC_SUMMARY_INTERVAL_Sseconds, 60 by default, with what it consumed and delivered since the previous one; the log lines section states every attribute:shgrep '"msg":"summary"' collector.log | tail -1json{"time":"2026-07-09T14:23:00.514Z","level":"INFO","msg":"summary","service":"tally-openstack-collector","interval_seconds":60,"connected":true,"consumed":14,"skipped":37,"unparseable":0,"delivered":12,"delivery_errors":0,"buffered":2,"oldest_buffered_seconds":3}