OpenStack collector (tally-openstack-collector)
tally-openstack-collector collects OpenStack usage events. It consumes oslo.messaging notifications from AMQP, maps them to Tally events, buffers them in a SQLite outbox, and posts them to the Reporting API from a loop of its own. Both loops retry what failed, so the process comes up while the broker or the Reporting API is unavailable and reports that state through the probes rather than through a failed start. A third loop logs a summary line every TALLY_OSC_SUMMARY_INTERVAL_S seconds. Beside them it serves the probes and the Prometheus exposition on the configured port; it serves no API of its own. The process is assembled in cmd/tally-openstack-collector/main.go.
Flags
| Flag | Type | Default | Description |
|---|---|---|---|
--dump | boolean | false | print the notifications the broker delivers as JSON lines instead of collecting them |
Modes
Collecting is the mode without a flag: the consumer, the sender, the summary loop, the outbox and the HTTP server all run, and the configuration gate asks for everything the pipeline reads.
--dump prints the notifications the broker delivers, one JSON line per delivery, and does nothing else: no HTTP, no outbox, no delivery. Its gate asks for TALLY_OSC_AMQP_URL alone. The mode is how the exchanges, the topics and the event types of a deployment are checked before the collector is pointed at it.
In what the dump prints, the value of every member whose name contains password, token, secret or connection_info, in any letter case and at any depth, is replaced by [redacted]. That holds for the payload of a parsed delivery and for the unparseable preview of a refused one, which is cut off after 512 bytes. A body that is not JSON, and an envelope whose oslo.message is not JSON, is reported by its size and not printed.
AMQP consumption
The queue, the consumer tag and the exchange kind are declared in internal/providers/openstack/osloamqp.go. The exchanges and the topics are configuration, and their defaults stand in internal/providers/openstack/config.go.
The collector's own queue is tally-notifications, declared durable, so notifications pile up in it while the collector is down and are consumed when it returns. It is bound to every topic the configuration names, on every configured exchange the broker carries. TALLY_OSC_EXCHANGES defaults to nova,neutron,openstack,glance and TALLY_OSC_TOPICS to notifications.info, which are the stock OpenStack settings: cinder sets no control_exchange and publishes on oslo's default, openstack. A topic is the routing key itself and not a prefix of one, because that is how oslo publishes.
TALLY_OSC_QUEUE_TYPE decides the arguments of that declare. The default is quorum, which sends x-queue-type=quorum and x-delivery-limit=-1, after the session has checked that the broker reports RabbitMQ 4.0 or newer. classic sends no arguments and so no queue type, which leaves the type to the broker: a virtual host whose default_queue_type is quorum creates a quorum queue with the broker's own delivery limit. A queue declared without arguments, by classic or by a collector up to v0.2.0, is refused under the default with a 406. That holds on a virtual host whose default_queue_type is quorum as well, because that queue carries no x-delivery-limit. Move the queue to another type has the order that drains the queue before the delete.
Each configured exchange is probed with a passive declare on a channel of its own, and none is created. An exchange the broker does not carry is skipped: the collector logs exchanges are missing on the broker, binding the others and retrying these with the missing ones, binds the others, and probes each missing one again after 1 s, doubling to 60 s. One that has appeared is bound and logged with an exchange appeared on the broker, bound it. What was published on an exchange before it was bound is not consumed.
A session fails over its exchanges, and the collector reconnects, in two cases:
- None of the configured exchanges exists. The error is
none of the exchanges in TALLY_OSC_EXCHANGES exists on the broker, followed by the list. TALLY_OSC_REQUIRE_EXCHANGESistrueand one exchange is missing. The error isthe exchange <name> does not exist on the broker, and TALLY_OSC_REQUIRE_EXCHANGES requires it, and the session ends before the queue is declared.
A session fails over its queue, and the collector reconnects, in two cases:
TALLY_OSC_QUEUE_TYPEisquorum, which it is by default, and the broker is older than RabbitMQ 4.0. The error isTALLY_OSC_QUEUE_TYPE=quorum needs RabbitMQ 4.0 or newer and the broker reports <version>: an older broker reads the delivery limit of -1 as a limit and drops a notification on its first requeue; set TALLY_OSC_QUEUE_TYPE=classic for this broker, and the session ends before the queue is declared. A broker whose version cannot be read is refused withTALLY_OSC_QUEUE_TYPE=quorum needs RabbitMQ 4.0 or newer and the broker reports no usable version: <value>; set TALLY_OSC_QUEUE_TYPE=classic for this broker.- The queue exists with another type than the setting declares. The error is
declaring the queue tally-notifications: the queue exists with other arguments than TALLY_OSC_QUEUE_TYPE=<type> declares, and a queue keeps the type it was declared with, followed by the broker's error. The collector does not delete the queue.
Octavia's control_exchange is octavia and the default leaves it out: a deployment that runs none would report it missing for as long as the collector runs. A deployment with octavia sets TALLY_OSC_EXCHANGES=nova,neutron,openstack,glance,octavia. Until it does, octavia's notifications reach no queue of this collector and show up in none of its counters.
The consumer registers under the tag tally-openstack-collector and acknowledges nothing automatically. A delivery is acknowledged after the outbox has committed the event it maps to, so a failed acknowledgement costs a redelivery and never an event. A delivery the outbox refuses is requeued instead, and stays on the broker until a buffer that works takes it.
Broker permissions
The collector issues four operations on the broker, in the virtual host the AMQP URL names:
| Operation | Resource | Permission |
|---|---|---|
passive exchange.declare | each exchange in TALLY_OSC_EXCHANGES | any one of configure, write or read on RabbitMQ 4.2.9, 4.3.1 and 4.3.6; configure on 4.3.0; none on 3.13.7, 4.0.9 and 4.1.8 |
queue.declare | tally-notifications, and in --dump mode a server-named amq.gen- queue | configure on the queue |
queue.bind | the queue and each exchange | write on the queue, read on the exchange, and a topic read pattern that matches the topic where topic permissions are set |
basic.consume | the queue | read on the queue |
The collector publishes nothing, declares no exchange and deletes nothing.
The releases named for the passive declare are the ones it was run against. Where a release checks it, an exchange outside the account's patterns is answered with a 403 whether it exists or not, such as ACCESS_REFUSED - configure access to exchange 'designate' in vhost '/' refused for user 'tally'. The collector reports that as an error and not as a missing exchange, so the session fails with declaring the exchange <name>: followed by the broker's error until the pattern lists the exchange. With the permission in place, a missing exchange is answered with the 404 that makes the collector skip it.
On RabbitMQ 4.3.0 read permission is not enough, so an account that holds read alone on the service exchanges fails every session there.
Bounds
Two sizes bound a notification on its way into the buffer, both stated as constants in internal/providers/openstack/osloamqp.go.
A delivery whose body is larger than 1 MiB is acknowledged unread and counted as unparseable. The bound comes before the parse: the parse is what an oversized body would take the process down in, and a process that dies there never acknowledges the delivery.
A notification whose mapped event is larger than 64 KiB is acknowledged and counted as skipped. An event that large is one the ingest endpoint refuses for its size, so buffering it would put it in front of every event behind it.
At the far end a 413 halves the batch until it fits, and the batch size then grows back gradually rather than jumping to TALLY_OSC_BATCH_MAX, so a Reporting API or an ingress with a smaller body limit settles on a size that fits. Only an event past the 64 KiB bound is dropped when it is refused alone. A smaller one is kept and retried like any other refusal: below that bound the 413 describes the destination and not the event.
Refused items
An item the API refuses does not fail its batch. The answer to POST /api/v1/events names each refused item with its index, its event id and the reason, the collector logs one warning per item, and the batch is deleted from the outbox with the rest of it. Refused items are not retried.
What the API keeps of a refused item depends on the reason. An item that failed validation is stored with the raw body it was submitted as and is readable through GET /api/v1/rejected-events. An item outside the credential's scope leaves an audit row of action events.scope_violation and no dead-letter row, so it does not show up in that view; the collector's log and the audit_log table are where it is found.
Mapping itself never fails. A notification whose payload the table did not understand still becomes an event, gets refused at ingestion, and lands in the dead-letter view with the reason it broke.
HTTP routes
Three routes are served on TALLY_OSC_HTTP_PORT, none of them with a credential. Each probe answers in plain text.
GET /readyz answers 200 with ok while the consumer holds a connection to the broker and the outbox answers. A skipped exchange does not fail readiness: a consumer bound to the exchanges the broker carries holds its connection. It answers 503 with the collector is not ready otherwise, which takes the pod out of the Service's endpoints while leaving it running. What failed goes to the log rather than into the body.
GET /healthz weighs the outbox alone, and only against time. It answers 200 with ok while the outbox answers, and it keeps answering 200 while the outbox has been unusable for fewer seconds than TALLY_OSC_UNHEALTHY_THRESHOLD_S. Past that threshold it answers 503 with the collector has been unhealthy for too long. The probe does not consult the broker.
Both probes bound their outbox check at 2 seconds, so a probe never outlasts the request that asked it.
GET /metrics serves the Prometheus exposition. A deployment that sets TALLY_METRICS_ENABLED to false is answered 404 there. The consumer and the sender keep counting either way.
Log lines
The collector logs JSON to stdout, one object per line, with time, level, msg and service, at the level TALLY_LOG_LEVEL names.
A healthy start logs listening with port, then the AMQP session is established, consuming with queue, exchanges and topics. exchanges lists the exchanges the queue was bound to in this session; a skipped exchange is absent and is reported by an exchange appeared on the broker, bound it once it is bound. The line is logged again after every reconnect.
summary is logged every TALLY_OSC_SUMMARY_INTERVAL_S seconds, the first one interval after the start, and whether or not anything happened. It carries these attributes:
| Attribute | Meaning |
|---|---|
interval_seconds | The configured interval. |
connected | Whether the consumer holds an AMQP session at the time of the line. |
consumed | Notifications mapped to an event and buffered. |
skipped | Notifications the mapping table produced no event for. |
unparseable | Deliveries whose body could not be parsed. |
delivered | Events the Reporting API accepted. |
delivery_errors | Delivery attempts the Reporting API did not accept. |
buffered | Events waiting in the outbox at the time of the line. |
oldest_buffered_seconds | Age of the oldest of them, 0 while the outbox is empty, and absent while the buffer cannot be read. |
The five counts, consumed to delivery_errors, cover the time since the previous summary line.
The line reads as follows:
connectedisfalse: the consumer holds no session, and the lastthe AMQP session ended, reconnectingwarning carries the reason.connectedistrueandconsumed,skippedandunparseableare 0: the consumer acknowledged no notification in the interval. Either none reached the queue, or they wait on the broker: behind a paused consumer, which has loggedthe outbox is at its bound, pausing the consumerand showsbufferedat 90 % ofTALLY_OSC_BUFFER_MAX_EVENTSor above, or behind an outbox that refuses them, which logsbuffering an event failed, requeueing the notification.skippedis above 0 andconsumedis 0: notifications arrive and none is of a mapped type.bufferedandoldest_buffered_secondsrise from line to line whiledeliveredstays 0: the Reporting API takes nothing, anddelivery_errorscounts the attempts.
TALLY_LOG_LEVEL=WARN hides both lines, and --dump logs neither.
Signals and exit status
SIGINT and SIGTERM begin a graceful shutdown. The HTTP server stops accepting connections, in-flight requests get 10 seconds, the consumer stops reading from the broker, the sender finishes the attempt it is in, the summary loop ends, and the outbox is closed once all three have returned. What is buffered stays in the file, where the next start picks it up, and the process exits 0.
Every other failure exits 1: a configuration that was refused, an outbox that could not be opened, and a listener that ended with anything but a closed server.
See also
The collector settings page lists every variable with its default. The notification mapping page states which event type maps to which Tally event, and the metrics page states the series the exposition carries.