Run a simulated month against the dev cluster
This guide publishes one generated OpenStack month onto the broker the dev cluster's collector consumes from, paces it, stops it again, and reads the month back through the Reporting API. What such a month holds is in the simulated OpenStack world.
Before you start
- A dev cluster from
make up. The stack posts into its Reporting API, andmake simulator-uprebuilds nothing that runs in the cluster and applies no migration, so a change to the Reporting API or to the migrations takes anothermake upfirst. - Docker with
docker compose, for the three containers of the stack: the broker, the collector image, and the simulator. - A shell at the repository root, for the
maketargets and forgo run. - The month to render, as
SIM_PERIOD. It has to lie in the past:runrefuses a month that has not ended. - The simulator command line page, which states the routes of the control endpoint this guide calls.
- The simulator settings page, which lists every
TALLY_SIM_variablemake simulator-upwrites intodeploy/compose/.env.
Start the stack
Bring the dev cluster up, then start the simulator stack for the month:
shmake up make simulator-up SIM_PERIOD=2026-07SIM_CLOUDdefaults toos-sim,SIM_SEEDto 1, andSIM_FACTORto 744. A factor of 744 puts a 31-day month on the bus in an hour.SIM_FAULTSis empty, which is every fault switch off; it takes the switch names, comma-separated, as inSIM_FAULTS=held-back, which switch on fault switches covers.SIM_REGISTER_PROJECTSisfalse, andtrueregisters the month's tenants and Gardener projects with the dev registry before the first notification goes out, which register simulated projects covers; any other value is refused withERROR: SIM_REGISTER_PROJECTS must be true or falsebefore an image is built.SIM_GARDEN_CLOUDdefaults togarden-simand is the cloud the two Gardener rows are registered under.Take the seven URLs the stack prints:
http://127.0.0.1:15672, the broker's management UI, guest/guesthttp://127.0.0.1:8090/metrics, the collectorhttp://127.0.0.1:8091/clock, the simulator's control endpointhttp://127.0.0.1:8091/metrics, the simulator's inventory, the database exporter stand-in of the metric serieshttps://api.tally.127-0-0-1.nip.io:8443/api/v1, the Reporting APIhttps://otlp.tally.127-0-0-1.nip.io:8443/v1/metrics, the OTLP endpoint the series are pushed tohttps://vm.tally.127-0-0-1.nip.io:8443/targets, the scrape targets of the dev cluster, where theopenstack-db-exporterjob goes green while the run publishes
Watch the first publish. The simulator waits for a consumer on the collector's
tally-notificationsqueue before it, and--wait-for-collectorbounds that wait, two minutes by default, with0disabling it. The stack's collector runs withTALLY_OSC_REQUIRE_EXCHANGES=true, which is what makes a consumer on the queue mean a bound queue. A wait that runs out ends the run with an error naming the fix: start the collector first, or pass--wait-for-collector 0to publish anyway.
Pace the month
Finish the month at once instead of waiting the factor out. The route rebases the clock on the virtual instant it has reached and answers the clock document:
shcurl -X PUT -d '{"factor": 0}' http://127.0.0.1:8091/clockjson{"virtual_now":"2026-07-09T14:22:00Z","factor":0,"published":52,"total":15741,"held":0,"holding":false,"period_from":"2026-07-01T00:00:00Z","period_to":"2026-08-01T00:00:00Z"}Let the notifications a run with
SIM_FAULTS=held-backkeeps back out. The answer is the document as it stood the moment before the release, withheld0 andholdingfalse:shcurl -X POST http://127.0.0.1:8091/releasejson{"virtual_now":"2026-08-01T00:00:00Z","factor":744,"published":15657,"total":15741,"held":0,"holding":false,"period_from":"2026-07-01T00:00:00Z","period_to":"2026-08-01T00:00:00Z"}Build a backlog on the durable queue and drain it again. The messages are persistent, so a backlog survives a broker restart as well:
shdocker compose -f deploy/compose/compose.yaml stop collector docker compose -f deploy/compose/compose.yaml start collectorRerun a run whose publish the broker did not confirm, with the same seed, period, and cloud. Such a publish ends the run with exit status 1, a rerun renders the same message ids, and ingestion deduplicates whatever was already delivered. SIGINT and SIGTERM stop a run with exit status 0, and what went out stays out.
Stop the stack
Remove the containers, the outbox volume, and
deploy/compose/.env, so the nextsimulator-upstarts from a stack that carries nothing of the last one:shmake simulator-downIt resets the stack and not the dev reporting database: what a run already delivered stays ingested, and no subcommand of
tally-reporting-admindeletes an event. Running the same period again under another seed or another cloud adds a second, disjoint set of rows beside the first and inflates the usage the API reports.Start the period over by dropping the ingested data first, with
make down && make up, or with:shTALLY_REPORTING_DB_URL='postgres://tally:tally-dev-password@db.tally.127-0-0-1.nip.io:5432/tally_reporting?sslmode=disable' \ go run ./cmd/tally-reporting-admin migrate-down-to 0 make migrate
Read the month
Issue an admin token and list what the month booked under the cloud:
shtoken="$(TALLY_REPORTING_DB_URL='postgres://tally:tally-dev-password@db.tally.127-0-0-1.nip.io:5432/tally_reporting?sslmode=disable' \ go run ./cmd/tally-reporting-admin create-api-token --role admin \ --description 'openstack simulator')" curl --cacert tally-ca.crt -H "Authorization: Bearer $token" \ 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/resources?cloud=os-sim'A run started with
SIM_REGISTER_PROJECTS=truecarries a token already:deploy/compose/.envholds it asTALLY_SIM_API_TOKEN. Which tokens the admin CLI issues, and how one is ended, is in issue and revoke credentials.
Check the result
Ask the collector whether it holds its broker connection and its outbox:
shcurl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8090/readyztext200Read the counters on
http://127.0.0.1:8090/metrics:tally_collector_consumed_totalgrows per event type while the month goes out, and carries the threeoctavia.loadbalancer.*.endseries among the others.tally_collector_skipped_totalclimbs per type beside it, and it climbs faster: 13929 of the month's 15741 notifications are ones the mapping claims nothing for.- Neither counter carries an
event_type="other"series. The month's 87 types stay inside the bound of 100 label values the two of them share. tally_collector_unparseable_totalstays 0. Anything else is a rendered body the collector could not read.tally_collector_delivered_totalrises above 0 within two minutes at factor 744. A counter that stays at 0 means the events sit in the outbox and the Reporting API is not taking them.
Hold the totals against the seed. Seed 1 over
2026-07renders 15741 notifications, 1812 of them billable, and 87 distinctevent_typevalues. The shape of a month is the seed's alone, so those counts hold on every cloud. They are the counts of a run with every fault switch off.Read the collector's log. It shows neither an
x509error nor a401when the CA and the token are right:shdocker compose -f deploy/compose/compose.yaml logs collector