Simulate a month of OpenStack
In this lesson you start the simulator stack beside the cluster of lesson 1. It publishes one generated month of OpenStack notifications onto a broker the collector consumes from. You watch the month arrive in the Reporting API, finish it at once instead of waiting its pace out, and read it back.
At the end you hold the month of July 2026 of the simulated cloud os-sim in the reporting database, minus the share the simulator keeps back for a later track.
This lesson takes about 15 minutes.
Before you start
- The state lesson 1 leaves: the kind cluster from
make up,tally-ca.crtat the repository root,TALLY_REPORTING_DB_URLandTALLY_API_TOKENin the shell, and that shell at the repository root. - Docker running, with
docker composeon the path. The stack of this lesson runs beside the cluster, as three containers of its own.
If you closed that shell, restore it with this block:
export TALLY_REPORTING_DB_URL='postgres://tally:tally-dev-password@db.tally.127-0-0-1.nip.io:5432/tally_reporting?sslmode=disable'
TALLY_API_TOKEN="$(go run ./cmd/tally-reporting-admin create-api-token --role admin --description 'tutorial')"
export TALLY_API_TOKENA token is printed once, so a closed shell means a new token. The one lesson 1 minted stays valid until it is revoked the way issue and revoke credentials says. make simulator-up writes tally-ca.crt again if the file is missing.
Start the simulator stack
Start the stack for July 2026 with the held-back switch on:
shmake simulator-up SIM_PERIOD=2026-07 SIM_FAULTS=held-backThe month
2026-07is the one period this track uses everywhere. The cloudos-sim, the seed 1 and the pace, factor 744, are theMakefiledefaults, and run a simulated month covers the variables that change them.held-backis the one fault switch this track turns on, and the others are in switch on fault switches.The target first builds the collector and the simulator images, which is Docker build output, and prints
==> writing the dev CA to tally-ca.crt. The lines shown here start where it issues the credential:text==> issuing an ingest credential for os-sim created ingest_credentials 0ce1a769-1794-47e5-97c8-6241a087f49a the token above is printed this one time: store it now, it will not be shown again docker compose -f deploy/compose/compose.yaml up -dThe credential is the ingest token the collector reports under. It went into
deploy/compose/.envand is printed nowhere else. The id is your own.Compose then pulls the broker image and starts the three containers,
rabbitmq,collectorandsimulator, as a progress display the terminal redraws in place. These are the last lines:textSimulator stack is up: http://127.0.0.1:15672 broker UI, guest/guest http://127.0.0.1:8090/metrics collector http://127.0.0.1:8091/clock simulator control http://127.0.0.1:8091/metrics simulator inventory, the database exporter stand-in https://api.tally.127-0-0-1.nip.io:8443/api/v1 Reporting API https://otlp.tally.127-0-0-1.nip.io:8443/v1/metrics OTLP endpoint the series are pushed to https://vm.tally.127-0-0-1.nip.io:8443/targets scrape targets Finish the month at once: curl -X PUT -d '{"factor": 0}' http://127.0.0.1:8091/clock Release the notifications a run with SIM_FAULTS=held-back keeps back: curl -X POST http://127.0.0.1:8091/release Inspect the registry a run with SIM_REGISTER_PROJECTS=true registered into: curl --cacert tally-ca.crt \ -H "Authorization: Bearer $(grep TALLY_SIM_API_TOKEN deploy/compose/.env | cut -d= -f2)" \ 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/projects?cloud=os-sim' Reconcile the cloud the run serves: docs/how-to/simulator/reconcile-the-simulated-cloud.mdThe seven URLs have to match. The four hints below them are printed on every run. This track follows the first of them, the
PUT /clockof a later step, and never the second: thePOST /releasehint is what lets the held share out, which the billing track does.Three containers now run beside the cluster. The month of July 2026 of seed 1 on the cloud
os-sim, with its six tenants, goes onto the bus at factor 744, which puts the 31 days on the bus in an hour. Theheld-backswitch keeps 84 of the month's 1812 billable notifications off the bus, and the simulator holds them until a release.The collector and the simulator start together once the broker answers. The simulator declares the exchanges when it connects, and the collector consumes only once every one of them exists, so
docker compose -f deploy/compose/compose.yaml logs collectormay open with one or more WARN linesthe AMQP session ended, reconnectingwhoseerrorreadsthe exchange <name> does not exist on the broker, and TALLY_OSC_REQUIRE_EXCHANGES requires it, naming the first exchange ofTALLY_OSC_EXCHANGESthe simulator had not declared yet. The INFO linethe AMQP session is established, consumingfollows within seconds. That is the order the stack starts in and not a fault. A WARN line that keeps repeating a minute in is one, anddocker compose -f deploy/compose/compose.yaml logs simulatorthen says why the simulator stopped.ERROR: set SIM_PERIOD to the past month to simulate, e.g. make simulator-up SIM_PERIOD=2026-07in place of the build output means the period was left off the command. Adial tcperror from the admin CLI at the credential step means the cluster is not up, so run themake upof lesson 1 again.
Watch the clock
Read the simulator's clock:
shcurl -s http://127.0.0.1:8091/clockjson{"virtual_now":"2026-07-01T01:45:55Z","factor":744,"published":43,"total":15741,"held":84,"holding":false,"period_from":"2026-07-01T00:00:00Z","period_to":"2026-08-01T00:00:00Z"}A minute later the same call answers:
shcurl -s http://127.0.0.1:8091/clockjson{"virtual_now":"2026-07-01T14:10:23Z","factor":744,"published":649,"total":15741,"held":84,"holding":false,"period_from":"2026-07-01T00:00:00Z","period_to":"2026-08-01T00:00:00Z"}factor744,total15741,held84,period_from2026-07-01T00:00:00Zandperiod_to2026-08-01T00:00:00Zhave to match.virtual_nowandpublishedare your own and grow between the two reads, which lie twelve virtual hours apart.holdingis false while the month publishes. Connection refused on 8091 means the simulator is not running, anddocker compose -f deploy/compose/compose.yaml psshows the three containers whiledocker compose -f deploy/compose/compose.yaml logs simulatorsays why it stopped.
Watch the events arrive
Count what the Reporting API holds, grouped by cloud and resource type:
shcurl --cacert tally-ca.crt -H "Authorization: Bearer $TALLY_API_TOKEN" 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/stats/resources?group_by=cloud,resource_type'json{"items":[{"cloud":"os-sim","count":13,"resource_type":"floating_ip"},{"cloud":"os-sim","count":9,"resource_type":"image"},{"cloud":"os-sim","count":15,"resource_type":"instance"},{"cloud":"os-sim","count":2,"resource_type":"loadbalancer"},{"cloud":"os-sim","count":25,"resource_type":"volume"}]}A minute later:
shcurl --cacert tally-ca.crt -H "Authorization: Bearer $TALLY_API_TOKEN" 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/stats/resources?group_by=cloud,resource_type'json{"items":[{"cloud":"os-sim","count":13,"resource_type":"floating_ip"},{"cloud":"os-sim","count":9,"resource_type":"image"},{"cloud":"os-sim","count":20,"resource_type":"instance"},{"cloud":"os-sim","count":2,"resource_type":"loadbalancer"},{"cloud":"os-sim","count":29,"resource_type":"volume"}]}The counts are your own. They are the resources the collector has delivered at the moment of the call, and they grow while the month goes out. The five resource types are the month's. An empty
itemslist two minutes after the start means the collector has delivered nothing yet, which the counter read of the step "Wait for the collector to deliver" diagnoses.Read one row of the fleet:
shcurl --cacert tally-ca.crt -H "Authorization: Bearer $TALLY_API_TOKEN" 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/resources?cloud=os-sim&resource_type=instance&limit=1'json{"items":[{"cloud":"os-sim","created_at":"2026-07-01T03:17:02Z","deleted_at":null,"first_event_at":"2026-07-01T03:17:02Z","last_event_at":"2026-07-02T05:36:15Z","last_event_type":"compute.instance.power_off","last_payload":{"provider":{"oslo_event_type":"compute.instance.power_off.end"},"state":"shutoff"},"platform":"openstack","project_id":"d5a8024946ddf673277b9e2490643a2c","resource_id":"079faae9-9d39-426f-a963-769cb12aa629","resource_type":"instance","size":{"disk_gb":20,"flavor":"m1.small","ram_gb":2,"vcpus":1},"state":"shutoff"}],"next_cursor":"WyJvcy1zaW0iLCJpbnN0YW5jZSIsIjA3OWZhYWU5LTlkMzktNDI2Zi1hOTYzLTc2OWNiMTJhYTYyOSJd"}That is one row of the projection. Its
stateisshutoff, the state the last event left it in.project_idnames the tenant by id alone, the way the simulated cloud names its tenants, andsizecarriesvcpus,ram_gb,disk_gbandflavor.next_cursoris the handle for the next page. Which row comes first is your own while the month is still arriving.
Finish the month at once
Set the clock's factor to 0:
shcurl -X PUT -d '{"factor": 0}' http://127.0.0.1:8091/clockjson{"virtual_now":"2026-07-02T02:35:42Z","factor":0,"published":946,"total":15741,"held":84,"holding":false,"period_from":"2026-07-01T00:00:00Z","period_to":"2026-08-01T00:00:00Z"}The route rebases the clock on the virtual instant it has reached and answers the clock document.
factor0 has to match, andvirtual_nowandpublishedare your own. A 400 withfactor must be a JSON object with a number member "factor" that is zero or positivemeans the body was not the JSON object shown.Read the clock again once the rest of the month is out:
shcurl -s http://127.0.0.1:8091/clockOn the run the answer below came 10 seconds after the factor changed. A slower machine takes a few minutes.
json{"virtual_now":"2026-07-02T02:35:42Z","factor":0,"published":15657,"total":15741,"held":84,"holding":true,"period_from":"2026-07-01T00:00:00Z","period_to":"2026-08-01T00:00:00Z"}published15657,held84,holdingtrue andtotal15741 have to match.virtual_nowis your own: at factor 0 the clock no longer advances, so it stays at the instant the factor became 0 while the rest of the month goes out at once. The simulator now waits for a release the billing track sends, and this lesson sends none.
Wait for the collector to deliver
Sum the collector's counters over their
event_typelabels:shcurl -s http://127.0.0.1:8090/metrics | awk ' /^tally_collector_consumed_total/ { c += $2 } /^tally_collector_skipped_total/ { s += $2 } /^tally_collector_unparseable_total/ { u += $2 } /^tally_collector_buffer_depth/ { d = $2 } END { print "consumed", c, "skipped", s, "unparseable", u, "depth", d }'textconsumed 1728 skipped 13929 unparseable 0 depth 0Repeat the read until
depthreads 0, about three minutes after the month went out. The four numbers have to match. 1728 is the month's 1812 billable notifications minus the 84 held ones, 13929 is the notifications the mapping claims nothing for, and the depth is the outbox the collector drains into the Reporting API.Read what the Reporting API took:
shcurl -s http://127.0.0.1:8090/metrics | grep ^tally_collector_delivered_totaltexttally_collector_delivered_total{cloud="os-sim",platform="openstack"} 1728The 1728 delivered are the 1728 consumed. A depth that does not fall while a
tally_collector_delivered_totalstays at 0 means the Reporting API refuses the batches, anddocker compose -f deploy/compose/compose.yaml logs collectorshows the401or thex509line. The cure ismake simulator-downand thenmake simulator-up SIM_PERIOD=2026-07 SIM_FAULTS=held-backagain, which issues a fresh credential and writes the CA file again.
Read the month back
Count the whole month, the deleted resources included:
shcurl --cacert tally-ca.crt -H "Authorization: Bearer $TALLY_API_TOKEN" 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/stats/resources?group_by=cloud,resource_type&status=all'json{"items":[{"cloud":"os-sim","count":16,"resource_type":"floating_ip"},{"cloud":"os-sim","count":9,"resource_type":"image"},{"cloud":"os-sim","count":668,"resource_type":"instance"},{"cloud":"os-sim","count":5,"resource_type":"loadbalancer"},{"cloud":"os-sim","count":169,"resource_type":"volume"}]}Every count has to match. They are the seed's: 16 floating IPs, 9 images, 668 instances, 5 load balancers and 169 volumes.
status=allcounts the deleted resources too, which the defaultactiveleaves out.Read the first event the month stored:
shcurl --cacert tally-ca.crt -H "Authorization: Bearer $TALLY_API_TOKEN" 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/events?cloud=os-sim&limit=1'json{"items":[{"cloud":"os-sim","event_id":"7d16f9f3-740a-4670-9898-bc32b30fd3f9","event_type":"image.create","payload":{"provider":{"oslo_event_type":"image.upload"},"size":{"size_gb":3.5},"state":"active"},"platform":"openstack","project_id":"005be5adeef3d87e280d03d9d57c38b4","received_at":"2026-09-07T12:37:37.784207Z","resource_id":"d45cc80a-b7ea-4988-8873-16b54ec1a39d","resource_type":"image","source":"collector","timestamp":"2026-07-01T00:21:32Z"}],"next_cursor":"WyIyMDI2LTA3LTAxVDAwOjIxOjMyWiIsIjdkMTZmOWYzLTc0MGEtNDY3MC05ODk4LWJjMzJiMzBmZDNmOSJd"}It is an
image.createat2026-07-01T00:21:32Zwith itssize.event_id,timestamp,resource_id,project_idand the payload are the seed's and have to match.received_atis your own, the wall-clock instant the collector delivered the event.
See the metrics land in the store
Count the instances that pushed traffic during the month:
shcurl --cacert tally-ca.crt -G 'https://vm.tally.127-0-0-1.nip.io:8443/api/v1/query' --data-urlencode 'query=count(count_over_time(ceilometer_network_outgoing_bytes_total{cloud="os-sim"}[31d]))' --data-urlencode 'time=2026-08-01T00:00:00Z'json{"status":"success","data":{"resultType":"vector","result":[{"metric":{},"value":[1785542400,"667"]}]},"stats":{"seriesFetched": "667","executionTimeMsec":2}}667 has to match, the instances of the month that pushed traffic. A smaller count means the simulator is still pushing the month's series, which it goes on doing after the last notification is out; repeat the read until it stands at 667, which takes about another minute. The series carry July 2026 timestamps, which is why the query is asked at the end of the month with
timeand over a window that spans the month, and why lesson 4 sets the dashboards' time range to July.The plain instant query, without the window and without
time, is evaluated at the wall clock, where the month's series ended a month ago:shcurl --cacert tally-ca.crt -G 'https://vm.tally.127-0-0-1.nip.io:8443/api/v1/query' --data-urlencode 'query=count(ceilometer_network_outgoing_bytes_total{cloud="os-sim"})'json{"status":"success","data":{"resultType":"vector","result":[]},"stats":{"seriesFetched": "0","executionTimeMsec":0}}An empty answer within 30 seconds of the push is the store's latency offset instead, described under check the result of the dashboards guide.
What you learned
- The month holds three classic projects, two Gardener tenants, one CI tenant and one external network: the world and the workload.
- The
held-backswitch keeps one in 20 of the billable transitions off the bus until a release lets them out: the fault switches. - The collector consumes the broker, keeps what it read in an outbox and delivers it to the Reporting API in batches: at-least-once over an outbox.
- The stats and the resource row you read come from the projection the events fold into: the projection.
- The traffic and the inventory series reach the store over OTLP, timestamped in the simulated month: what is pushed.
Where to go next
Meter and rate your first month turns the month you just ingested into rated usage.
It starts from the state this lesson leaves behind: everything lesson 1 left, plus the compose stack running with the simulator holding 84 notifications, so GET /clock answers holding true, the collector idle with an empty outbox, and the reporting database holding the month of seed 1 on os-sim minus that held share. The project registry is empty. Nothing of this lesson registers the six tenants, and a lesson of the billing track does that.