Reconcile a cloud
This guide gives the Reporting API what it needs to reconcile one OpenStack cloud: a clouds.yaml it can read, an account that may list every project's resources, an entry naming the cloud, and the call that runs a sync. What a run observes and how it corrects the projection is in how reconciliation observes a cloud.
Before you start
- The Reporting API deployed and reachable, and the shared internal token it guards the internal routes with.
- A
clouds.yamlfor the cloud, with an admin-scoped account, in a Secret the Reporting API pod mounts. Asecure.yamlbeside it is merged over it, which is where an entry's password belongs when the clouds.yaml itself is not a secret. - The
openstackclient on your machine, reading the same file, for the account check. - The Reporting API settings page, which names
TALLY_REPORTING_CLOUDS_CONFIG(the clouds file the API reads at startup),TALLY_REPORTING_SYNC_ALLOW_AT,TALLY_REPORTING_SYNC_BUDGET_SandTALLY_REPORTING_SYNC_SETTLE_S.
Restrict the account to read requests
This section is optional. It replaces the account's password in the Secret with an application credential that every API answers for the GET requests of a sync alone (a credential that only reads). It needs an entry on your machine that authenticates the same account with its password, called os-prod-eu1-password in the steps.
Set
service_typeunder[keystone_authtoken]in the configuration of each API the sync reads to that API's type in the catalog. List the types:shopenstack --os-cloud os-prod-eu1-password catalog list -c Name -c Typetext+----------+---------------+ | Name | Type | +----------+---------------+ | nova | compute | | cinder | block-storage | | neutron | network | | glance | image | | octavia | load-balancer | | keystone | identity | +----------+---------------+On a Kolla deployment cinder is the one API to change. Add the setting to
/etc/kolla/config/cinder.confand reconfigure cinder:ini[keystone_authtoken] service_type = block-storageshkolla-ansible reconfigure -i multinode --tags cinderCreate the credential as the account itself:
shopenstack --os-cloud os-prod-eu1-password application credential create \ --role admin \ --access-rules '[ {"service": "compute", "method": "GET", "path": "/**/servers/detail"}, {"service": "compute", "method": "GET", "path": "/**/flavors/detail"}, {"service": "compute", "method": "GET", "path": "/**/"}, {"service": "block-storage", "method": "GET", "path": "/**/volumes/detail"}, {"service": "block-storage", "method": "GET", "path": "/**/types"}, {"service": "network", "method": "GET", "path": "/**/floatingips"}, {"service": "image", "method": "GET", "path": "/**/images"}, {"service": "load-balancer", "method": "GET", "path": "/**/lbaas/loadbalancers"} ]' \ -f yaml -c id -c secret \ tally-reconciliationtextid: 6b1f0c2d3e4f5a6b7c8d9e0f1a2b3c4d secret: <the secret, shown this once>The
serviceof a rule is theservice_typeof step 1, not the name a client calls the API by. Theload-balancerrule belongs to a cloud whose entry setsinclude_octavia; leave it out otherwise.include_octaviais also what gives a load balancer its listener and pool counts, which octavia's notifications leave out. Keystone shows the secret once.Keystone cannot change the rules of an existing credential. A credential created with seven rules, without the
typesone, ends every runfailedon the volume type listing: create a new one with all eight and replace the entry with it in step 3.Replace the entry the Secret carries.
clouds.yaml:yamlclouds: os-prod-eu1: auth_type: v3applicationcredential auth: auth_url: https://keystone.example.com:5000/v3 application_credential_id: 6b1f0c2d3e4f5a6b7c8d9e0f1a2b3c4d region_name: RegionOne interface: publicsecure.yaml:yamlclouds: os-prod-eu1: auth: application_credential_secret: <the secret>The entry names no user, no password and no project, because the credential carries its project. The
os-prod-eu1-passwordentry stays on your machine and never reaches the Secret.
Continue with the sections below, using the new os-prod-eu1 entry.
Mount the clouds file
The prod overlay does this from secrets/clouds.yaml, as deploy the stack to a cluster shows, and step 2 checks it there as well.
Mount the Secret at
/etc/openstack/, the directory a Kubernetes Secret volume mounts at, and setOS_CLIENT_CONFIG_FILEto the file it carries. The variable makes that file the only location searched:yamlenv: - name: OS_CLIENT_CONFIG_FILE value: /etc/openstack/clouds.yamlCheck that the pod runs with the mount. The image carries the binary alone, so no
kubectl execcan list the directory, and the pod spec says what the container mounts:shkubectl get pod -l app.kubernetes.io/name=reporting-api \ -o jsonpath='{range .items[*]}{.status.phase}{" "}{.spec.containers[0].volumeMounts[*].mountPath}{"\n"}{end}'textRunning /etc/tally/reconciliation /etc/openstack /run/secrets/tally /run/secrets/tally-internal /var/run/secrets/kubernetes.io/serviceaccountThe kubelet starts the container only after it has projected every item its volumes name, so a
Runningpod that lists/etc/openstackhas the file. Whether the file authenticates is what check the account and the first sync show.
Check the account
Run the three requests the adapter depends on, against the same file it reads:
shopenstack --os-cloud os-prod-eu1 token issue openstack --os-cloud os-prod-eu1 server list --all-projects openstack --os-cloud os-prod-eu1 server list --all-projects --deletedtext+------------+----------------------------------+ | Field | Value | +------------+----------------------------------+ | expires | 2026-07-09T15:22:00+0000 | | project_id | 9c4a1b2d3e4f5061728394a5b6c7d8e9 | | user_id | 1f0e9d8c7b6a5948372615043f2e1d0c | +------------+----------------------------------+ +--------------------------------------+--------+--------+ | ID | Name | Status | +--------------------------------------+--------+--------+ | 1a2b3c4d-5e6f-4a7b-8c9d-0e1f2a3b4c5d | web-01 | ACTIVE | | 2b3c4d5e-6f7a-4b8c-9d0e-1f2a3b4c5d6e | db-01 | ACTIVE | +--------------------------------------+--------+--------+ +--------------------------------------+---------+---------+ | ID | Name | Status | +--------------------------------------+---------+---------+ | 3c4d5e6f-7a8b-4c9d-8e0f-2a3b4c5d6e7f | web-00 | DELETED | +--------------------------------------+---------+---------+The second call is the request every run probes the cloud with, and the third is the deleted-servers listing a missed delete is dated from. Both have to answer with more than the account's own project. A cloud that refuses the probe ends the run with an error naming the clouds.yaml entry, before a single listing follows.
With a credential restricted to read requests, a 401 on the second or third call comes from an API that cannot validate the credential's rules. That is repaired in step 1 of restrict the account to read requests.
Name the cloud
The prod overlay does this from reconciliation/clouds-config.yaml, as configure reconciliation shows.
Add one entry per cloud to the file
TALLY_REPORTING_CLOUDS_CONFIGnames:yamlclouds: - cloud: os-prod-eu1 platform: openstack adapter: openstack adapter_config: os_cloud: os-prod-eu1 include_octavia: truecloud,platform,adapterandadapter_configare the members of an entry, andos_cloudandinclude_octaviaare the two settings this adapter takes. Parsing is strict: a key the adapter does not know is refused with the key named.Restart the Reporting API. It reads its clouds once, at startup:
shkubectl rollout restart deployment/reporting-api kubectl rollout status deployment/reporting-apitextdeployment "reporting-api" successfully rolled out
Trigger a sync
Run one sync of one cloud:
shcurl -sS -X POST \ -H "Authorization: Bearer $TALLY_REPORTING_INTERNAL_TOKEN" \ https://tally-reporting.internal/internal/sync/os-prod-eu1json{"sync_run_id": "...", "stats": {"created": 3, "updated": 1, "deleted": 2}}POST /internal/sync/{cloud}is not part of the public API (the Reporting API reference). It is guarded by the shared internal token, the value ofTALLY_REPORTING_INTERNAL_TOKENor of the fileTALLY_REPORTING_INTERNAL_TOKEN_FILEnames, presented as a bearer token, and it takes no other credential. Whatever drives the sync schedule calls it, one call per configured cloud.In the prod overlay, CronJob
tally-syncof thereconciliationcomponent is that caller: it posts for the cloud ofcollector.envevery 10 minutes, and the answer is its Job's log. Run it once by hand with a Job of your own, and delete that Job afterwards:shkubectl create job --from=cronjob/tally-sync sync-now kubectl wait --for=condition=complete job/sync-now --timeout=2m kubectl logs job/sync-now kubectl delete job sync-nowA Job created while a scheduled run holds the cloud is answered 409 and fails, and the next scheduled run is not affected.
Tell a sync the instant it runs at (development only)
TALLY_REPORTING_SYNC_ALLOW_AT exists for a development deployment that reconciles a simulated cloud, whose clock is not the wall clock. A production deployment keeps the default, where such a request is refused before a run starts. With the setting on, every caller of the internal endpoint chooses the instant the run's sync_runs row starts at and the timestamp every correction the run books carries, and the endpoint takes the one shared static bearer token and nothing narrower — so the audit row of a back-dated run carries the instant its caller picked.
Set
TALLY_REPORTING_SYNC_ALLOW_ATon the deployment that is to take the instant. It isfalseby default:shkubectl set env deployment/reporting-api TALLY_REPORTING_SYNC_ALLOW_AT=truetextdeployment.apps/reporting-api env updatedName the instant the run happens at in a JSON body whose one optional member is an RFC 3339 timestamp:
shcurl -sS -X POST \ -H "Authorization: Bearer $TALLY_REPORTING_INTERNAL_TOKEN" \ -H 'Content-Type: application/json' \ -d '{"at": "2026-07-09T14:22:00Z"}' \ https://tally-reporting.internal/internal/sync/os-prod-eu1json{"sync_run_id": "...", "stats": {"created": 3, "updated": 1, "deleted": 2}}A call with no body, one whose body is empty, and one whose body carries no
atare the call of the section above: the run reads the wall clock.A deployment that sets nothing answers a body carrying
atwith 400. The refusal comes before the syncer runs, so the request leaves nosync_runsrow behind:shcurl -sS -X POST \ -H "Authorization: Bearer $TALLY_REPORTING_INTERNAL_TOKEN" \ -H 'Content-Type: application/json' \ -d '{"at": "2026-07-09T14:22:00Z"}' \ https://tally-reporting.internal/internal/sync/os-prod-eu1json{"type":"urn:tally:error:validation","title":"Validation failed","status":400,"detail":"this deployment does not take a sync instant; TALLY_REPORTING_SYNC_ALLOW_AT is off"}
Check the result
A run that finished is answered 200 with its id and the corrections it booked:
shcurl -sS -X POST \ -H "Authorization: Bearer $TALLY_REPORTING_INTERNAL_TOKEN" \ https://tally-reporting.internal/internal/sync/os-prod-eu1json{"sync_run_id": "6c1f0a94-3b52-4d7e-8f10-2a5b7c9d0e13", "stats": {"created": 3, "updated": 1, "deleted": 2}}A cloud the configuration does not name is answered 404. A cloud another run is holding is answered 409, and that lock lives in the database, so it holds across replicas. A run that recorded any error at all is answered 500, and its
sync_runsrow holds the reasons:shcurl -sS -o /dev/null -w '%{http_code}\n' -X POST \ -H "Authorization: Bearer $TALLY_REPORTING_INTERNAL_TOKEN" \ https://tally-reporting.internal/internal/sync/os-prod-eu2text404Read the last runs of the cloud back from the database:
sqlSELECT started_at, completed_at, status, stats FROM sync_runs WHERE cloud = 'os-prod-eu1' ORDER BY started_at DESC LIMIT 5;textstarted_at | completed_at | status | stats ----------------------------+----------------------------+-----------+-------------------------------------------------------------------------------------------------------- 2026-07-09 14:22:00.512+00 | 2026-07-09 14:22:07.118+00 | completed | {"errors": [], "created": 3, "deleted": 2, "updated": 1, "deferred": {"recent": 0, "transitional": 0}} 2026-07-09 13:22:00.401+00 | 2026-07-09 13:22:01.930+00 | failed | {"errors": ["unknown setting \"include_octavia_lb\""]}deferredcounts the corrections the run left to a later run:transitionalfor a resource the platform was still changing,recentfor one changed insideTALLY_REPORTING_SYNC_SETTLE_Sbefore the run. A count that stays above zero for the same cloud points at a resource stuck in transition.The Reporting API's log carries the same reasons on the request that triggered the run, together with the
sync_run_idthe row is found by.
Give a large cloud a longer budget
A run is bounded by TALLY_REPORTING_SYNC_BUDGET_S, 45 seconds by default. A cloud whose enumeration does not fit that budget completes no run until you raise it.
Read how long the runs take:
sqlSELECT started_at, completed_at, completed_at - started_at AS took, status, stats->'errors' FROM sync_runs WHERE cloud = 'os-prod-eu1' ORDER BY started_at DESC LIMIT 5;textstarted_at | completed_at | took | status | ?column? ----------------------------+----------------------------+--------------+--------+--------------------------------------------------------------------------------------------------------- 2026-07-09 14:22:00.512+00 | 2026-07-09 14:22:45.530+00 | 00:00:45.018 | failed | ["listing the resources of os-prod-eu1: the run ended while enumerating server: context deadline exceeded"]A run that ended on the budget is
failed, took about the budget, and carriescontext deadline exceededin its errors.Set the budget and restart:
shkubectl set env deployment/reporting-api TALLY_REPORTING_SYNC_BUDGET_S=300 kubectl rollout status deployment/reporting-apitextdeployment.apps/reporting-api env updated deployment "reporting-api" successfully rolled outThe value is in seconds, applies to every cloud of the deployment, and is refused at startup outside 1 to 86400.
Give whatever calls the route a client timeout of the budget plus 20 seconds,
--max-time 320for a budget of 300:shcurl -sS --max-time 320 -X POST \ -H "Authorization: Bearer $TALLY_REPORTING_INTERNAL_TOKEN" \ https://tally-reporting.internal/internal/sync/os-prod-eu1A client that gives up earlier closes the connection, which cancels the run and records it
failed. CronJobtally-synctakes its timeout fromSYNC_TIMEOUT_S, 65 seconds by default:shkubectl set env cronjob/tally-sync SYNC_TIMEOUT_S=320The next apply of the overlay sets the value back to 65, so a deployment that keeps the longer budget patches the CronJob in its overlay as well.
Keep the budget under the interval of the sync schedule. A call that arrives while a run holds the cloud is answered 409, and
TallySyncStalefires when no run completes within 30 minutes. CronJobtally-syncends a Job after 540 seconds, so a budget above 520 is cut off there.Size the connection pool for the clouds you sync at the same time. A running sync keeps one connection of
TALLY_REPORTING_DB_MAX_CONNSchecked out for as long as it runs and takes a second one for each write, so a longer budget holds that connection longer. RaiseTALLY_REPORTING_DB_MAX_CONNSby the number of clouds synced at the same time, or stagger the schedule so that their runs do not overlap.