Skip to content

Set up your local Tally ​

In this lesson you clone the repository, create a kind cluster with the Tally dev overlay on it, trust the certificate authority that cluster issues its certificates from, mint an API token, and make your first call against the Reporting API.

At the end you have a running Tally answering on https://api.tally.127-0-0-1.nip.io:8443, an admin token in a variable of your shell, and the CA certificate in tally-ca.crt at the repository root. That is the state the rest of this track starts from.

This lesson takes about 30 minutes, most of it make up moving images.

Before you start ​

  • Docker running, Docker Desktop on macOS or Docker Engine on Linux, given enough of the machine to run the whole stack on one node.
  • git, kind, kubectl, Go, jq and curl on the path, and docker with its compose plugin. No version is named here. The clone below carries make check-tools, which calls every one of them, prints what each answered and what the Docker engine was given, and ends on the ones that are missing or answering an error. The next section runs it, before make up: it touches no cluster, and a tool that is not there is cheaper to find there than half an hour into make up.
  • The Go on the path has to satisfy the go line of go.mod, which is what make check-tools compares it against. That Go downloads the toolchain the toolchain line names on its first call inside the clone, which is the make check-tools of the next section.
  • No kind cluster named tally on the machine. kind get clusters prints No kind clusters found. when there is none. If it prints tally, tear that cluster down with the make down of Tear down your local Tally first.
  • 13 GB of free disk for Docker, measured with docker system df across images, volumes and build cache. The eight images the stack runs come to 3.3 GB and are held twice from here on, once by Docker and once inside the node, which is what lets make up copy them onto the node instead of leaving it to fetch them.

Get the code ​

  1. Clone the repository and change into it:

    sh
    git clone https://github.com/B42Labs/tally.git
    cd tally
    text
    Cloning into 'tally'...
    remote: Enumerating objects: 3555, done.        
    remote: Counting objects: 100% (727/727), done.        
    remote: Compressing objects: 100% (407/407), done.        
    Receiving objects: 100% (3555/3555), 3.29 MiB | 119.00 KiB/s
    Receiving objects: 100% (3555/3555), 3.30 MiB | 111.00 KiB/s, done.
    Resolving deltas: 100% (1866/1866)
    Resolving deltas: 100% (1866/1866), done.

    The object counts and the transfer speed are your own. Every command from here on runs from this directory, the repository root.

  2. Check the tools of Before you start:

    sh
    make check-tools

    Every line has to read ok. A missing line names a tool that is not on the path and a broken one a tool that is there and answering an error, and both are to be settled before the next section: make up reaches the same tool minutes in and stops with a half-created cluster behind it. A warn line for the Docker engine costs time rather than the run.

    On a machine whose Go is older than the toolchain line of go.mod, the go probe is the first Go command inside the clone, and that Go downloads the toolchain there. go: downloading go1.27.1 (<os>/<arch>) then stands on a line of its own immediately above ok go go1.27.1, at or above the go 1.26.0 of go.mod, with your own platform in place of <os>/<arch>. The download is where that probe's time goes. A Go already at go1.27.1 prints no such line.

Create the cluster and bring the stack up ​

  1. Bring the stack up:

    sh
    make up

    The target creates the kind cluster tally and says so with ==> creating kind cluster tally, installs cert-manager v1.21.1 and Envoy Gateway v1.8.3 and waits for their rollouts, applies the dev certificate authority, builds the four images (tally-reporting, tally-engine, tally-openstack-collector and tally-openstack-simulator), puts those and the eight images the stack runs on the node, applies the dev overlay, and applies the two migration chains, the reporting one and the engine one. Several hundred lines go by. These are the last ones:

    text
    Stack is up:
      https://api.tally.127-0-0-1.nip.io:8443/api/v1        Reporting API
      https://vm.tally.127-0-0-1.nip.io:8443                VictoriaMetrics
      https://grafana.tally.127-0-0-1.nip.io:8443           Grafana
      https://vmalert.tally.127-0-0-1.nip.io:8443/vmalert/  vmalert
      https://alertmanager.tally.127-0-0-1.nip.io:8443      Alertmanager
      https://otlp.tally.127-0-0-1.nip.io:8443              OTLP/HTTP
      db.tally.127-0-0-1.nip.io:5432                        TimescaleDB
    
    Trust the dev CA with: make -s ca > tally-ca.crt

    The seven URLs and the last line have to match exactly. Everything above them differs in timing and in the ids Kubernetes hands out.

  2. Leave it alone while it moves the images. A new node carries none of them, so the step before the overlay puts all twelve there: the four built above, and the eight the stack runs. An image your Docker already holds is copied straight onto the node, and only an image it lacks is fetched from the network, which is what the ==> pulling lines are:

    text
    ==> pulling busybox:1.37
    ==> loading busybox:1.37 onto the node

    Which of the eight are fetched is your machine's, and a second make up moves none of them: an image the node already carries is skipped.

  3. Let it wait if a readiness wait expires. One is given five minutes, and an expired one is repeated rather than ending the run:

    text
    ==> TimescaleDB is not ready after 300s; the node may still be pulling an image, waiting again (2/6)

    An error: timed out waiting for the condition above such a line is that wait expiring, not a fault. Six waits is the budget one rollout gets, half an hour; make up WAIT_ATTEMPTS=12 doubles it, and WAIT_TIMEOUT changes how long one wait lasts. cert-manager and Envoy Gateway install from their own manifests, so the node still fetches their images itself, and those are the waits most likely to repeat.

  4. Read the error if make up stopped instead of printing that block. A rollout that never comes up spends the budget and ends the run:

    text
    ERROR: TimescaleDB did not become ready in 6 waits of 300s.
           make up stops here, so the stack is incomplete. Stopping before
           the migration chain leaves the Reporting API at 0/1 until a later
           make up applies it: it never migrates on its own.
           kubectl --context kind-tally get pods -A, and the events of
           the pod that is not ready, say why it is not.
           make up is safe to run again: it reuses the cluster and carries on.
    make: *** [up] Error 1
  5. Wait until every pod but reporting-api reads Running before the next call:

    sh
    kubectl --context kind-tally -n tally get pods
    text
    NAME                             READY   STATUS    RESTARTS   AGE
    alertmanager-0                   1/1     Running   0          37m
    grafana-6c58589c8b-fg4tm         2/2     Running   0          37m
    otel-collector-599c64db4-vbh2z   1/1     Running   0          37m
    reporting-api-6b8ffd9598-smsw8   0/1     Running   0          37m
    timescaledb-0                    1/1     Running   0          37m
    victoriametrics-0                1/1     Running   0          37m
    vmalert-7699459f4-gf2nc          1/1     Running   0          37m

    The names carry ids of their own and the ages are your machine's. reporting-api reads 0/1 here, and waiting does not change it: a make up that stopped on a rollout never reached its migration step, and the API holds itself unready while its database carries no schema. Its log says so once per readiness probe:

    sh
    kubectl --context kind-tally -n tally logs deployment/reporting-api | grep readiness | tail -1
    text
    {"time":"2026-09-07T18:57:55.783841918Z","level":"WARN","msg":"readiness probe could not use the database","service":"tally-reporting","request_id":"4bc504a2-732d-4902-8c80-d04cd5e5ed2e","error":"reading the schema version: ERROR: relation \"goose_db_version\" does not exist (SQLSTATE 42P01)"}

    goose_db_version is the table a migration chain records itself in, and make up applies the two chains after the rollouts it waited on. So run make up again once the other pods read Running: against a cluster that already exists it reuses it, applies the overlay again, applies both chains and prints the block above. It moves no image it moved before, so a repeated call is short.

    ==> kind cluster tally already exists as the first line of a first make up means a cluster from an earlier run of this track is still on the machine. Tear it down with the make down of Tear down your local Tally and start over. On a repeated call after a timeout that line is the expected one.

  6. Read the pods once make up has printed its block:

    sh
    kubectl --context kind-tally -n tally get pods
    text
    NAME                             READY   STATUS      RESTARTS   AGE
    alertmanager-0                   1/1     Running     0          41m
    grafana-6c58589c8b-fg4tm         2/2     Running     0          41m
    otel-collector-599c64db4-vbh2z   1/1     Running     0          41m
    reporting-api-6b8ffd9598-smsw8   1/1     Running     0          41m
    tally-engine-29813460-mhpnt      0/1     Completed   0          4m18s
    timescaledb-0                    1/1     Running     0          41m
    victoriametrics-0                1/1     Running     0          41m
    vmalert-7699459f4-gf2nc          1/1     Running     0          41m

    reporting-api reads 1/1 now: the chain that call applied gave the readiness probe the schema it asks for. Every other long-running pod reads Running. The tally-engine-<id> pod is a Job of the hourly scheduler and appears once the clock has passed an hour mark. Its first tick moves the month that has ended into its grace window and reads Completed; a later tick meters that month, finds no pricing model and reads Error, which is expected until lesson 3 imports one.

Trust the dev CA ​

  1. Write the cluster's CA certificate to a file:

    sh
    make -s ca > tally-ca.crt

    The target prints nothing. tally-ca.crt now holds the certificate of the certificate authority cert-manager created for this cluster, 562 bytes on the run. It is not installed into the operating system or into a browser: cert-manager creates a new CA on every make up, so an installed one is stale after the next tear-down. curl is handed the file per call with --cacert instead.

  2. Call the Reporting API with the file and no credential:

    sh
    curl --cacert tally-ca.crt https://api.tally.127-0-0-1.nip.io:8443/api/v1/resources
    json
    {"type":"urn:tally:error:unauthorized","title":"Unauthorized","status":401,"detail":"the request carries no bearer token"}

    The API answered and refused, which is what a verified connection without a credential looks like. type urn:tally:error:unauthorized and status 401 have to match. The answer is an RFC 9457 problem document, the one shape every error of this API has, described under errors.

  3. Make the same call without the file:

    sh
    curl https://api.tally.127-0-0-1.nip.io:8443/api/v1/resources
    text
    curl: (60) SSL certificate problem: unable to get local issuer certificate
    More details here: https://curl.se/docs/sslcerts.html
    
    curl failed to verify the legitimacy of the server and therefore could not
    establish a secure connection to it. To learn more about this situation and
    how to fix it, please visit the web page mentioned above.

    This is what the file is for. The (60) line has to match, and a curl that is not the one macOS ships may word the lines after it differently.

Mint an API token ​

  1. Point the admin CLI at the dev reporting database:

    sh
    export TALLY_REPORTING_DB_URL='postgres://tally:tally-dev-password@db.tally.127-0-0-1.nip.io:5432/tally_reporting?sslmode=disable'

    The command prints nothing. The URL reaches the dev database through the Gateway's TCP listener, the path a psql on your machine takes as well, and tally-reporting-admin reads it from the environment.

  2. Mint the token and export it:

    sh
    TALLY_API_TOKEN="$(go run ./cmd/tally-reporting-admin create-api-token --role admin --description 'tutorial')"
    export TALLY_API_TOKEN
    text
    created api_tokens f2f68e0c-7228-4c75-a0da-8e344dabd56c
    the token above is printed this one time: store it now, it will not be shown again

    The two lines are the CLI's notices on stderr. The token itself went to stdout and from there into the variable, and it is printed this one time. The id is your own. go run builds the binary before it runs it, and make up already built this one for the migration chain, so the call is quick.

    The token carries the role admin, the role that every operation of this API accepts; the other two roles and how a token is revoked are in issue and revoke credentials.

    A dial tcp error naming db.tally.127-0-0-1.nip.io:5432 in place of the two notices means make up did not finish. Run it again.

Make your first call ​

  1. Ask the API for the resources it knows:

    sh
    curl --cacert tally-ca.crt -H "Authorization: Bearer $TALLY_API_TOKEN" 'https://api.tally.127-0-0-1.nip.io:8443/api/v1/resources'
    json
    {"items":[],"next_cursor":null}

    The fleet is empty because nothing has reported into it yet, which lesson 2 changes. items has to be empty and next_cursor null. A 401 here means the variable is empty in this shell: the token is printed once, so mint another one with the two commands above.

What you learned ​

  • The cluster runs the Reporting API, its database, the metrics store and the dashboards behind one Gateway, the arrangement the system at a glance draws.
  • tally-reporting-admin, the command that minted the token, is one of the six binaries, and you ran it with go run instead of from an image.
  • An API token carries a role, and the role decides what it may do: who may write what.
  • GET /api/v1/resources reads the projection, which stays empty until something reports: the projection.

Where to go next ​

Simulate a month of OpenStack fills the empty fleet you just read.

It starts from the state this lesson leaves behind: the kind cluster tally with the dev overlay running, tally-ca.crt at the repository root, and TALLY_REPORTING_DB_URL and TALLY_API_TOKEN in the shell. Keep that shell open. The token is printed once, and lesson 2 says how to mint another one if the shell was closed.