Use another Gateway API implementation
This guide deploys the prod overlay to a cluster whose Gateway API implementation is not Envoy Gateway. The base is plain Gateway API. What needs Envoy Gateway is in the kustomize component deploy/kubernetes/components/envoy-gateway, and this guide takes that component out of the overlay.
The steps were checked with Traefik v3.7.13 from Helm chart 41.6.1, set with providers.kubernetesGateway.enabled=true and gateway.enabled=false, on kind v0.32.0 with the standard channel of Gateway API v1.6.2. The check covered the Gateway, the four HTTPRoutes and the GRPCRoute of the overlay: the Gateway is accepted and programmed, every route attaches, plain HTTP is redirected, and a hostname no route names is answered 404. It did not cover certificate issuance through the listener, OTLP traffic or a LoadBalancer.
Before you start
- A Gateway API implementation on the cluster, with a GatewayClass and the v1 kinds
Gateway,HTTPRouteandGRPCRoute. - cert-manager with its Gateway API support switched on, which is
config.gatewayAPI.enabled=truein its Helm chart. Install it after the Gateway API resource types exist, because it looks for them at startup. - Everything else Deploy the stack to a cluster asks for, apart from
helmformake prod-addons. - The sections "Cut the release that publishes the images", "Set the domain", "Name the cloud", "Configure reconciliation" and "Write the secrets" of that guide, done.
kubectl apply -kreads the.envfiles,collector.env,secrets/clouds.yamlandreconciliation/clouds-config.yaml.
Take the component out of the overlay
In
deploy/kubernetes/overlays/prod/kustomization.yaml, remove theenvoy-gatewayline of thecomponentsentry and keep the other three: the collector, the migration Job and the sync:yamlcomponents: - ../../components/openstack-collector - ../../components/migrations - ../../components/reconciliationIn the same file, remove the patch that deletes
HTTPRouteFilter alertmanager-deny-writes:yaml- patch: | apiVersion: gateway.envoyproxy.io/v1alpha1 kind: HTTPRouteFilter metadata: name: alertmanager-deny-writes $patch: deleteThe component declares that filter. With the patch left in,
kubectl kustomize deploy/kubernetes/overlays/prodstops:texterror: no resource matches strategic merge patch "HTTPRouteFilter.v1alpha1.gateway.envoyproxy.io/alertmanager-deny-writes.[noNs]": no matches for Id HTTPRouteFilter.v1alpha1.gateway.envoyproxy.io/alertmanager-deny-writes.[noNs]; failed to find unique target for patch HTTPRouteFilter.v1alpha1.gateway.envoyproxy.io/alertmanager-deny-writes.[noNs]
Name your class and ports
List the classes of the cluster:
shkubectl --context <ctx> get gatewayclassOn a cluster with Traefik:
textNAME CONTROLLER ACCEPTED AGE traefik traefik.io/gateway-controller True 12sAdd a replacement of
/spec/gatewayClassNameto the JSON patch onGateway tallyindeploy/kubernetes/overlays/prod/kustomization.yaml, with the name of your class:yaml- target: group: gateway.networking.k8s.io kind: Gateway name: tally patch: | - op: remove path: /spec/listeners/2 - op: replace path: /spec/gatewayClassName value: traefikWhere the implementation binds a listener to a port of its own, replace the ports of the
httpand thehttpslistener in the same patch. Traefik's chart serves its entry points on 8000 and 8443:yaml- op: replace path: /spec/listeners/0/port value: 8000 - op: replace path: /spec/listeners/1/port value: 8443With a port that matches no entry point, Traefik refuses the listener. The Gateway then reports
AcceptedasFalsewith the reasonPortUnavailableunderstatus.listeners, and this message:textCannot find entryPoint for Gateway: no matching entryPoint for port 80 and protocol "HTTP"
Replace what the component did
Limit both OTLP hostnames,
otlp.andotlp-grpc., to 10 requests a second per client address, with the means of your implementation. What is exposed and what is not says what the limit protects.Answer
/api/datasources/proxyon the Grafana hostname with a 403, with the means of your implementation. What the route publishes says what the refusal is for.
The component's other two objects need no replacement: the cluster declares the GatewayClass, and the overlay does not publish Alertmanager.
Nothing else on this page fails without the two steps. Checks 4 to 6 under Check the result are what shows that both are in place.
Deploy
make prod-addons and make prod-up are not used here. The first installs Envoy Gateway, and the second waits on the LoadBalancer Service Envoy Gateway creates.
Check that each of the six
.envfiles carries every key of its example, and that no value is empty or still a placeholder.make prod-uprefuses such a file, andkubectl apply -kapplies it: the pod that mounts a key the Secret lacks does not start, and with an emptyadmin-passwordGrafana keeps its default admin password. The loop prints the file and the key of every key a file lacks, and thegrepthose of every value to fill in:shfor example in deploy/kubernetes/overlays/prod/secrets/*.env.example; do for key in $(sed -n 's/^\([^#=][^=]*\)=.*/\1/p' "$example"); do grep -q "^$key=" "${example%.example}" || echo "${example%.example}:$key" done done grep -EH '^[^#=]+=([[:space:]]*$|.*<[a-z-]+>)' \ deploy/kubernetes/overlays/prod/secrets/*.env | cut -d= -f1Both print nothing when every key is there and every value is filled in. A line such as this one names a key to copy from the example, or a value to fill in, before the next step:
textdeploy/kubernetes/overlays/prod/secrets/tally-grafana.env:admin-passwordCheck that
collector.envnames the cloud.make prod-uprefuses an emptyTALLY_OSC_CLOUDand one with whitespace around it, andkubectl apply -kapplies both: the collector exits withchecking the configuration: TALLY_OSC_CLOUD: must be seton the first, and the ConfigMap keeps the whitespace of the second, so the sync asks for a cloud the clouds config does not name. The command counts the lines that set a value without whitespace around it:shgrep -Ec '^TALLY_OSC_CLOUD=[^[:space:]]+$' deploy/kubernetes/overlays/prod/collector.envtext1Check the input of reconciliation.
make prod-uprefuses each of the three, andkubectl apply -kapplies them: a placeholder fails every sync with a 500, a cloud other than the one ofcollector.envis answered 404, and an emptycloudends the Reporting API at startup. The first command counts the placeholders left insecrets/clouds.yaml, the second thecloudlines that name the cloud ofcollector.env, and the third theos_cloudlines that name an entry:shgrep -Ec '<[a-z-]+>' deploy/kubernetes/overlays/prod/secrets/clouds.yaml grep -cxF " - cloud: $(sed -n 's/^TALLY_OSC_CLOUD=//p' deploy/kubernetes/overlays/prod/collector.env)" \ deploy/kubernetes/overlays/prod/reconciliation/clouds-config.yaml grep -Ec '^ os_cloud: [^[:space:]]+$' deploy/kubernetes/overlays/prod/reconciliation/clouds-config.yamltext0 1 1Delete the migration Job of the previous deploy once it has finished. A Job's pod template is immutable, so the apply of a new tag fails on the old Job. The block is the one of deploy from a pipeline: it leaves a Job that is still running alone, because a migration that is killed can leave a half-built index, and stops on it:
sh( state="$(kubectl --context <ctx> -n tally get job tally-migrate --ignore-not-found \ -o jsonpath='{.metadata.name}:{.status.conditions[?(@.status=="True")].type}')" \ || { echo "cluster read failed; not deleting" >&2; exit 1; } case "$state" in ''|*Complete*|*Failed*) kubectl --context <ctx> -n tally delete job tally-migrate --ignore-not-found ;; *) echo "tally-migrate has not finished; wait for it before deploying" >&2 exit 1 ;; esac )Apply the certificate issuer and the overlay:
shkubectl --context <ctx> apply -f deploy/kubernetes/overlays/prod/issuers.yaml kubectl --context <ctx> apply -k deploy/kubernetes/overlays/prodWait for the migration Job, which applies both chains in the cluster, and then for the Reporting API:
shkubectl --context <ctx> -n tally wait --for=condition=complete job/tally-migrate --timeout=30m kubectl --context <ctx> -n tally rollout status deployment/reporting-apiThe Job reads nothing of the Gateway. When it fails,
kubectl --context <ctx> -n tally logs -l batch.kubernetes.io/job-name=tally-migrate --all-containers --prefix --tail=-1prints why.Issue the ingest credential into the Secret the collector waits for, as issue an ingest credential shows. The collector pod stays in
ContainerCreatinguntil then. The credential needs the migrated database of the step before, and nothing of the Gateway.Point the domain at the address your implementation publishes the Gateway on, and wait for the certificate, as point the domain at the cluster shows.
Check the result
The overlay renders no object of Envoy Gateway:
shkubectl kustomize deploy/kubernetes/overlays/prod | grep -c envoyproxytext0The Gateway is programmed. The
PROGRAMMEDcolumn readsTrue:shkubectl --context <ctx> -n tally get gateway tallyRun the checks of Deploy the stack to a cluster. Its last check runs
make prod-up, so leave that one out.Each OTLP hostname answers a burst from one address with 429. Each command sends 40 requests at once, without a credential:
shseq 40 | xargs -P 40 -I{} curl -s -o /dev/null -w '%{http_code}\n' -X POST \ https://otlp.tally.demo.b42labs.com/v1/metrics | sort | uniq -c seq 40 | xargs -P 40 -I{} curl -s -o /dev/null -w '%{http_code}\n' -X POST \ https://otlp-grpc.tally.demo.b42labs.com/ | sort | uniq -cThe dev stack, behind Envoy Gateway and the component's limit, prints:
text10 401 30 429 10 415 30 429401and415are the collector's own answers, to a request without a credential and to one that is not gRPC. An output without429means that hostname has no limit and every request reached the collector: with the limit taken away, the dev stack prints40 401and40 415.The limit counts per client address. The
429of the check above shows that a limit is there, and no more: one bucket for the whole hostname answers a burst from one address the same way, and so does a per-address limit behind a LoadBalancer that hides the client's address. Run this on one machine, then on two machines with different public addresses at the same time. It sends ten such bursts, one a second:shfor i in $(seq 10); do seq 40 | xargs -P 40 -I{} curl -s -o /dev/null -w '%{http_code}\n' -X POST \ https://otlp.tally.demo.b42labs.com/v1/metrics sleep 1 done | sort | uniq -cThe dev stack, asked from two addresses at once, prints for each what one address gets alone:
text100 401 300 429The counts are not exact. A bucket refills while a burst is still arriving, so a slower path gets more
401, and two runs from one address differ by a few: with each burst spread over 200 ms, the dev stack prints between 110 and 116. Judge by the sum of the two machines, because one of two that share a bucket can keep most of its count. Under a limit per client address the two get between them about twice what one gets alone. Machines that share a bucket get between them about what one gets alone, and any sender holds that bucket empty for every publisher without a credential. With the component's rule for distinct addresses taken away, the dev stack prints26 401for one address and82 401for the other, 108 between them and not 200. Repeat the check with theotlp-grpc.URL of check 4, where415takes the place of401.The Gateway refuses the datasource proxy before Grafana sees the request:
shcurl -sS -o /dev/null -w '%{http_code}\n' \ https://grafana.tally.demo.b42labs.com/api/datasources/proxy/uid/x/text403401is Grafana's own answer to that request, and means the prefix is published.