Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Nodes

The per-host schema, what each field decides, and how to add a node end to end. At the end of Adding a node the host is provisioned, on the mesh, and part of the cluster.

ansible/nodes/<hostname>/host.yml is the source of truth for a host, and carries everything reviewable in a diff. Its one identifying value, the address, lives in the nodes map of config/sops/ops.sops.yaml. ansible/inventory/host_vars/<hostname>/ holds a symlink to host.yml, which is how Ansible picks it up.

Do not confuse ansible/nodes/ with the repository-root nodes/. This one is provisioning data: identity, address, and how to reach and bootstrap the host. The other is workload definition, what runs once the host exists. See Node apps.

Schema

node:
  hostname: <hostname>
  os: fedora
  workflow: k8s
  k8s_role: controller
  mesh: true
  public_ingress: true
  ip: "{{ nodes[inventory_hostname].ip }}"
  initial_user: fedora
  initial_port: 22
FieldMeaning
workflowk8s, podman or none. Branches later setup steps; k8s nodes are the ones playbooks/k8s.yml installs k3s on, podman nodes get the standalone container plane
k8s_rolecontroller or worker. Only with workflow: k8s. A k3s server is a worker too unless tainted
meshOptional, default false. Joins the NetBird mesh
public_ingressOptional, default false. Opens 443 in firewalld and marks this host as the one PUBLIC_IP and MESH_IP in config/sops/cluster.sops.yaml describe
app_tierOptional, default false. This host carries apps under nodes/<hostname>.k8s/, so it gets an infisical-node-<hostname> tier
ipThe node’s public address. Also becomes its Kubernetes ExternalIP, via k3s’s node-external-ip
initial_user / initial_portFirst-contact login, the provider default, before the admin account exists

ip stays a reference, never a literal, because a real address is an identifying value and this repository is public. The nodes map comes from config/sops/ops.sops.yaml:

nodes:
  kenaz:
    ip: 203.0.113.10
    mesh_ip: ""

Keeping the reference inside the node: dict rather than moving the field out of it is deliberate: every hostvars[x].node.ip in the roles and the k3s config template keeps working unchanged, and Jinja resolves the indirection at use time. See Secrets.

Exactly one host should be a controller. A second controller makes etcd a two-member cluster with quorum two, which is worse for availability than a single controller, not better.

workflow decides two things beyond which playbook touches the host: which four roles playbooks/setup.yml runs under its podman tag, and which NetBird group the peer joins. ansible/roles/netbird derives that group name from this field, so a new workflow value needs a matching netbird_group in tofu/netbird/groups.tf before the first join, or the peer enrols into a group that does not exist and the setup key is rejected. podman is brokkr; see The standalone Podman plane.

mesh is orthogonal to workflow: opt in for any node, cloud or local, that needs mesh reachability. There is no mesh IP to store. Once joined, the node is addressed as <hostname>.<mesh_dns_domain>, and NetBird’s own resolver keeps that correct across re-keys. The node’s Kubernetes InternalIP comes from k8s_cluster’s node-ip, which is the address recorded in node.mesh_ip, not from that name.

public_ingress and mesh are both read generically. Nothing in roles/netbird, roles/firewall_ingress or roles/flux_bootstrap branches on a hostname, so on this plane moving public ingress is a one-line inventory change.

The other planes are not so lucky. public_ingress gates the firewall_ingress role and nothing else. No tofu module and nothing in the Flux tree reads it, so the edge node is named independently in five more places, with nothing detecting disagreement between them. Moving public ingress means changing all six together:

WhereWhat names the edge node
ansible/nodes/<host>/host.ymlpublic_ingress: true, which opens 443 and lowers the port floor
config/sops/cluster.sops.yamlPUBLIC_IP and MESH_IP must be that node’s addresses
tofu/bunny/refs.envTF_VAR_edge_public_ip reads ["nodes"][<host>]["ip"]
tofu/netbird/refs.envTF_VAR_mesh_ip reads ["nodes"][<host>]["mesh_ip"]
infra/traefik-edge/app/helmrelease.yamlnodeSelector, required: the hostIPs exist only on that node
infra/traefik-internal/app/helmrelease.yamlnodeSelector, required for the same reason

Miss a nodeSelector and that Traefik pod sits Pending. Miss PUBLIC_IP and the edge never binds. Miss tofu/bunny/refs.env and the public A record points at the wrong host; miss tofu/netbird/refs.env and every internal hostname does.

Adding a node

mkdir -p ansible/nodes/<hostname> ansible/inventory/host_vars/<hostname>
$EDITOR ansible/nodes/<hostname>/host.yml

# add a `<hostname>:` entry under `nodes`, with its `ip` and an empty `mesh_ip`
just ops sops config/sops/ops.sops.yaml

ln -s ../../../nodes/<hostname>/host.yml ansible/inventory/host_vars/<hostname>/host.yml
# then add `<hostname>: {}` under all.hosts in ansible/inventory/hosts.yml
just ans setup <hostname>
just ans k8s   # workflow: k8s only

just ans setup writes the node’s mesh address back into nodes.<hostname>.mesh_ip. Commit that change before running just ans k8s, which reads it.

Verify: just ans ping reaches the new host, ssh <hostname> netbird status reports connected, and for a workflow: k8s node, just ks nodes lists it Ready.

A workflow: podman node skips just ans k8s entirely, and just ans setup is the whole procedure: it converges the runtime, pushes the secrets, installs the git reconciler and starts the containers. Verify with ssh <hostname> podman ps instead. Its own prerequisites, the NetBird group and the ansible.secrets.<hostname> references, are in The standalone Podman plane.

If the node will run its own tenant apps under nodes/<hostname>.k8s/, it also needs its own Infisical operator tier, which is three files and one defaults entry, listed in Cluster infrastructure.