Deploy a highly available Telegraf Controller cluster
Deploy multiple Telegraf Controller nodes against a shared PostgreSQL database, put them behind a load balancer, and let one node lead while the others stand by. This guide covers configuring the nodes, routing traffic with the health endpoints, and tuning PostgreSQL for quick failover.
Available with Telegraf Enterprise
High availability is only available with Telegraf Enterprise. If you are interested in learning more about Telegraf Enterprise, contact us.
- Prerequisites
- Provision shared PostgreSQL and secrets
- Enable high availability on each node
- Apply the license
- Put the cluster behind a load balancer
- Tune PostgreSQL for fast failover
- Verify leadership and failover
Prerequisites
- A valid Telegraf Enterprise license. You apply the license once; every node reads it from the shared database. See Apply a license.
- A PostgreSQL database that every node can reach over the network. You can use self-managed PostgreSQL or a PostgreSQL-compatible managed service. High availability does not support SQLite. See Requirements and constraints.
- Two or more hosts to run the Telegraf Controller binary.
- A load balancer that can route traffic based on an HTTP health check.
Provision shared PostgreSQL and secrets
Every node connects to the same PostgreSQL database and signs sessions with the same secret.
Provision a PostgreSQL database and note its connection string. Connect nodes directly to PostgreSQL, or use a connection pooler in session-pooling mode. Transaction-pooling mode breaks leader election.
Generate a single
SESSION_SECRETto share across all nodes. Reuse the same value on every node so a session stays valid regardless of which node serves the request.openssl rand -hex 32
Keep DATABASE_URL and SESSION_SECRET identical on every node
A mismatched DATABASE_URL points a node at the wrong database, and a
mismatched SESSION_SECRET invalidates sessions when the load balancer routes
a user to a different node. The API, web interface, and heartbeat ports can
differ between nodes.
Enable high availability on each node
On every node, set HA_ENABLED=true, point DATABASE_URL at the shared
PostgreSQL database, and set the shared SESSION_SECRET. Start the binary the
same way you would a single node. You apply the license separately, once, in the
next step.
Add the high-availability variables to each node’s systemd unit file (typically
/etc/systemd/system/telegraf-controller.service):
[Service]
Environment=HA_ENABLED=true
Environment=DATABASE_URL=postgresql://POSTGRES_USER:POSTGRES_PASSWORD@POSTGRES_HOST:5432/telegraf_controller
Environment=SESSION_SECRET=SHARED_SECRETReload systemd and restart the service on each node:
sudo systemctl daemon-reload
sudo systemctl restart telegraf-controllerExport the high-availability variables, then start the binary on each node:
export HA_ENABLED=true
export DATABASE_URL="postgresql://POSTGRES_USER:POSTGRES_PASSWORD@POSTGRES_HOST:5432/telegraf_controller"
export SESSION_SECRET="SHARED_SECRET"
telegraf_controller --no-interactiveSet the high-availability variables, then start the binary on each node:
$env:HA_ENABLED="true"
$env:DATABASE_URL="postgresql://POSTGRES_USER:POSTGRES_PASSWORD@POSTGRES_HOST:5432/telegraf_controller"
$env:SESSION_SECRET="SHARED_SECRET"
./telegraf_controller.exe --no-interactiveReplace the following:
POSTGRES_USERandPOSTGRES_PASSWORD: the credentials for the shared PostgreSQL database.POSTGRES_HOST: the hostname or address of the shared PostgreSQL database, reachable from every node.SHARED_SECRET: the shared session secret generated withopenssl rand -hex 32, identical on every node.
Optionally, set
HA_POLL_INTERVAL_MS
to change how quickly settings, token, and license changes propagate between
nodes. It defaults to 5000 (5 seconds).
Apply the license
Apply your Telegraf Enterprise license to one of your Telegraf Controller nodes. Use either method:
- Set
LICENSE_FILE_PATHon one node at first startup to seed the license into the shared database. - Apply the license through the user interface or API after the cluster is running.
Every node reads the license from the shared database and becomes licensed within one poll interval, without a restart. For details, see Apply a license.
After a license is present, one node acquires the leader lock and the rest stand by. Confirm the cluster’s state with the health endpoints.
Put the cluster behind a load balancer
Run the cluster behind a load balancer that health-checks each node and routes around any node that fails. Route each class of traffic using the node’s unauthenticated health endpoints:
- Web interface and API traffic: route to nodes that return
200fromGET /health/readyon the API port (default8888). If you serve the web interface on a separateui-port, route that port as its own pool, health-checked withGET /. - Agent heartbeat traffic: route to nodes that return
200fromGET /healthon the heartbeat port (default8000).
Because sessions are validated with the shared SESSION_SECRET, sticky sessions
are not required.
For the full list of health endpoints and example configurations for HAProxy, NGINX, and cloud load balancers such as AWS Elastic Load Balancing, see Configure a load balancer.
Tune PostgreSQL for fast failover
When a leader shuts down gracefully, it releases the advisory lock and a standby takes over almost immediately, typically within a second or two. When a leader fails abruptly (a crash, a killed process, or a lost host), PostgreSQL must first notice that the leader’s connection is gone before another node can acquire the lock. With default operating-system keepalive settings, that can take much longer than a graceful handoff.
To bound failover time, shorten PostgreSQL’s TCP keepalive settings so the server detects dropped connections sooner:
ALTER SYSTEM SET tcp_keepalives_idle = 10;
ALTER SYSTEM SET tcp_keepalives_interval = 5;
ALTER SYSTEM SET tcp_keepalives_count = 3;
SELECT pg_reload_conf();With settings in this range, failover after an abrupt failure completes in well under a minute. Adjust the values to match your availability requirements and network conditions. Throughout either kind of failover, the surviving nodes keep ingesting agent heartbeats.
Verify leadership and failover
Identify the current leader by querying each node’s leader endpoint. Exactly one node returns
200:curl -s http://node-1.example.com:8888/health/leader curl -s http://node-2.example.com:8888/health/leaderConfirm that each licensed node reports ready:
curl -s http://node-1.example.com:8888/health/readyTest failover by stopping the leader (for example,
sudo systemctl stop telegraf-controlleron the leader host). Within a few seconds, another node’sGET /health/leaderreturns200, and the load balancer continues to serve web interface, API, and heartbeat traffic from the surviving nodes.
Was this page helpful?
Thank you for your feedback!
Support and feedback
Thank you for being part of our community! We welcome and encourage your feedback and bug reports for Telegraf and this documentation. To find support, use the following resources:
Customers with an annual or support contract can contact InfluxData Support.