Configure a load balancer for high availability

A high-availability cluster runs behind a load balancer that health-checks each node and routes around any node that fails. Telegraf Controller exposes unauthenticated health endpoints for this purpose. This page describes those endpoints and shows worked configurations for common load balancers.

Available with Telegraf Enterprise

High availability is only available with Telegraf Enterprise. If you are interested in learning more about Telegraf Enterprise, contact us.

Health endpoints

Each node exposes health endpoints designed as load-balancer probes. They are unauthenticated, so a load balancer can probe them without credentials. The /health/* endpoints listen on the API port (--port, default 8888). The Rust heartbeat server exposes a separate /health endpoint on the heartbeat port (--heartbeat-port, default 8000).

EndpointPortStatus codesBodyPurpose
GET /health/liveAPI port200{"status":"ok"}Process liveness. Does not check the database. Use as a liveness probe.
GET /health/readyAPI port200, 503{"ready":true|false}Node can serve web interface and API traffic. Use as the readiness probe for UI/API nodes.
GET /health/leaderAPI port200, 503{"leader":true|false}Returns 200 only on the current leader. Use to locate the leader.
GET /healthHeartbeat port200, 503Status textRouting hint for heartbeat traffic. Heartbeat ingestion is always accepted, even on 503.

When you serve the web interface on a separate port with ui-port, the /health/* endpoints stay on the API port. The web interface port serves only static files and has no health endpoint; health-check it with an HTTP GET /, which returns the web interface with 200.

Route traffic to healthy nodes

By default, a Telegraf Controller node serves the web interface and API together on the API port and accepts agent heartbeats on the heartbeat port. You can also serve the web interface on its own port with ui-port. Route each class of traffic to nodes that pass the matching health check:

TrafficPortHealth check
Web interface and API (combined)API port (--port, default 8888)GET /health/ready returns 200
Web interface (separate port)UI port (--ui-port)GET / returns 200
API (with a separate UI port)API port (--port, default 8888)GET /health/ready returns 200
Agent heartbeatHeartbeat port (--heartbeat-port, default 8000)GET /health returns 200

Every licensed node serves web interface and API traffic, so distribute requests across all healthy nodes. Sessions are validated with the shared SESSION_SECRET, so sticky sessions are not required. Every node ingests heartbeats even when the heartbeat GET /health returns 503, so that check is a routing hint rather than a gate.

Use GET /health/leader for diagnostics and for tools that need to reach the leader directly. Do not use it as the readiness probe for web interface and API traffic, because standby nodes serve that traffic too.

Serving the web interface on a separate port

When you serve the web interface on its own port, route it as a separate pool. Because browsers reach the API through the load balancer’s external address, also set public-api-url and public-ui-url so the web interface calls the correct external API URL and the API’s CORS checks allow the web interface origin. See Public URLs and CORS.

Example configurations

The following examples serve the web interface and API together on the API port and route agent heartbeat traffic separately, for a three-node cluster. Replace the node addresses, ports, and health-check thresholds with values that match your deployment. If you serve the web interface on its own port, add a third pool for the UI port, health-checked with GET /.

Open-source HAProxy performs active HTTP health checks and stops routing to a node that fails them. Configure one backend per traffic class:

frontend fe_web
    bind *:8888
    default_backend be_api

frontend fe_heartbeat
    bind *:8000
    default_backend be_heartbeat

backend be_api
    balance roundrobin
    option httpchk
    http-check send meth GET uri /health/ready
    http-check expect status 200
    server node1 
NODE_1_IP
:8888 check inter 3s fall 3 rise 2 server node2
NODE_2_IP
:8888 check inter 3s fall 3 rise 2 server node3
NODE_3_IP
:8888 check inter 3s fall 3 rise 2 backend be_heartbeat balance roundrobin option httpchk http-check send meth GET uri /health http-check expect status 200 server node1
NODE_1_IP
:8000 check inter 3s fall 3 rise 2 server node2
NODE_2_IP
:8000 check inter 3s fall 3 rise 2 server node3
NODE_3_IP
:8000 check inter 3s fall 3 rise 2

Replace NODE_1_IP, NODE_2_IP, and NODE_3_IP with the addresses of your Telegraf Controller nodes.

Open-source NGINX proxies traffic and passively removes a node after failed responses, but it does not actively probe the health endpoints. For active health checks against /health/ready, use NGINX Plus.

Open-source NGINX (passive checks):

upstream tc_api {
    server 
NODE_1_IP
:8888
max_fails=3 fail_timeout=10s;
server
NODE_2_IP
:8888
max_fails=3 fail_timeout=10s;
server
NODE_3_IP
:8888
max_fails=3 fail_timeout=10s;
} upstream tc_heartbeat { server
NODE_1_IP
:8000
max_fails=3 fail_timeout=10s;
server
NODE_2_IP
:8000
max_fails=3 fail_timeout=10s;
server
NODE_3_IP
:8000
max_fails=3 fail_timeout=10s;
} server { listen 8888; location / { proxy_pass http://tc_api; proxy_set_header Host $host; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; } } server { listen 8000; location / { proxy_pass http://tc_heartbeat; proxy_set_header Host $host; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; } }

NGINX Plus (active checks):

upstream tc_api {
    zone tc_api 64k;
    server 
NODE_1_IP
:8888
;
server
NODE_2_IP
:8888
;
server
NODE_3_IP
:8888
;
} upstream tc_heartbeat { zone tc_heartbeat 64k; server
NODE_1_IP
:8000
;
server
NODE_2_IP
:8000
;
server
NODE_3_IP
:8000
;
} match tc_ready { status 200; } match tc_health { status 200; } server { listen 8888; location / { proxy_pass http://tc_api; health_check uri=/health/ready match=tc_ready interval=3s fails=3 passes=2; } } server { listen 8000; location / { proxy_pass http://tc_heartbeat; health_check uri=/health match=tc_health interval=3s fails=3 passes=2; } }

Replace NODE_1_IP, NODE_2_IP, and NODE_3_IP with the addresses of your Telegraf Controller nodes.

On AWS, create a target group per traffic class and configure each target group’s health check to probe the node’s health endpoint. This applies to both Application Load Balancers (ALB) and Network Load Balancers (NLB).

Target groupTargets (port)Health check protocolHealth check pathSuccess codes
Web interface and APINodes on 8888HTTP/health/ready200
Agent heartbeatNodes on 8000HTTP/health200

Create a target group for each traffic class with the AWS CLI:

# Web interface and API target group
aws elbv2 create-target-group \
  --name telegraf-controller-api \
  --protocol HTTP \
  --port 8888 \
  --vpc-id 
VPC_ID
\
--health-check-protocol HTTP \ --health-check-path /health/ready \ --matcher HttpCode=200 # Agent heartbeat target group aws elbv2 create-target-group \ --name telegraf-controller-heartbeat \ --protocol HTTP \ --port 8000 \ --vpc-id
VPC_ID
\
--health-check-protocol HTTP \ --health-check-path /health \ --matcher HttpCode=200

Replace VPC_ID with the ID of the VPC that contains your nodes, then register each node with both target groups.

Other cloud load balancers follow the same pattern. Configure the backend service or health probe to send an HTTP GET to each node’s health endpoint and treat only 200 as healthy:

  • Web interface and API: HTTP health check on the API port (8888) at path /health/ready, expecting 200.
  • Agent heartbeat: HTTP health check on the heartbeat port (8000) at path /health, expecting 200.

For example, use a Google Cloud health check or an Azure Load Balancer health probe with these settings. Consult your provider’s documentation for the exact field names.


Was this page helpful?

Thank you for your feedback!