🏗️infrastructure / nginx

Nginx Load Balancer Misconfiguration

junior20 min🏢 B2B SaaS (enterprise clients behind corporate NAT)
NginxLoad Balancingupstreamkeepaliveip_hash

Tech Stack

  • Nginx 1.24
  • Node.js 20
  • 5 backend nodes
  • Prometheus
  • Grafana
🚨

Incident Scenario

Your team deployed a new feature yesterday. Today the SRE reports one backend node is running at 94% CPU while the other four are nearly idle.
Response times from that one node are climbing. The load balancer is Nginx.
Traffic is 800 RPS total. No code changed since last week. What's wrong?

🔍 Investigation Artifacts

Reveal artifacts one by one. Each clue brings you closer to the root cause.

📊
Clue #1 · graph
Request Distribution Across Backend Nodes
💡 The distribution pattern tells you something about HOW the load balancer is routing requests
📋
Clue #2 · log
Nginx Access Log (last 50 requests)
💡 Look at the source IP addresses — what do they have in common?
⚙️
Clue #3 · config
Nginx Upstream Configuration
💡 What does ip_hash do when most clients share the same source IP?
📋
Clue #4 · log
Node Resource Metrics
💡 Does the resource usage match the traffic distribution?