Fixing AWS ELB Health Checks That Pass While Your App Serves Errors
Your monitoring dashboard looks healthy.
Every target behind the AWS Elastic Load Balancer (ELB) is marked as Healthy.
CPU utilization is normal.
Memory usage is acceptable.
No instances have been removed from the target group.
Yet customers continue reporting:
- HTTP 500 errors
- Failed API requests
- Timeouts
- Empty responses
- Broken application features
The load balancer insists everything is working.
Your users strongly disagree.
This situation is surprisingly common because an ELB health check only verifies what you ask it to verify. If the health endpoint checks little more than whether a web server responds with HTTP 200, it may completely miss failures affecting databases, caches, message queues, third-party APIs, or internal application logic.
This guide explains why AWS ELB health checks can produce false confidence and how to build health endpoints that accurately reflect your application's real operational state.
What You'll Learn
After reading this guide, you'll understand:
- How ELB health checks work.
- Why healthy targets may still serve failing requests.
- Common causes of false-positive health checks.
- How to design meaningful health endpoints.
- Debugging techniques for production environments.
- Best practices for reliable application monitoring.
Understanding ELB Health Checks
An Elastic Load Balancer periodically sends requests to each registered target.
Typical checks verify:
- HTTP status code
- HTTPS response
- TCP connection
- Response timeout
- Success threshold
- Failure threshold
If the configured conditions are met, the target remains in service.
However, passing a health check does not guarantee the application is functioning correctly.
The Health Endpoint Is Often Too Simple
Many applications expose a route such as:
/health
or
/ping
that simply returns:
HTTP/1.1 200 OK
While this confirms that the web server is running, it says nothing about whether the application can actually process user requests.
Your Application Depends on More Than HTTP
Modern applications rely on multiple services.
Examples include:
- Databases
- Redis
- Object storage
- Authentication providers
- Search engines
- Message queues
- External APIs
If any critical dependency fails, users may experience errors even though the health endpoint continues returning HTTP 200.
Database Connectivity Issues
A common scenario is:
- Web server is running.
- Health endpoint returns 200.
- Database connection pool is exhausted.
Every API request requiring database access now fails.
Unless the health endpoint verifies database connectivity, the load balancer has no reason to mark the target unhealthy.
Background Workers May Be Failing
Many SaaS applications use background jobs for:
- Email delivery
- Image processing
- Report generation
- Queue processing
- Payment reconciliation
The web application may appear healthy while background workers have completely stopped processing jobs.
User-facing functionality gradually breaks despite healthy ELB targets.
Caching Problems
Redis or Memcached failures can dramatically affect application behavior.
Possible symptoms include:
- Increased latency
- Session failures
- Authentication issues
- API timeouts
Again, the load balancer cannot detect these issues unless they are reflected in the health check.
Third-Party Service Failures
Your application may depend on:
- Payment gateways
- Email providers
- SMS services
- Identity providers
- External APIs
An outage affecting one of these services may render key application features unusable while your infrastructure itself remains operational.
Verify the Health Check Path
Ensure the configured path actually exercises meaningful application logic.
Avoid endpoints that simply return:
200 OK
without verifying the components required for normal request processing.
Different Types of Health Checks
Many production systems distinguish between:
Liveness Checks
Answers:
"Is the application process running?"
Readiness Checks
Answers:
"Can this instance safely receive traffic?"
Deep Health Checks
Answers questions such as:
- Is the database reachable?
- Is Redis responding?
- Can required services be contacted?
- Are migrations complete?
- Are queues operational?
Separating these checks improves both reliability and troubleshooting.
Check Response Codes Carefully
Some frameworks return HTTP 200 even when internal errors occur.
Verify that failures correctly return:
- HTTP 503
- HTTP 500
- Appropriate error responses
Otherwise, the load balancer continues routing traffic to unhealthy instances.
Review ELB Logs and Application Logs Together
Correlate:
- Request timestamps
- Target IP addresses
- HTTP status codes
- Response latency
- Backend errors
Viewing infrastructure and application logs together often reveals patterns that neither source shows independently.
Watch for Partial Failures
Applications rarely fail completely.
Instead, only specific endpoints break.
Examples:
Working:
- Homepage
- Login page
- Static assets
Failing:
- Checkout
- File uploads
- Search
- Reports
A health endpoint that exercises only one route cannot detect these partial failures.
Don't Ignore Connection Pools
Applications often fail because of:
- Database pool exhaustion
- Thread exhaustion
- File descriptor limits
- Memory pressure
- Network saturation
The web server may still respond successfully to lightweight health requests while production traffic experiences widespread failures.
Consider Deployment Timing
Rolling deployments can create temporary inconsistencies.
Examples include:
- Incompatible database schema
- Incomplete migrations
- Missing environment variables
- Version mismatches
- Stale configuration
Health checks should validate that initialization completed successfully before accepting traffic.
Real-World Example
An e-commerce platform runs behind an AWS Application Load Balancer with a health check configured to request /health. The endpoint simply returns HTTP 200 OK if the web server is running.
During a routine deployment, a database migration introduces an unexpected issue that prevents new database connections from being established. Existing health check requests continue succeeding because they never query the database, so every instance remains marked as healthy.
Meanwhile, customers attempting to browse products or place orders receive repeated HTTP 500 errors. After investigating logs and updating the health endpoint to verify database connectivity before returning a successful response, the load balancer automatically removes affected instances during future failures, preventing users from being routed to broken application servers.
Build Meaningful Health Endpoints
A production-ready health endpoint should verify only the dependencies required for the application to serve requests successfully.
Examples include:
- Database connectivity
- Cache availability
- Critical configuration
- Required secrets
- Storage access
- Internal service communication
Avoid checking optional services that should not remove an otherwise functional instance from rotation.
Avoid Expensive Health Checks
Health endpoints should remain lightweight.
Avoid:
- Complex SQL queries
- Large API calls
- Heavy filesystem operations
- Long-running computations
Health checks run frequently, so unnecessary work can create additional load.
Monitor User Experience Too
Infrastructure health alone is insufficient.
Track:
- Error rate
- Request latency
- API success rate
- Availability
- User transactions
- Synthetic monitoring
User-focused monitoring detects problems that infrastructure health checks may miss.
Best Practices Checklist
For reliable ELB health checks:
β Use dedicated health endpoints
β Verify critical dependencies
β Separate liveness and readiness checks
β Return accurate HTTP status codes
β Keep health checks lightweight
β Monitor application logs
β Correlate ELB logs with backend logs
β Validate deployments before accepting traffic
β Test health checks under failure conditions
β Monitor real user experience
Common Mistakes to Avoid
Avoid:
β Returning HTTP 200 unconditionally
β Checking only whether the web server is running
β Ignoring database connectivity
β Including expensive operations in health endpoints
β Treating liveness and readiness as identical
β Assuming healthy infrastructure means healthy applications
β Failing to test dependency failures
Healthy Infrastructure Doesn't Always Mean Healthy Software
A load balancer can only evaluate the signals it receives. If those signals measure process availability rather than application readiness, infrastructure dashboards may remain green while customers experience widespread failures.
Designing meaningful health checks helps bridge the gap between infrastructure health and actual service availability.
Build Health Checks That Reflect Reality
Effective health checks should answer the same question your users care about: Can the application successfully handle requests right now? By validating critical dependencies, using accurate status codes, and separating different types of health probes, you enable your load balancer to make better routing decisions and improve overall system resilience.
Health checks are not merely operational conveniencesβthey are an essential part of building reliable, highly available cloud applications.
Frequently Asked Questions (FAQ)
Why does my ELB report healthy targets while users receive 500 errors?
Most often, the configured health endpoint verifies only that the web server responds with a successful HTTP status. It may not check essential dependencies such as databases, caches, or external services that are required to process real user requests.
Should my health endpoint test every dependency?
No. It should validate only the dependencies required for the application to safely receive production traffic. Optional or non-critical services should generally not cause an instance to be marked unhealthy.
What's the difference between liveness and readiness checks?
A liveness check determines whether the application process is still running. A readiness check determines whether the application has completed initialization and can successfully serve incoming requests. Many production systems implement both.
Can expensive health checks cause problems?
Yes. Health endpoints are called frequently by load balancers. Performing heavy database queries, complex computations, or long-running network requests can increase system load and even contribute to performance issues.
Wrapping Summary
AWS ELB health checks are invaluable for maintaining high availability, but they are only as effective as the endpoints they monitor. A simplistic health check that always returns HTTP 200 OK may provide a false sense of confidence while critical application dependencies fail behind the scenes.
By designing meaningful health endpoints, distinguishing between liveness and readiness, validating essential services, and combining infrastructure monitoring with application-level observability, you can ensure that healthy targets truly represent healthy applicationsβand deliver a more reliable experience for your users.
π€ Share this article
Sign in to saveRelated Articles
Comments (0)
No comments yet. Be the first!