Fixing AWS ELB Health Checks That Pass While Your App Serves Errors

October 10, 2026 6 min read

Your monitoring dashboard looks healthy.

Every target behind the AWS Elastic Load Balancer (ELB) is marked as Healthy.

CPU utilization is normal.

Memory usage is acceptable.

No instances have been removed from the target group.

Yet customers continue reporting:

  • HTTP 500 errors
  • Failed API requests
  • Timeouts
  • Empty responses
  • Broken application features

The load balancer insists everything is working.

Your users strongly disagree.

This situation is surprisingly common because an ELB health check only verifies what you ask it to verify. If the health endpoint checks little more than whether a web server responds with HTTP 200, it may completely miss failures affecting databases, caches, message queues, third-party APIs, or internal application logic.

This guide explains why AWS ELB health checks can produce false confidence and how to build health endpoints that accurately reflect your application's real operational state.


What You'll Learn

After reading this guide, you'll understand:

  • How ELB health checks work.
  • Why healthy targets may still serve failing requests.
  • Common causes of false-positive health checks.
  • How to design meaningful health endpoints.
  • Debugging techniques for production environments.
  • Best practices for reliable application monitoring.

Understanding ELB Health Checks

An Elastic Load Balancer periodically sends requests to each registered target.

Typical checks verify:

  • HTTP status code
  • HTTPS response
  • TCP connection
  • Response timeout
  • Success threshold
  • Failure threshold

If the configured conditions are met, the target remains in service.

However, passing a health check does not guarantee the application is functioning correctly.


The Health Endpoint Is Often Too Simple

Many applications expose a route such as:

/health

or

/ping

that simply returns:

HTTP/1.1 200 OK

While this confirms that the web server is running, it says nothing about whether the application can actually process user requests.


Your Application Depends on More Than HTTP

Modern applications rely on multiple services.

Examples include:

  • Databases
  • Redis
  • Object storage
  • Authentication providers
  • Search engines
  • Message queues
  • External APIs

If any critical dependency fails, users may experience errors even though the health endpoint continues returning HTTP 200.


Database Connectivity Issues

A common scenario is:

  • Web server is running.
  • Health endpoint returns 200.
  • Database connection pool is exhausted.

Every API request requiring database access now fails.

Unless the health endpoint verifies database connectivity, the load balancer has no reason to mark the target unhealthy.


Background Workers May Be Failing

Many SaaS applications use background jobs for:

  • Email delivery
  • Image processing
  • Report generation
  • Queue processing
  • Payment reconciliation

The web application may appear healthy while background workers have completely stopped processing jobs.

User-facing functionality gradually breaks despite healthy ELB targets.


Caching Problems

Redis or Memcached failures can dramatically affect application behavior.

Possible symptoms include:

  • Increased latency
  • Session failures
  • Authentication issues
  • API timeouts

Again, the load balancer cannot detect these issues unless they are reflected in the health check.


Third-Party Service Failures

Your application may depend on:

  • Payment gateways
  • Email providers
  • SMS services
  • Identity providers
  • External APIs

An outage affecting one of these services may render key application features unusable while your infrastructure itself remains operational.


Verify the Health Check Path

Ensure the configured path actually exercises meaningful application logic.

Avoid endpoints that simply return:

200 OK

without verifying the components required for normal request processing.


Different Types of Health Checks

Many production systems distinguish between:

Liveness Checks

Answers:

"Is the application process running?"


Readiness Checks

Answers:

"Can this instance safely receive traffic?"


Deep Health Checks

Answers questions such as:

  • Is the database reachable?
  • Is Redis responding?
  • Can required services be contacted?
  • Are migrations complete?
  • Are queues operational?

Separating these checks improves both reliability and troubleshooting.


Check Response Codes Carefully

Some frameworks return HTTP 200 even when internal errors occur.

Verify that failures correctly return:

  • HTTP 503
  • HTTP 500
  • Appropriate error responses

Otherwise, the load balancer continues routing traffic to unhealthy instances.


Review ELB Logs and Application Logs Together

Correlate:

  • Request timestamps
  • Target IP addresses
  • HTTP status codes
  • Response latency
  • Backend errors

Viewing infrastructure and application logs together often reveals patterns that neither source shows independently.


Watch for Partial Failures

Applications rarely fail completely.

Instead, only specific endpoints break.

Examples:

Working:

  • Homepage
  • Login page
  • Static assets

Failing:

  • Checkout
  • File uploads
  • Search
  • Reports

A health endpoint that exercises only one route cannot detect these partial failures.


Don't Ignore Connection Pools

Applications often fail because of:

  • Database pool exhaustion
  • Thread exhaustion
  • File descriptor limits
  • Memory pressure
  • Network saturation

The web server may still respond successfully to lightweight health requests while production traffic experiences widespread failures.


Consider Deployment Timing

Rolling deployments can create temporary inconsistencies.

Examples include:

  • Incompatible database schema
  • Incomplete migrations
  • Missing environment variables
  • Version mismatches
  • Stale configuration

Health checks should validate that initialization completed successfully before accepting traffic.


Real-World Example

An e-commerce platform runs behind an AWS Application Load Balancer with a health check configured to request /health. The endpoint simply returns HTTP 200 OK if the web server is running.

During a routine deployment, a database migration introduces an unexpected issue that prevents new database connections from being established. Existing health check requests continue succeeding because they never query the database, so every instance remains marked as healthy.

Meanwhile, customers attempting to browse products or place orders receive repeated HTTP 500 errors. After investigating logs and updating the health endpoint to verify database connectivity before returning a successful response, the load balancer automatically removes affected instances during future failures, preventing users from being routed to broken application servers.


Build Meaningful Health Endpoints

A production-ready health endpoint should verify only the dependencies required for the application to serve requests successfully.

Examples include:

  • Database connectivity
  • Cache availability
  • Critical configuration
  • Required secrets
  • Storage access
  • Internal service communication

Avoid checking optional services that should not remove an otherwise functional instance from rotation.


Avoid Expensive Health Checks

Health endpoints should remain lightweight.

Avoid:

  • Complex SQL queries
  • Large API calls
  • Heavy filesystem operations
  • Long-running computations

Health checks run frequently, so unnecessary work can create additional load.


Monitor User Experience Too

Infrastructure health alone is insufficient.

Track:

  • Error rate
  • Request latency
  • API success rate
  • Availability
  • User transactions
  • Synthetic monitoring

User-focused monitoring detects problems that infrastructure health checks may miss.


Best Practices Checklist

For reliable ELB health checks:

βœ… Use dedicated health endpoints

βœ… Verify critical dependencies

βœ… Separate liveness and readiness checks

βœ… Return accurate HTTP status codes

βœ… Keep health checks lightweight

βœ… Monitor application logs

βœ… Correlate ELB logs with backend logs

βœ… Validate deployments before accepting traffic

βœ… Test health checks under failure conditions

βœ… Monitor real user experience


Common Mistakes to Avoid

Avoid:

❌ Returning HTTP 200 unconditionally

❌ Checking only whether the web server is running

❌ Ignoring database connectivity

❌ Including expensive operations in health endpoints

❌ Treating liveness and readiness as identical

❌ Assuming healthy infrastructure means healthy applications

❌ Failing to test dependency failures


Healthy Infrastructure Doesn't Always Mean Healthy Software

A load balancer can only evaluate the signals it receives. If those signals measure process availability rather than application readiness, infrastructure dashboards may remain green while customers experience widespread failures.

Designing meaningful health checks helps bridge the gap between infrastructure health and actual service availability.


Build Health Checks That Reflect Reality

Effective health checks should answer the same question your users care about: Can the application successfully handle requests right now? By validating critical dependencies, using accurate status codes, and separating different types of health probes, you enable your load balancer to make better routing decisions and improve overall system resilience.

Health checks are not merely operational conveniencesβ€”they are an essential part of building reliable, highly available cloud applications.


Frequently Asked Questions (FAQ)

Why does my ELB report healthy targets while users receive 500 errors?

Most often, the configured health endpoint verifies only that the web server responds with a successful HTTP status. It may not check essential dependencies such as databases, caches, or external services that are required to process real user requests.

Should my health endpoint test every dependency?

No. It should validate only the dependencies required for the application to safely receive production traffic. Optional or non-critical services should generally not cause an instance to be marked unhealthy.

What's the difference between liveness and readiness checks?

A liveness check determines whether the application process is still running. A readiness check determines whether the application has completed initialization and can successfully serve incoming requests. Many production systems implement both.

Can expensive health checks cause problems?

Yes. Health endpoints are called frequently by load balancers. Performing heavy database queries, complex computations, or long-running network requests can increase system load and even contribute to performance issues.


Wrapping Summary

AWS ELB health checks are invaluable for maintaining high availability, but they are only as effective as the endpoints they monitor. A simplistic health check that always returns HTTP 200 OK may provide a false sense of confidence while critical application dependencies fail behind the scenes.

By designing meaningful health endpoints, distinguishing between liveness and readiness, validating essential services, and combining infrastructure monitoring with application-level observability, you can ensure that healthy targets truly represent healthy applicationsβ€”and deliver a more reliable experience for your users.

πŸ“€ Share this article

Sign in to save

Comments (0)

No comments yet. Be the first!

Leave a Comment

Sign in to comment with your profile.

πŸ“¬ Weekly Newsletter

Stay ahead of the curve

Get the best programming tutorials, data analytics tips, and tool reviews delivered to your inbox every week.

No spam. Unsubscribe anytime.