Picture a partner API, say the one that prices your shipping, that starts accepting connections and never answering. Not down. Just silent. In our test, that silence took a Spring Boot app with default settings completely offline in 10 seconds, health check included, and the app stayed offline after the traffic stopped.
Our founder has had to fix thread exhaustion in a Java HTTP client before, and adding missing timeouts was part of that fix. For this post we built a test app from scratch and measured five setups side by side, including two that surprised us, plus a few follow-up tests.
- 10 s until the app stopped answering anything, at 20 requests a second
- 0/180 health checks answered after the traffic stopped
- 2.0 s to a fast error instead, with two lines of config
In short
- With the JDK’s HTTP client, which Spring Boot 4.1 uses unless another client is on the classpath,
RestClientgets no timeouts at all. A partner that accepts connections but never answers held every Tomcat worker thread: at 20 requests a second, the whole app stopped answering after 10 seconds and was still down 45 seconds after the traffic stopped. - Two properties fixed it:
spring.http.clients.connect-timeoutandspring.http.clients.read-timeout. Every call failed fast, in 2 seconds (a 503, with a small exception handler), and the app stayed up. Upgrading from Spring Boot 3.4 or 3.5? The oldspring.http.client.*names are silently ignored in 4.1. - Virtual threads kept the health check green, but all 600 partner calls were still hanging when the test ended: green health checks, stuck customers.
- With Apache HttpClient on the classpath, timeouts weren’t enough: its pool allows 5 connections per host and waits up to 3 minutes for one. Raise the pool and cap the wait.
The experiment: one app, one silent partner, five setups
A Spring Boot 4.1.1 app on Java 21 with two endpoints. /quote calls a partner API through a RestClient built from Spring Boot’s auto-configured RestClient.Builder. With no other HTTP client library on the classpath, Spring Boot builds it on the JDK’s own HttpClient and sets neither a connect nor a read timeout. /ping returns “ok” and touches nothing, so if it stops answering, the whole app is down. The partner was a stub that accepts the connection, reads the request and never replies: one of the nastiest ways an API can fail. A server that refuses the connection fails fast. A silent one drags you down with it.
We sent 20 requests a second to /quote for 30 seconds, 600 in all, and checked /ping every 250 ms for 75 seconds, counting it as down if it took longer than 2 seconds. Each test request opened its own connection to the app.
| Setup | Health check | Partner call |
|---|---|---|
| Spring Boot defaults | Down No answer from 10 s to the end of the run | Hung 600 of 600 unanswered at the end: 200 stuck on the partner, 400 queued for a thread |
| Timeouts: connect 1 s, read 2 s | Up 300 of 300 checks answered | Fast 503 600 of 600 in 2.0 s |
| Defaults + virtual threads | Up 300 of 300 checks answered | Hung 600 of 600 still waiting at the end |
| Apache HttpClient + the same timeouts | Down No answer from 11.5 s to the end of the run | Queued 190 got a 503 after a median wait of about 35 s; 410 still waiting at the end |
| Apache HttpClient, bigger pool, 1 s pool wait | Up 300 of 300 checks answered | Fast 503 600 of 600 in 2.0 s |
Why the whole app goes down
Spring Boot’s embedded Tomcat serves requests with a pool of 200 worker threads. Each /quote request holds its thread while it waits on the partner, and with no read timeout, it waits for as long as the partner keeps quiet. At 20 requests a second, all 200 threads are taken after 10 seconds. From then on, every new request queues for a thread, /ping included.
/ping, queues for a thread that never comes free.The part that hurts most: stopping the traffic doesn’t help. A thread comes back only when the partner answers or hangs up, and ours never did. A thread dump at the end of the run showed all 200 worker threads parked inside the HTTP call. If /ping were your liveness probe, an orchestrator would restart the pod, and at this traffic the fresh one would be stuck again about 10 seconds later, for as long as the partner stays silent.
The fix: two lines
# Applies to every client built from Spring Boot's RestClient.Builder
spring.http.clients.connect-timeout=1s
spring.http.clients.read-timeout=2s
Upgrade trap
Spring Boot 4 renamed these from spring.http.client.*, the names Spring Boot 3.4 and 3.5 used. The old names are still listed as deprecated, but 4.1 no longer reads them: with spring.http.client.read-timeout=2s, our call still hung, and nothing in the log said why.
Search your configuration for the old names, or add spring-boot-properties-migrator for one release. With it on the classpath, our app logged a warning naming each renamed key and applied it, so the call timed out after 2 seconds.
Inject Spring Boot’s RestClient.Builder rather than calling RestClient.create(), which builds its own client and ignores these properties. Then turn the timeout into an answer:
@RestController
class QuoteController {
private final RestClient partner;
// Inject Spring Boot's builder: RestClient.create() would ignore the timeout properties.
QuoteController(RestClient.Builder builder, @Value("${partner.url}") String partnerUrl) {
this.partner = builder.baseUrl(partnerUrl).build();
}
@GetMapping("/quote")
String quote() {
return partner.get().uri("/rates").retrieve().body(String.class);
}
// A timed-out call becomes a quick, honest 503 instead of a hang.
@ExceptionHandler(ResourceAccessException.class)
ResponseEntity<String> partnerUnavailable() {
return ResponseEntity.status(503).body("Rates are unavailable right now. Please try again shortly.");
}
}
With those in place, every one of the 600 calls got its 503 in 2.0 seconds, /ping answered all 300 checks, and about 40 calls were waiting on the partner at any moment: 20 requests a second times 2 seconds.
The read timeout did the work in that run, because our silent partner accepted every connection. The connect timeout is for a partner that never answers the connection attempt itself. Unlike a refusal, which fails at once, the attempt simply gets no reply. We simulated it with a server whose connection queue was full, so new attempts went unanswered. With no timeouts, the call was still hanging after 20 seconds. With the JDK client, whose read timeout covers the whole request, connecting included, the read timeout alone already returned a 503 in 2 seconds, and adding the connect timeout cut that to 1. With Apache HttpClient (more on that below), the read timeout alone was still hanging after 20 seconds; only the connect timeout stopped it. Set both.
Plot twist
Virtual threads kept the app answering every health check, with no timeouts at all. Every single partner call still hung: 600 of 600.
Virtual threads hide the problem
With spring.threads.virtual.enabled=true, Tomcat runs each request on a virtual thread, so there’s no pool of 200 to run out of. /ping answered every time. But nothing ever finished: when the test ended, 600 requests were still waiting and 600 connections to the partner were still open.
Your health checks would stay green while every customer who asked for a quote stared at a spinner. Virtual threads make waiting cheap. They don’t make it end. You still need the timeouts.
Gotcha
Add Apache HttpClient to your project, directly or through another library, and Spring Boot quietly switches to it. Your 2-second read timeout stops protecting the app.
The Apache HttpClient trap
Spring Boot picks the HTTP client by what’s on the classpath: Apache HttpClient first, then Jetty, Reactor Netty and finally the JDK’s own client. Apache does have a default read timeout of its own, but it’s 3 minutes: at 20 requests a second, all 200 threads are gone long before that. Its connection pool allows 5 connections per host and 25 in total, and a request waits up to 3 minutes for a free connection. The read timeout doesn’t cover that wait.
With the same 2-second timeouts, only 5 calls reached the partner at a time. A thread dump showed 195 of the 200 worker threads waiting for a pooled connection, and the app went down at 11.5 seconds. The calls that did fail took a median of about 35 seconds to do it. A bigger pool and a short wait for a connection fixed it:
import java.util.concurrent.TimeUnit;
import org.springframework.boot.http.client.HttpComponentsClientHttpRequestFactoryBuilder;
import org.springframework.boot.http.client.autoconfigure.ClientHttpRequestFactoryBuilderCustomizer;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
@Configuration
class PartnerClientConfig {
@Bean
ClientHttpRequestFactoryBuilderCustomizer<HttpComponentsClientHttpRequestFactoryBuilder> partnerPool() {
return builder -> builder
.withConnectionManagerCustomizer(pool -> pool.setMaxConnPerRoute(50).setMaxConnTotal(200))
.withDefaultRequestConfigCustomizer(request -> request.setConnectionRequestTimeout(1, TimeUnit.SECONDS));
}
}
Same traffic, same timeouts: 600 fast 503s, every health check answered, and about 40 connections to the partner at a time. The 1-second cap never came into play there: about 40 connections fit in a pool of 50. To test the cap, we shrank the pool to 10. Of the 600 calls, 440 got their 503 after waiting 1 second for a connection, the other 160 got theirs in 2 to 3 seconds, and every health check was still answered.
Picking the numbers
- Read timeout: a little above the partner’s slowest normal response. Take it from your metrics, not from their marketing page.
- Connect timeout: short. A healthy server normally accepts a connection in milliseconds, so 1 or 2 seconds is generous.
- Do the arithmetic: requests per second times the read timeout is how many threads can be waiting on the partner. We measured about 40 at 20 a second and 2 seconds. Keep it well under your 200.
- Pool size: with Apache HttpClient, at least that many connections per host, and a connection-request timeout of a second or two.
Beyond timeouts
Timeouts make a silent partner cost you two seconds a request instead of everything. Two patterns go further. A circuit breaker stops calling a partner that keeps failing and answers straight away for a while, and a bulkhead caps how many threads one partner may hold at once. Resilience4j provides both. Spring Framework 7, which Spring Boot 4 is built on, also has a @ConcurrencyLimit annotation, enabled with @EnableResilientMethods. Set policy = REJECT: by default it makes extra callers wait, which is the problem you’re trying to avoid. Where you can, serve the last good answer from a cache, and alert on the partner’s response time so you hear about it before your customers do.
Checklist
- Set
spring.http.clients.connect-timeoutandspring.http.clients.read-timeout. Coming from Spring Boot 3.4 or 3.5, rename the oldspring.http.client.*keys. - Build clients from Spring Boot’s
RestClient.Builder, so the settings apply. - Turn a timeout into a fast, honest error, or a cached answer.
- Using Apache HttpClient? Size the pool and cap the wait for a connection.
- Don’t count on virtual threads: they keep the app up, not your users’ requests.
- Point a test at a partner that never answers, before a real one does it for you.