ADR-056: Chat API Client Cancellation and Timeout Handling
Status: Partially Implemented (Phase 0 complete) Date: 2025-12-12 (Proposed), 2025-12-14 (Phase 0 Implemented) Context: Fixing Issue #2 - Client cancellation doesn't stop server processing
Problem Statement
When a client disconnects (Ctrl+C) from /v1/chat/completions, two critical issues occur:
Issue 1: Client Request Hangs
Symptom: Client request never completes, hangs indefinitely Impact: Poor UX, resource leaks on client side
Issue 2: Server Keeps Processing
Symptom: LangChain4j LLM call continues running after client disconnect Impact: Wasted compute, wasted API costs, zombie threads
Error During Shutdown
ERROR [io.qua.ver.htt.run.QuarkusErrorHandler] HTTP Request to /v1/chat/completions failed
[Error Occurred After Shutdown]: java.lang.NullPointerException:
Cannot invoke "io.quarkus.arc.ArcContainer.getActiveContext..." because "<local2>" is null
Root Cause Analysis
1. Non-Streaming Requests (Blocking Call)
Location: ChatApiResource.java:91
ChatResponse response = agentService.chat(request, agentContext);
Problems:
- ❌ Blocking call with no timeout
- ❌ No way to detect client disconnect
- ❌ No way to interrupt ongoing LLM call
2. LangChain4j Integration (No Cancellation)
Location: ChatAgentService.java:282
String agentResponse = chatAgent.chat(userMessage);
Problems:
- ❌ Synchronous LangChain4j call has no timeout
- ❌ No cancellation propagation mechanism
- ❌ Long-running LLM calls can't be interrupted
3. Streaming Requests (Output Failure)
Location: ChatApiResource.java:169-208
StreamingOutput stream = output -> {
try (var writer = new java.io.OutputStreamWriter(output)) {
agentService.chatStream(request, agentContext, new StreamCallback() {
@Override
public void onChunk(StreamChunk chunk) {
writer.write("data: " + json + "\n\n"); // ⚠️ Fails after disconnect
writer.flush();
}
});
}
};
Problems:
- ⚠️ Writer fails when client disconnects (IOException)
- ❌ But LangChain4j subscribe() keeps running
- ❌ No cancellation propagated to
chatAgent.streamChat()
4. SmallRye Context Manager NPE (CRITICAL - Newly Discovered)
Location: ChatAgentService.java:289 → JaxRsHttpClient.execute() (LangChain4j)
Root Cause Stack:
java.lang.NullPointerException: Cannot invoke
"io.smallrye.context.SmallRyeContextManager.defaultThreadContext()"
because the return value of "io.smallrye.context.SmallRyeContextManagerProvider.getManager()" is null
at SmallRyeThreadContext.getCurrentThreadContextOrDefaultContexts
at DefaultContextPropagationInterceptor.getThreadContext
at Infrastructure.decorate
at UniOnItem.transform
at ClientSendRequestHandler.createRequest
at JaxRsHttpClient.execute (LangChain4j)
at OllamaClient.chat
at ChatAgent.chat
at ChatAgentService.executeChat:289
Problem:
- SmallRye Context Manager becomes null BEFORE Arc container during shutdown
- LangChain4j JAX-RS client tries to propagate reactive context
- NPE occurs INSIDE the LLM call, not in exception handling
- Our
isShuttingDown()check in catch block is TOO LATE
Impact: This is the actual production NPE - occurs before the Arc container NPE
5. Shutdown Exception Mapper Failure (SECONDARY)
Location: ExceptionMappers.java (during shutdown)
Problem:
- Arc container is null during shutdown
- Exception mapper tries to access CDI beans
- Results in NullPointerException logged as error
Impact: This NPE only occurs if SmallRye NPE (#4) gets through
Decision
Implement multi-layer cancellation and timeout handling:
1. Add Request Timeout (All Requests)
- Default: 120 seconds for LLM calls
- Configurable via
chat-api.agent.request-timeout-seconds - Wrap LangChain4j calls in CompletableFuture with timeout
2. Detect Client Disconnect (Streaming Only)
- Detect IOException when writing to output stream
- Cancel LangChain4j reactive stream subscription
- Clean up resources immediately
3. Add Graceful Shutdown Handling
- Reject new requests during shutdown
- Wait for in-flight requests (up to 30 seconds)
- Improve exception mapper to handle shutdown state
4. Add Interruption Support (Best Effort)
- For blocking calls: use Thread.interrupt() on timeout
- For reactive calls: use subscription.cancel()
- Document that actual LLM API call may not be interruptible
Implementation
Phase 0: Prevent LLM Calls During Shutdown (IMPLEMENTED - v1.66.0)
File: src/main/java/org/idempiere/cli/chatapi/agent/ChatAgentService.java
Problem: SmallRye Context Manager NPE occurs INSIDE the LLM call before our catch block runs.
Solution: Check for shutdown BEFORE calling chatAgent.chat() or chatAgent.streamChat().
private ChatResponse executeChat(ChatRequest request, AgentContext context,
ModelRouter.RoutingResult routing, String requestId) {
String model = request.model() != null ? request.model() : defaultModel;
try {
// Check if shutting down BEFORE making LLM call
// This prevents SmallRye Context Manager NPE during shutdown
if (isShuttingDown()) {
LOG.debugf("Request %s rejected - service is shutting down", requestId);
throw new RuntimeException("Service is shutting down");
}
// Extract the last user message to send to the agent
ChatMessage lastMessage = request.messages().get(request.messages().size() - 1);
String userMessage = lastMessage.content();
// Call the LangChain4j ChatAgent (SAFE - shutdown checked above)
String agentResponse = chatAgent.chat(userMessage);
// ... rest of method
} catch (Exception e) {
// Still check in catch block for errors DURING execution
if (isShuttingDown()) {
LOG.debugf("Agent execution interrupted during shutdown for request %s", requestId);
throw new RuntimeException("Service is shutting down", e);
}
// ... normal error handling
}
}
private void executeStreamingChat(ChatRequest request, AgentContext context,
ModelRouter.RoutingResult routing, String requestId,
StreamCallback callback) {
try {
// Check if shutting down BEFORE making LLM streaming call
if (isShuttingDown()) {
LOG.debugf("Streaming request %s rejected - service is shutting down", requestId);
callback.onError(new RuntimeException("Service is shutting down"));
return;
}
// Safe to call LangChain4j streaming (shutdown checked above)
chatAgent.streamChat(userMessage).subscribe().with(...);
} catch (Exception e) {
// ... error handling
}
}
/**
* Enhanced shutdown detection - checks BOTH Arc and SmallRye managers.
*/
private boolean isShuttingDown() {
try {
// Check Arc CDI container
var container = Arc.container();
if (container == null) {
return true;
}
// Check SmallRye Context Manager (used by LangChain4j JAX-RS client)
var contextManager = io.smallrye.context.SmallRyeContextManagerProvider.getManager();
if (contextManager == null) {
return true;
}
return false;
} catch (Exception e) {
// If we can't access containers, assume we're shutting down
return true;
}
}
Result:
- ✅ Prevents SmallRye Context Manager NPE
- ✅ Prevents Arc container NPE in ExceptionMapper
- ✅ Graceful "Service is shutting down" error message
- ✅ No stack traces in logs during shutdown
- ✅ All 4 unit tests pass (ShutdownNPETest.java)
Phase 1: Add Timeout Wrapper (PROPOSED)
File: src/main/java/org/idempiere/cli/chatapi/agent/ChatAgentService.java
@ConfigProperty(name = "chat-api.agent.request-timeout-seconds", defaultValue = "120")
int requestTimeoutSeconds;
private ChatResponse executeChat(ChatRequest request, AgentContext context,
ModelRouter.RoutingResult routing, String requestId) {
// Wrap in CompletableFuture with timeout
CompletableFuture<String> future = CompletableFuture.supplyAsync(() -> {
try {
ChatMessage lastMessage = request.messages().get(request.messages().size() - 1);
return chatAgent.chat(lastMessage.content());
} catch (Exception e) {
throw new CompletionException("Agent execution failed", e);
}
});
try {
// Wait with timeout
String agentResponse = future.get(requestTimeoutSeconds, TimeUnit.SECONDS);
// Build response (existing code)
// ...
} catch (TimeoutException e) {
future.cancel(true); // Interrupt the thread
LOG.warnf("Request %s timed out after %d seconds", requestId, requestTimeoutSeconds);
throw new RuntimeException("Request timed out after " + requestTimeoutSeconds + " seconds");
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new RuntimeException("Request interrupted", e);
} catch (ExecutionException e) {
throw new RuntimeException("Agent execution failed", e.getCause());
}
}
Phase 2: Detect Client Disconnect in Streaming
File: src/main/java/org/idempiere/cli/chatapi/agent/ChatAgentService.java
private void executeStreamingChat(ChatRequest request, AgentContext context,
ModelRouter.RoutingResult routing, String requestId,
StreamCallback callback) {
String model = request.model() != null ? request.model() : defaultModel;
try {
ChatMessage lastMessage = request.messages().get(request.messages().size() - 1);
String userMessage = lastMessage.content();
int inputTokens = estimateInputTokens(request);
final int[] outputTokenCount = {0};
callback.onChunk(StreamChunk.start(requestId, model, ChatMessage.Role.ASSISTANT));
// Track subscription for cancellation
AtomicReference<Subscription> subscriptionRef = new AtomicReference<>();
AtomicBoolean clientDisconnected = new AtomicBoolean(false);
chatAgent.streamChat(userMessage)
.subscribe().with(
subscription -> subscriptionRef.set(subscription),
// On each token
token -> {
if (clientDisconnected.get()) {
// Client disconnected, stop processing
return;
}
try {
callback.onChunk(StreamChunk.content(requestId, model, 0, token));
outputTokenCount[0] += token.length() / 4;
} catch (Exception e) {
// Client disconnect detected (IOException from writer)
LOG.debugf("Client disconnected for request %s, cancelling stream", requestId);
clientDisconnected.set(true);
// Cancel subscription to stop LangChain4j processing
Subscription sub = subscriptionRef.get();
if (sub != null) {
sub.cancel();
}
}
},
// On error
error -> {
if (!clientDisconnected.get()) {
LOG.errorf(error, "Streaming chat failed for request %s", requestId);
callback.onError(error);
}
},
// On complete
() -> {
if (clientDisconnected.get()) {
LOG.debugf("Stream completed after client disconnect for request %s", requestId);
return;
}
TokenUsage usage = TokenUsage.of(
inputTokens,
outputTokenCount[0],
model,
costCalculator.calculateCost(model, inputTokens, outputTokenCount[0], 0)
);
callback.onChunk(StreamChunk.finish(requestId, model, "stop", usage));
callback.onComplete(usage);
costGuard.recordSpending(
usage.estimatedCostUsd(),
context.getClientId(),
context.getUserId()
);
}
);
} catch (Exception e) {
LOG.errorf(e, "Failed to start streaming chat for request %s", requestId);
callback.onError(e);
}
}
Phase 3: Improve Shutdown Handling
File: src/main/java/org/idempiere/cli/chatapi/api/ExceptionMappers.java
@Provider
public class GenericExceptionMapper implements ExceptionMapper<Exception> {
@Override
public Response toResponse(Exception exception) {
// Check if shutting down BEFORE accessing CDI
if (isShuttingDown()) {
return Response.status(Response.Status.SERVICE_UNAVAILABLE)
.entity(ChatResponse.error(
"Service is shutting down",
"shutdown_error",
null
))
.build();
}
// Normal exception handling (can access CDI beans safely)
// ...
}
private boolean isShuttingDown() {
try {
var container = Arc.container();
return container == null;
} catch (Exception e) {
return true;
}
}
}
Phase 4: Add Configuration
File: src/main/resources/application.properties
# Chat API Agent Configuration
chat-api.agent.request-timeout-seconds=120
chat-api.agent.shutdown-grace-period-seconds=30
Trade-offs
Timeout Approach
Pros:
- ✅ Prevents indefinite hangs
- ✅ Protects against slow/unresponsive LLM providers
- ✅ Clear error message for users
Cons:
- ❌ May interrupt legitimate long-running requests
- ❌ Some LLM API calls may not be interruptible (API-side)
- ❌ Need to tune timeout value
Decision: Use 120 seconds default (2 minutes), make configurable
Client Disconnect Detection
Pros:
- ✅ Immediately stops wasted processing
- ✅ Reduces API costs
- ✅ Frees up threads faster
Cons:
- ❌ Best-effort only (LangChain4j may not support cancellation)
- ❌ Partial results are lost
Decision: Implement best-effort cancellation, document limitations
Thread Interruption
Pros:
- ✅ Standard Java mechanism
- ✅ Works for well-behaved blocking code
Cons:
- ❌ Not all code respects interruption
- ❌ LangChain4j HTTP client may ignore interrupt
- ❌ Can leave resources in inconsistent state
Decision: Use future.cancel(true) but document it's best-effort
Configuration
| Property | Default | Description |
|---|---|---|
chat-api.agent.request-timeout-seconds |
120 |
Maximum time for LLM request |
chat-api.agent.shutdown-grace-period-seconds |
30 |
Time to wait for in-flight requests during shutdown |
Testing
Test 1: Non-Streaming Timeout
# Send request that takes > 120 seconds
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-4","messages":[{"role":"user","content":"..."}]}'
# Expected: 408 Request Timeout after 120 seconds
Test 2: Streaming Client Disconnect
# Start streaming request, Ctrl+C after 5 seconds
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-4","messages":[...],"stream":true}'
# Expected: Server logs "Client disconnected", cancels processing
Test 3: Graceful Shutdown
# Start long request, then Ctrl+C server
curl -X POST http://localhost:8080/v1/chat/completions &
sleep 2
pkill -INT java # Send SIGINT to server
# Expected: Server waits up to 30 seconds for request to complete
Limitations
-
LangChain4j API Calls Not Interruptible
- Anthropic/OpenAI HTTP calls may continue even after timeout/cancellation
- Best effort: we stop processing locally but API call may complete
- API costs still incurred for interrupted requests
-
Reactive Streams Cancellation
- LangChain4j may not support cancellation on all providers
- Subscription.cancel() is advisory, not guaranteed
-
Partial Results Lost
- On timeout/disconnect, partial LLM responses are discarded
- No mechanism to resume or recover partial progress
Success Criteria
Phase 0 (IMPLEMENTED - v1.66.0)
- ✅ No SmallRye Context Manager NPE during shutdown
- ✅ No Arc Container NPE during shutdown
- ✅ Graceful "Service is shutting down" error message
- ✅ All 4 unit tests pass (ShutdownNPETest.java)
- ✅ Production error eliminated (verified with actual Ctrl+C test)
Phase 1-4 (PROPOSED - Future)
- ⏳ Client requests complete (success or error) within timeout period
- ⏳ Server logs indicate client disconnect detection
- ⏳ Thread pool doesn't grow unbounded during load test with client disconnects
Monitoring
Add metrics to track:
chat_api.requests.timeout.count- Number of timed-out requestschat_api.requests.client_disconnect.count- Client disconnects detectedchat_api.requests.duration.p99- 99th percentile request duration
References
- Issue #2: Client Cancellation Bug Report (User bug report 2025-12-12)
- ADR-048: iDempiere AI Hub Integration
- ADR-054: AI Tool Architecture Clarity
- LangChain4j Reactive Streams
- Quarkus Graceful Shutdown
Version: v1.66.0 (Phase 0 implemented), v1.67.0+ (Phase 1-4 proposed) Status: Partially Implemented
Implementation History
- 2025-12-12: ADR created, problem documented
- 2025-12-14: Phase 0 implemented and tested
- Root cause identified: SmallRye Context Manager NPE occurs BEFORE Arc container NPE
- Fix: Check
isShuttingDown()BEFORE calling LangChain4j - Enhanced:
isShuttingDown()now checks both Arc and SmallRye managers - Tests: ShutdownNPETest.java with 4 passing tests
- Result: Production NPE eliminated
Next Steps
Phase 1-4 (timeout wrapper, client disconnect detection, etc.) remain as proposed enhancements for future releases.