ADR-052: Chat API Production Readiness - Error Handling, Shutdown, and HTTP Standards

Status

ACCEPTED - Production-Ready Configuration

Date

2025-12-12

Context and Problem Statement

The Chat API (/v1/chat/completions) is the REST API backend for iDempiere UI clients (ZK/Angular) to provide conversational AI assistance. Several production-readiness issues were identified:

Problems

  1. Graceful Shutdown: Server produced error spam (NullPointerException, RejectedExecutionException) when stopping, making it difficult to distinguish real errors from shutdown noise.

  2. CORS Configuration: Restricted to single origin (${IDEMPIERE_URL:http://localhost:8080}), making it difficult to use with reverse proxies, load balancers, or multiple frontend deployments.

  3. HTTP Status Codes: All chat errors returned 400 BAD_REQUEST regardless of error type, violating REST best practices and making it harder for clients to handle errors appropriately.

  4. Developer Experience: No logging of internal error details, making debugging difficult for developers.

  5. Documentation Accuracy: Command documentation showed non-existent server chat-api subcommand.

Requirements

Decision

1. Implement Graceful Shutdown

Add graceful shutdown handling to avoid error spam and allow in-flight requests to complete:

Configuration (application.properties):

# Wait up to 10 seconds for in-flight requests to complete before forcing shutdown
quarkus.shutdown.timeout=10s

# Suppress shutdown-related errors (cosmetic - happens during termination)
quarkus.log.category."io.vertx.ext.web.RoutingContext".level=OFF
quarkus.log.category."io.quarkus.vertx.http.runtime.QuarkusErrorHandler".level=WARN

Shutdown Hook (ApiServerCommand.java):

// Register shutdown hook for graceful termination
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
    System.out.println("\nShutting down gracefully...");
}));

CDI Container Check (ExceptionMappers.java and ChatAgentService.java):

// Check if CDI container is still active (not shutting down)
var container = Arc.container();
if (container == null) {
    // Container is shutting down, log at debug level only
    LOG.debugf("Error during shutdown (ignored): %s", e.getMessage());
    return Response.status(Response.Status.SERVICE_UNAVAILABLE)
        .entity(new ErrorResponse("Service is shutting down", "shutdown_error", null, null))
        .build();
}

Agent Service Shutdown Detection (ChatAgentService.java):

private boolean isShuttingDown() {
    try {
        var container = Arc.container();
        return container == null;
    } catch (Exception e) {
        return true;  // Assume shutting down if can't access container
    }
}

// Used in exception handlers
catch (Exception e) {
    if (isShuttingDown()) {
        LOG.debugf("Chat processing interrupted during shutdown for request %s", requestId);
        return ChatResponse.error("Service is shutting down", "shutdown_error", requestId);
    }
    LOG.errorf(e, "Chat processing failed for request %s", requestId);
    // ... normal error handling
}

2. Open CORS Configuration

Change from restricted single-origin to open CORS policy:

Before:

%chat-api.quarkus.http.cors.origins=${IDEMPIERE_URL:http://localhost:8080}

After:

%chat-api.quarkus.http.cors.origins=*

Rationale:

3. Standard HTTP Status Code Mapping

Implement dynamic status code mapping based on error type, following OpenAI API conventions:

Implementation (ChatApiResource.java):

private Response.Status mapErrorToStatus(ChatResponse.ErrorInfo error) {
    return switch (error.type()) {
        case "invalid_request_error" -> Response.Status.BAD_REQUEST;           // 400
        case "authentication_error" -> Response.Status.UNAUTHORIZED;            // 401
        case "permission_error", "boundary_violation" -> Response.Status.FORBIDDEN;  // 403
        case "not_found_error" -> Response.Status.NOT_FOUND;                   // 404
        case "rate_limit_error", "quota_exceeded" -> Response.Status.TOO_MANY_REQUESTS; // 429
        case "server_error", "internal_error" -> Response.Status.INTERNAL_SERVER_ERROR; // 500
        case "service_unavailable", "shutdown_error" -> Response.Status.SERVICE_UNAVAILABLE; // 503
        default -> Response.Status.INTERNAL_SERVER_ERROR;                      // 500 (fallback)
    };
}

Status Code Matrix:

Error Type HTTP Status Use Case
invalid_request_error 400 Bad Request Invalid parameters, missing fields, malformed JSON
authentication_error 401 Unauthorized Missing or invalid API key
permission_error 403 Forbidden Insufficient permissions
boundary_violation 403 Forbidden Agent boundary violation (e.g., write in READ_ONLY mode)
not_found_error 404 Not Found Model doesn't exist
rate_limit_error 429 Too Many Requests Rate limiting enforced
quota_exceeded 429 Too Many Requests Daily/monthly quota exceeded
server_error 500 Internal Server Error LLM provider error, unexpected failure
internal_error 500 Internal Server Error Internal processing error
service_unavailable 503 Service Unavailable No providers available
shutdown_error 503 Service Unavailable Server shutting down

4. Developer Error Logging

Add comprehensive error logging with context for debugging:

Implementation (ChatApiResource.java):

private void logErrorForDevelopers(ChatResponse.ErrorInfo error, Response.Status status, ChatRequest request) {
    String logMessage = String.format(
        "[%d %s] error_type=%s, error_code=%s, message=\"%s\", model=%s, param=%s",
        status.getStatusCode(),
        status.getReasonPhrase(),
        error.type(),
        error.code(),
        error.message(),
        request.model(),
        error.param()
    );

    // Log level based on error type
    switch (error.type()) {
        case "invalid_request_error", "not_found_error" ->
            LOG.infof("Client error: %s", logMessage);  // INFO - expected client errors

        case "authentication_error", "permission_error", "boundary_violation" ->
            LOG.warnf("Security error: %s", logMessage); // WARN - security issues

        case "rate_limit_error", "quota_exceeded" ->
            LOG.infof("Rate limit: %s", logMessage);    // INFO - rate limiting working

        case "server_error", "internal_error" ->
            LOG.errorf("Server error: %s", logMessage); // ERROR - unexpected issues

        case "service_unavailable", "shutdown_error" ->
            LOG.debugf("Service unavailable: %s", logMessage); // DEBUG - temporary states

        default ->
            LOG.errorf("Unknown error type: %s", logMessage); // ERROR - investigate
    }
}

Log Output Examples:

INFO  Client error: [400 Bad Request] error_type=invalid_request_error, error_code=model_not_found, message="Model 'gpt-99' not found", model=gpt-99, param=model

WARN  Security error: [403 Forbidden] error_type=boundary_violation, error_code=read_only_violation, message="Write operation not allowed", model=llama3.2, param=null

ERROR Server error: [500 Internal Server Error] error_type=server_error, error_code=llm_provider_error, message="Ollama connection failed", model=llama3.2, param=null

5. Documentation Updates

Fixed Command (CHAT-API-GUIDE.md):

# CORRECT (there is no 'chat-api' subcommand)
java -Dquarkus.profile=chat-api -jar idempiere-hub-runner.jar server api

Created Documentation:

Consequences

Positive

✅ Clean Shutdown: No error spam, clear "Shutting down gracefully..." message, 10-second grace period for requests

✅ Flexible Deployment: CORS open to all origins, works with reverse proxies, load balancers, and multiple frontends

✅ REST Compliant: Standard HTTP status codes following RFC 9110, OpenAI API compatible

✅ Better Client Experience: Clients can handle errors appropriately based on status codes (retry on 5xx, don't retry on 4xx)

✅ Developer Friendly: Comprehensive error logging with context, appropriate log levels, easy debugging

✅ Documentation Accuracy: Correct commands, complete reference guides

Negative

⚠️ CORS Security: Open CORS (*) means any origin can call the API. Organizations with strict security requirements may need to restrict this in production.

Mitigation: Document CORS restriction for production deployments:

# Production: Restrict to specific origins
%chat-api.quarkus.http.cors.origins=https://erp.company.com,https://app.company.com

Neutral

📝 Logging Volume: Error logging adds more log entries, but uses appropriate levels (INFO for expected errors, ERROR only for unexpected issues)

📝 Code Complexity: Status code mapping adds ~50 lines, but improves maintainability and clarity

Implementation Details

Files Modified

Configuration:

Code:

Tests:

Documentation:

Testing Verification

Graceful Shutdown:

java -Dquarkus.profile=chat-api -jar target/idempiere-hub-runner.jar server api
# Press Ctrl+C → Should show "Shutting down gracefully..." without errors

Status Codes:

# 400 - Invalid request
curl -i -X POST http://localhost:8081/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"invalid","messages":[]}'

# 404 - Model not found
curl -i http://localhost:8081/v1/models/non-existent

Error Logging:

# Check logs show formatted error messages with appropriate levels
tail -f logs/application.log | grep -E "(Client error|Security error|Server error)"

Standards Compliance

HTTP Status Codes (RFC 9110)

✅ All status codes are standard:

❌ NO custom or non-standard codes

OpenAI API Compatibility

Our status code usage matches OpenAI API conventions:

Scenario Our Code OpenAI Code
Success 200 200
Invalid request 400 400
Auth failed 401 401
Permission denied 403 403
Not found 404 404
Rate limited 429 429
Server error 500 500
Unavailable 503 503

References:

Alternatives Considered

1. Keep Restricted CORS

Rejected: Too inflexible for real-world deployments with proxies/load balancers.

2. Use Only 400/500 Status Codes

Rejected: Violates REST principles, makes client error handling harder, not OpenAI compatible.

3. Custom HTTP Status Codes

Rejected: Non-standard, breaks HTTP clients, violates RFC 9110.

4. No Error Logging

Rejected: Makes debugging impossible for developers.

5. Log All Errors as ERROR Level

Rejected: Creates noise, makes it hard to find real issues (client errors are expected).

References

External Standards

Internal Documentation

Implementation Files


Decision Date: 2025-12-12 Implemented: 2025-12-12 Status: ✅ Production Ready

Path: /docs/developers/architecture/idempiere-hub/052-chat-api-production-readiness