ADR-058: Logging and Configuration Architecture

Status: 🟢 Phase 4 Complete (All Phases Implemented) Date: 2025-12-15 (Updated: 2025-12-15) Context: Architectural review by quarkus-reactive-architect expert Related: ADR-048 (Hub Architecture), ADR-054 (Tool Architecture), ADR-055 (Ecosystem Review)

Phase 3 Completion: All throwable exceptions migrated to structured error codes. See Phase 3 Completion Summary below. Phase 4 Completion: Structured logging, metrics, and distributed tracing implemented. See Phase 4 Completion Summary below.


Phase 3 Completion Summary

Completed: December 15, 2025 Migration Method: Manual refactoring + 3 automated batch migrations (Python scripts)

Overall Statistics

Batch Migration Efficiency

Three batch migrations were performed using Python regex scripts:

  1. RestADService (53 exceptions) - Previous session

    • Pattern: IOException | InterruptedException → OperationalException(IDEMPIERE_API_UNAVAILABLE)
    • Efficiency: ~50x speedup vs one-by-one
  2. PackCommand HTTP Errors (7 exceptions) - Commit 20595ad

    • Pattern: IOException("HTTP error...") → OperationalException(IDEMPIERE_API_UNAVAILABLE)
    • Automation saved ~1 hour manual work
  3. Config Validation (3 exceptions) - Commit 7b90b1e

    • Files: WorkflowCommand, ServerCommand, CacheCommand
    • Pattern: IllegalStateException("API not configured") → ClientException(DATABASE_NOT_CONFIGURED)

Error Code Distribution

Error Code Usage Exception Type Use Cases
IDEMPIERE_API_UNAVAILABLE (1004) 58 OperationalException HTTP API failures, REST client errors
FILE_SYSTEM_ERROR (5005) 10 ServerException File I/O operations, write failures
DATABASE_NOT_CONFIGURED (2001) 5 ClientException Missing API/DB configuration
DB_CONNECTION_TIMEOUT (1002) 4 OperationalException Database connection issues
INVALID_INPUT (2003) 3 ClientException JSON parsing, validation errors
RESOURCE_NOT_FOUND (2008) 2 ClientException Missing resources, processes
RESOURCE_TEMPORARILY_UNAVAILABLE (1009) 8 OperationalException Terminal I/O, embedding model warmup
UNEXPECTED_ERROR (5000) 3 ServerException Cryptographic failures, serialization

Key Files Migrated

Services Package:

Command Package:

Wizard Package:

RAG Package:

Generator Package:

Migration Commits

  1. dd779f6 - SqlADService + PackOutADService (2 exceptions)
  2. c073912 - DockerPostgresService (1 exception)
  3. 803291b - WindowGenerator (1 exception, semantic fix)
  4. 20595ad - PackCommand batch migration (7 exceptions)
  5. 7b90b1e - API config checks batch (3 exceptions)
  6. 894d291 - Final 8 exceptions (Wizard + RAG packages) - Phase 3 COMPLETE 🎉
  7. Current - Documentation updates

Success Criteria - All Achieved ✅


Phase 4 Completion Summary

Completed: December 15, 2025 Implementation Method: Quarkus extension integration + custom observability services

Overall Statistics

Features Implemented

1. JSON Logging (Production Profile)

2. MDC (Mapped Diagnostic Context) - Request ID Tracking

3. OpenTelemetry Distributed Tracing

4. Prometheus Metrics

5. Exception Mappers with Metrics

Files Created

Observability Services:

Exception Mappers (Chat API):

Configuration Sections Added

JSON Logging:

quarkus.log.console.json=false
%prod.quarkus.log.console.json=true
%prod.quarkus.log.console.json.fields.service_name.value=idempiere-hub

OpenTelemetry:

quarkus.application.name=idempiere-hub
quarkus.otel.enabled=false
%chat-api.quarkus.otel.enabled=true
%chat-api.quarkus.otel.traces.sampler.ratio=0.1

Prometheus Metrics:

quarkus.micrometer.enabled=true
quarkus.micrometer.export.prometheus.enabled=true

Metrics Example

# HELP hub_errors_total Total number of errors by error code
# TYPE hub_errors_total counter
hub_errors_total{code="1003",name="LLM_PROVIDER_UNAVAILABLE",severity="WARN",retryable="true"} 42.0
hub_errors_total{code="2003",name="INVALID_INPUT",severity="INFO",retryable="false"} 15.0
hub_errors_total{code="5005",name="FILE_SYSTEM_ERROR",severity="ERROR",retryable="false"} 3.0

JSON Log Example

{
  "timestamp": "2025-12-15T23:30:00.123Z",
  "sequence": 1234,
  "loggerClassName": "org.idempiere.cli.chatapi.ChatAgentService",
  "loggerName": "org.idempiere.cli.chatapi.ChatAgentService",
  "level": "WARN",
  "message": "[1003] Ollama server not available - ensure Ollama is running",
  "threadName": "executor-thread-1",
  "threadId": 42,
  "mdc": {
    "request_id": "123e4567-e89b-12d3-a456-426614174000"
  },
  "service_name": "idempiere-hub",
  "version": "1.71.0",
  "environment": "production"
}

Success Criteria - All Achieved ✅

Benefits Achieved

  1. Production Observability: JSON logs ready for log aggregation platforms
  2. Distributed Tracing: Request correlation across services via request_id and OpenTelemetry
  3. Error Monitoring: Prometheus metrics dashboard shows top errors by code
  4. Performance Tuning: Metrics for HTTP endpoints, clients, JVM, system
  5. Zero Overhead in CLI: Observability disabled by default, enabled per profile
  6. Ready for Cloud: OpenTelemetry OTLP exporter configurable via environment variables

Executive Summary

The iDempiere Hub has scattered configuration (75+ @ConfigProperty usages) and inconsistent error handling patterns across its 3 interfaces (CLI, Chat API, MCP Server). While good foundations exist (ToolResult, CliErrorCode, ToolExecutionException), they're not used consistently, leading to log pollution (100+ line stack traces for operational errors) and configuration duplication (Anthropic config appears 3 times).

Critical Issues

  1. Configuration Scattered: 75+ @ConfigProperty usages, no single source of truth, Anthropic duplicated 3x
  2. Log Pollution: 100+ line stack traces for operational errors (Ollama down, connection refused)
  3. Error Code Inconsistency: CLI uses 38 codes, Chat API uses 8 codes, services use none
  4. No Error Classification: Operational errors (WARN) mixed with bugs (ERROR)
  5. Multiple Error Patterns: ToolResult, ToolExecutionException, RuntimeException, CliErrorCode (4 different approaches)

Proposed Solution

4-phase implementation to systematically address configuration management and error handling:

Phase Focus Effort Impact
Phase 1 Configuration consolidation (@ConfigMapping) 1 week Single source of truth, type-safe config
Phase 2 Unified error codes (HubErrorCode enum) 1 week Consistent error codes across all interfaces
Phase 3 Refactor error handling (typed exceptions) 2 weeks 80% reduction in log volume, clear error classification
Phase 4 Structured logging (JSON, optional) 1 week Production observability, metrics dashboard

Context

Current State Assessment

Configuration Landscape:

Configuration Patterns (4 in use):

  1. @ConfigProperty (75+ usages) - scattered, no type safety
  2. @ConfigMapping (1 usage) - type-safe, best practice
  3. Environment variables - not discoverable
  4. Hardcoded values - not configurable

Logging Patterns:

  1. Simple error logging - LOG.error("msg") - no context
  2. Error with exception - LOG.errorf(e, "msg") - shows full stack trace for operational errors
  3. Warn with error code - LOG.warnf("[%d] %s", code, msg) - only in ExceptionMappers
  4. Error codes - CLI only (CliErrorCode)

Error Handling Patterns:

  1. Catch + Log + Rethrow RuntimeException - stack trace logged twice
  2. Wrap in ToolExecutionException - better, but Chat API tools only
  3. Return ToolResult - best for tools, no exceptions
  4. Check + Early Return - CLI only

Example Problem (Connection Error):

Before (100+ lines):

jakarta.ws.rs.ProcessingException: io.netty.channel.AbstractChannel$AnnotatedConnectException: Connection refused: localhost/127.0.0.1:11434
    at org.jboss.resteasy.reactive.client.handlers.ClientSendRequestHandler$2.accept(...)
    at org.jboss.resteasy.reactive.client.handlers.ClientSendRequestHandler$2.accept(...)
    [... 60+ more netty/vertx frames ...]

After (1 line):

WARN  [1002] LLM provider 'ollama' not available - check that the service is running (ensure Ollama is started: ollama serve)

Decision

Phase 1: Configuration Consolidation

Decision: Migrate all configuration to @ConfigMapping interfaces organized by domain.

Structure:

src/main/java/org/idempiere/cli/config/
├── LlmProviderConfig.java      # Multi-provider LLM config
├── DatabaseConfig.java          # Database connection pools
├── ChatApiConfig.java           # Chat API-specific (guardrails, observability)
├── CliConfig.java               # CLI-specific settings
├── McpConfig.java               # MCP server-specific settings
└── ObservabilityConfig.java    # Logging, metrics, tracing (optional Phase 4)

Example: LlmProviderConfig.java

@ConfigMapping(prefix = "idempiere.hub.llm")
public interface LlmProviderConfig {

    @WithDefault("anthropic")
    String defaultProvider();

    @WithDefault("120s")
    Duration requestTimeout();

    @WithDefault("300s")
    Duration streamingTimeout();

    Map<String, ProviderConfig> providers();

    interface ProviderConfig {
        boolean enabled();
        Optional<String> baseUrl();
        Optional<String> apiKey();
        Optional<String> defaultModel();
        Optional<Duration> timeout();
    }
}

application.properties reorganization:

# ==================== Section 1: Quarkus Core (Global) ====================
quarkus.banner.enabled=false
quarkus.log.level=WARN
quarkus.log.category."org.idempiere.cli".level=INFO

# ==================== Section 2: Shared Infrastructure (74%) ====================
# Used by CLI, Chat API, MCP Server

# 2.1 LLM Providers
idempiere.hub.llm.default-provider=anthropic
idempiere.hub.llm.request-timeout=120s
idempiere.hub.llm.streaming-timeout=300s

idempiere.hub.llm.providers.anthropic.enabled=true
idempiere.hub.llm.providers.anthropic.api-key=${ANTHROPIC_API_KEY:not-configured}
idempiere.hub.llm.providers.anthropic.default-model=claude-sonnet-4-5-20250929

idempiere.hub.llm.providers.ollama.enabled=true
idempiere.hub.llm.providers.ollama.base-url=http://localhost:11434
idempiere.hub.llm.providers.ollama.default-model=llama3.2

# 2.2 Database (see DatabaseConfig)
idempiere.hub.database.default.host=${IDEMPIERE_DB_HOST:localhost}
idempiere.hub.database.default.port=${IDEMPIERE_DB_PORT:5433}

# 2.3 RAG (existing RagConfig)
idempiere.hub.rag.enabled=true

# ==================== Section 3: CLI Profile (19%) ====================
%cli.quarkus.http.port=0

# ==================== Section 4: Chat API Profile (4%) ====================
%chat-api.quarkus.http.port=8081
%chat-api.idempiere.hub.chat-api.guardrails.enabled=true

# ==================== Section 5: MCP Profile (3%) ====================
%mcp.quarkus.http.port=8765

Benefits:


Phase 2: Unified Error Codes

Decision: Create HubErrorCode enum with systematic error classification and exception hierarchy.

Error Code Ranges:

Range Type Severity Log Level Stack Trace? Retryable?
1000-1999 Infrastructure Operational WARN No Yes
2000-2999 Client errors Expected INFO No No
5000-5999 Server errors Bugs ERROR Yes No

HubErrorCode Enum:

package org.idempiere.cli.error;

public enum HubErrorCode {

    // ==================== Infrastructure Errors (1000-1999) ====================

    DATABASE_UNREACHABLE(
        1001,
        ErrorSeverity.WARN,
        "Cannot connect to database",
        "Cannot connect to database at %s:%d",
        "Check IDEMPIERE_DB_HOST and IDEMPIERE_DB_PORT environment variables"
    ),

    LLM_PROVIDER_UNAVAILABLE(
        1002,
        ErrorSeverity.WARN,
        "LLM provider is not available",
        "LLM provider '%s' not available - check that the service is running",
        "For Ollama: ensure service is running (ollama serve)"
    ),

    LLM_REQUEST_TIMEOUT(
        1003,
        ErrorSeverity.WARN,
        "LLM request timed out",
        "LLM request timed out after %d seconds",
        "Increase timeout or check LLM service health"
    ),

    RAG_VECTORDB_UNAVAILABLE(
        1004,
        ErrorSeverity.WARN,
        "RAG vector database is not available",
        "RAG vector database (pgvector) is not available",
        "Install pgvector: CREATE EXTENSION vector"
    ),

    // ==================== Client Errors (2000-2999) ====================

    INVALID_ARGUMENTS(
        2001,
        ErrorSeverity.INFO,
        "Invalid arguments provided",
        "Invalid arguments: %s",
        "Check command syntax: idempiere-hub help <command>"
    ),

    TABLE_NOT_FOUND(
        2002,
        ErrorSeverity.INFO,
        "Table not found in database",
        "Table '%s' not found in database",
        "Create the table first with 'dict table add'"
    ),

    TABLE_ALREADY_EXISTS(
        2003,
        ErrorSeverity.INFO,
        "Table already exists",
        "Table '%s' already exists in database",
        "Use 'dict table update' to modify existing table"
    ),

    PERMISSION_DENIED(
        2004,
        ErrorSeverity.INFO,
        "Permission denied",
        "Permission denied: %s",
        "Check agent boundaries and user permissions"
    ),

    // ==================== Server Errors (5000-5999) ====================

    INTERNAL_ERROR(
        5000,
        ErrorSeverity.ERROR,
        "An unexpected error occurred",
        "Internal error: %s",
        "Check logs for details and report to maintainers"
    ),

    NULL_POINTER_ERROR(
        5001,
        ErrorSeverity.ERROR,
        "Null pointer exception",
        "Null pointer at %s",
        "This is a bug - please report with stack trace"
    );

    private final int code;
    private final ErrorSeverity severity;
    private final String description;
    private final String messageTemplate;
    private final String suggestion;

    public String format(Object... args) {
        return String.format(messageTemplate, args);
    }

    /**
     * Log this error with appropriate severity.
     * Infrastructure/Client errors: message only
     * Server errors: full stack trace
     */
    public void log(Logger logger, Object... args) {
        String message = format(args);
        switch (severity) {
            case ERROR -> logger.errorf("[%d] %s", code, message);
            case WARN -> logger.warnf("[%d] %s", code, message);
            case INFO -> logger.infof("[%d] %s", code, message);
        }
    }

    public void log(Logger logger, Throwable cause, Object... args) {
        String message = format(args);
        switch (severity) {
            case ERROR -> logger.errorf(cause, "[%d] %s", code, message);
            case WARN, INFO -> logger.warnf("[%d] %s - %s", code, message, cause.getMessage());
        }
    }
}

enum ErrorSeverity {
    ERROR,  // Unexpected failures, bugs → Full stack trace
    WARN,   // Operational issues (LLM down) → Message only
    INFO    // Client errors (invalid args) → Message only
}

Exception Hierarchy:

// Base exception
public abstract class HubException extends Exception {
    private final HubErrorCode errorCode;
    private final boolean retryable;

    public void log(Logger logger) {
        if (getCause() != null) {
            errorCode.log(logger, getCause(), getMessage());
        } else {
            errorCode.log(logger, getMessage());
        }
    }

    public ToolResult toToolResult() {
        return ToolResult.error(getMessage(), errorCode.name());
    }
}

// Operational errors (infrastructure down, timeouts) - WARN, retryable
public class OperationalException extends HubException {
    public OperationalException(HubErrorCode errorCode, String message) {
        super(errorCode, message, true);
    }

    public static OperationalException llmUnavailable(String provider) {
        return new OperationalException(
            HubErrorCode.LLM_PROVIDER_UNAVAILABLE,
            String.format("LLM provider '%s' not available", provider)
        );
    }
}

// Client errors (invalid input, not found) - INFO, not retryable
public class ClientException extends HubException {
    public ClientException(HubErrorCode errorCode, String message) {
        super(errorCode, message, false);
    }

    public static ClientException tableNotFound(String tableName) {
        return new ClientException(
            HubErrorCode.TABLE_NOT_FOUND,
            String.format("Table '%s' not found", tableName)
        );
    }
}

// Server errors (bugs, NPE) - ERROR, not retryable
public class ServerException extends HubException {
    public ServerException(HubErrorCode errorCode, String message, Throwable cause) {
        super(errorCode, message, cause);
    }

    public static ServerException internalError(String message, Throwable cause) {
        return new ServerException(
            HubErrorCode.INTERNAL_ERROR,
            message,
            cause
        );
    }
}

Phase 3: Refactor Error Handling

Decision: Replace ad-hoc RuntimeException catch-all blocks with typed exceptions.

Before (Current):

try {
    String response = chatAgent.chat(userMessage);
} catch (Exception e) {
    LOG.errorf(e, "Agent execution failed for request %s", requestId);
    throw new RuntimeException("Agent execution failed: " + e.getMessage(), e);
}

Issues:

After (Systematic):

private ChatResponse executeChat(...)
        throws OperationalException, ServerException {
    try {
        String response = chatAgent.chat(userMessage);
        // ... build response

    } catch (CompletionException e) {
        // Timeout - operational error
        if (e.getCause() instanceof TimeoutException) {
            throw new OperationalException(
                HubErrorCode.LLM_REQUEST_TIMEOUT,
                String.format("LLM request timed out after %d seconds", requestTimeoutSeconds)
            );
        }

        // Connection refused - operational error
        if (e.getCause() instanceof ProcessingException) {
            throw OperationalException.llmUnavailable(routing.provider());
        }

        // Unexpected error - server error
        throw ServerException.internalError("LLM request failed", e.getCause());
    }
}

At API Boundary:

@POST
public Response chat(ChatRequest request) {
    try {
        ChatResponse response = agentService.chat(request, context);
        return Response.ok(response).build();

    } catch (OperationalException e) {
        e.log(LOG);  // WARN: [1002] LLM provider 'ollama' not available...
        return Response.status(503).entity(e.toToolResult()).build();

    } catch (ClientException e) {
        e.log(LOG);  // INFO: [2002] Table 'Foo' not found...
        return Response.status(400).entity(e.toToolResult()).build();

    } catch (ServerException e) {
        e.log(LOG);  // ERROR with full stack trace
        return Response.status(500).entity(e.toToolResult()).build();
    }
}

Log Output:

WARN  [1002] LLM provider 'ollama' not available - check that the service is running

Instead of 100+ line stack trace.

Refactor Targets:


Phase 4: Structured Logging (Optional)

Decision: Enable JSON logging for production with error code metrics.

Configuration:

# Production profile
%prod.quarkus.log.console.json=true
%prod.quarkus.log.console.json.pretty-print=false
%prod.quarkus.log.console.json.additional-field."environment".value=production
%prod.quarkus.log.console.json.additional-field."service".value=idempiere-hub

Example JSON Log:

{
  "timestamp": "2025-12-15T10:30:00.123Z",
  "level": "WARN",
  "logger": "org.idempiere.cli.chatapi.agent.ChatAgentService",
  "message": "[1002] LLM provider 'ollama' not available - check that the service is running",
  "error_code": 1002,
  "provider": "ollama",
  "request_id": "abc123",
  "environment": "production",
  "service": "idempiere-hub"
}

Benefits:


Consequences

Positive

  1. Single Source of Truth: Configuration no longer duplicated (Anthropic 3x → 1x)
  2. Type Safety: Build-time validation prevents runtime config errors
  3. Consistent Error Codes: All interfaces (CLI, Chat API, MCP) use same codes
  4. Clean Logs: 80% reduction in log volume (no stack traces for operational errors)
  5. Clear Error Classification: WARN (operational) vs ERROR (bugs)
  6. Better HTTP Mapping: Operational (503), Client (400), Server (500)
  7. LLM-Friendly Errors: ToolResult includes error codes for AI context
  8. Production Ready: JSON logs, metrics, distributed tracing

Negative

  1. Migration Effort: 4-6 weeks total across all phases
  2. Breaking Changes: Exception signatures change (checked exceptions)
  3. Learning Curve: Team needs to learn new error classification
  4. External Impact: Deprecated facades (RegistryToolLogic, DatabaseQueryTool) for external consumers

Risks

  1. Scope Creep: Phase 3 touches 25+ files (services + tools)
  2. Test Coverage: Need comprehensive tests for error handling paths
  3. Backward Compatibility: External MCP/LangChain4j consumers may depend on current API

Mitigation

  1. Phased Rollout: Each phase is independent and testable
  2. Deprecation Period: Keep old patterns deprecated for 1-2 releases
  3. Documentation: Update USER_GUIDE.md with error code reference
  4. Testing: Add integration tests for each error scenario

Implementation Plan

Phase 1: Configuration Consolidation (Week 1)

Tasks:

Success Metrics:


Phase 2: Unified Error Codes (Week 2)

Tasks:

Success Metrics:


Phase 3: Refactor Error Handling (Week 3-4) ✅ COMPLETE

Step 1: ChatAgentService (Priority 1) ✅

Step 2: Services (Priority 2) ✅

Step 3: Commands & Additional Packages ✅

Step 4: ExceptionMappers ✅

Success Metrics: ✅ ALL ACHIEVED

Completion Date: December 15, 2025 Migration Method: Manual + 3 automated batch migrations (Python scripts)


Phase 4: Structured Logging (Week 5) ✅ COMPLETE

Tasks:

Success Metrics:

Completion Date: December 15, 2025 Implementation: Quarkus extensions + 5 custom observability classes


Alternatives Considered

Alternative 1: Keep Current Approach (Rejected)

Pros:

Cons:

Rejection Reason: Technical debt compounds over time, making future changes harder


Alternative 2: Partial Migration (Rejected)

Pros:

Cons:

Rejection Reason: Half-measures create more problems than they solve


Alternative 3: Big Bang Migration (Rejected)

Pros:

Cons:

Rejection Reason: Phased approach is safer and more manageable


Success Metrics

Phase 1 (Configuration)

Phase 2-3 (Error Handling)

Phase 4 (Observability)


References

Existing Error Code Implementations:

Immediate Fix (Completed):


Approval

Proposed By: quarkus-reactive-architect expert Reviewed By: Pending Approved By: Pending Implementation Start: TBD


Updates

Path: /docs/developers/architecture/idempiere-hub/058-logging-and-configuration-architecture