Architecture Decision Record: RAG Implementation for iDempiere Documentation

ADR-001: Retrieval-Augmented Generation for Editor.js Documentation

Status: Proposed
Date: 2024-12-11
Authors: CloudEmpiere Technical Team
Deciders: Norbert (CEO), Development Team


Context

CloudEmpiere maintains extensive iDempiere ERP documentation stored as Editor.js JSON format in PostgreSQL. Current challenges:

Goal: Enable natural language Q&A over documentation corpus using AI-powered semantic search.


Decision

Implement Retrieval-Augmented Generation (RAG) system using:

Technology Stack

Component Technology Justification
Framework Quarkus Existing stack, native performance
RAG Library LangChain4j 0.35+ Java-native, Quarkus integration, active development
Vector Database PostgreSQL + pgvector Existing infrastructure, no new DB needed
Embedding Model all-MiniLM-L6-v2 (local) 384 dims, fast, free, good Slovak support
LLM Claude Sonnet 4.5 Best reasoning, handles Slovak/English, structured output
Chunking Strategy Semantic block-based Preserves Editor.js structure

Architecture

┌─────────────────────────────────────────────────────────────┐
│ USER INTERFACE                                              │
│  • Admin Panel: Document ingestion UI                       │
│  • Client Portal: Q&A Chat Interface                        │
│  • API: POST /api/rag/query                                │
└─────────────────────────────────────────────────────────────┘
                          │
                          ▼
┌─────────────────────────────────────────────────────────────┐
│ QUARKUS APPLICATION LAYER                                   │
│                                                             │
│  ┌──────────────────┐  ┌──────────────────┐               │
│  │ RagQueryService  │  │ IngestionService │               │
│  └──────────────────┘  └──────────────────┘               │
│           │                      │                          │
│           │                      ▼                          │
│           │           ┌──────────────────┐                 │
│           │           │ EditorJsChunker  │                 │
│           │           └──────────────────┘                 │
│           │                      │                          │
│           └──────────────────────┴────────┐                │
│                                            ▼                │
│                              ┌────────────────────────┐    │
│                              │ LangChain4j Services  │    │
│                              │ • EmbeddingStore      │    │
│                              │ • ContentRetriever    │    │
│                              │ • AiServices          │    │
│                              └────────────────────────┘    │
└─────────────────────────────────────────────────────────────┘
                          │
                          ▼
┌─────────────────────────────────────────────────────────────┐
│ DATA LAYER                                                  │
│                                                             │
│  ┌────────────────────────┐  ┌────────────────────────┐   │
│  │ documents              │  │ embeddings             │   │
│  │ • id (PK)             │  │ • id (UUID, PK)        │   │
│  │ • title               │  │ • content (TEXT)       │   │
│  │ • editor_json (JSONB) │  │ • vector (VECTOR)      │   │
│  │ • created_at          │  │ • metadata (JSONB)     │   │
│  │ • updated_at          │  │   - document_id        │   │
│  └────────────────────────┘  │   - chunk_index        │   │
│                              │   - section_title      │   │
│                              │   - block_types        │   │
│                              └────────────────────────┘   │
│                                                             │
│  Indexes:                                                   │
│  • IVFFLAT on embeddings.vector (cosine similarity)        │
│  • GIN on documents.editor_json                            │
└─────────────────────────────────────────────────────────────┘
                          │
                          ▼
┌─────────────────────────────────────────────────────────────┐
│ EXTERNAL SERVICES                                           │
│  • Anthropic API (Claude)                                   │
└─────────────────────────────────────────────────────────────┘

Chunking Strategy

Parent Document Retrieval with semantic block-based chunking:

// Small chunks for precise retrieval
Chunk = {
  text: "Step 1: Navigate to Warehouse window...",
  vector: [0.234, -0.891, ...],
  metadata: {
    document_id: 123,
    section_id: "warehouse-setup",
    section_title: "Initial Warehouse Configuration",
    parent_content: "[FULL SECTION TEXT]",  // 2-3 pages
    chunk_index: 1,
    block_types: "header,paragraph,list"
  }
}

Chunking Rules:

Retrieval Flow:

  1. User asks question → embed query
  2. Find top 5 relevant small chunks (fast, precise)
  3. Extract parent sections (deduplicated)
  4. Send full sections to Claude (complete context)

Implementation Plan

Phase 1: Infrastructure (Week 1)

Phase 2: Core Services (Week 2)

Phase 3: API Layer (Week 3)

Phase 4: UI Integration (Week 4)

Phase 5: Optimization (Week 5-6)


Configuration

application.properties:

# Database
quarkus.datasource.jdbc.url=jdbc:postgresql://localhost:5432/cloudempiere
quarkus.datasource.username=${DB_USER}
quarkus.datasource.password=${DB_PASSWORD}

# LangChain4j - pgvector
quarkus.langchain4j.pgvector.datasource=<default>
quarkus.langchain4j.pgvector.table-name=embeddings
quarkus.langchain4j.pgvector.dimension=384
quarkus.langchain4j.pgvector.create-table=true
quarkus.langchain4j.pgvector.drop-table-first=false

# Embedding Model (Local)
quarkus.langchain4j.embedding-model.provider=all-minilm-l6-v2

# LLM - Claude
quarkus.langchain4j.anthropic.api-key=${ANTHROPIC_API_KEY}
quarkus.langchain4j.anthropic.chat-model.model-name=claude-sonnet-4-20250514
quarkus.langchain4j.anthropic.chat-model.max-tokens=4096
quarkus.langchain4j.anthropic.timeout=60s

# RAG Configuration
rag.retrieval.max-results=5
rag.retrieval.min-score=0.6
rag.chunk.max-size=1000
rag.chunk.overlap=200

Consequences

Positive

✅ User Experience

✅ Operational Efficiency

✅ Technical Benefits

✅ Business Value

Negative

⚠️ Costs

⚠️ Complexity

⚠️ Limitations

Risks & Mitigations

Risk Impact Mitigation
Poor answer quality High Implement feedback loop, tune retrieval params, improve chunking
High API costs Medium Set rate limits, cache common queries, monitor usage
Outdated embeddings Medium Automated re-indexing on doc updates (webhook)
Slovak language quality Medium Test extensively, provide fallback to English docs
Vector index performance Low IVFFLAT index, query optimization, consider disk-ANN if needed

Alternatives Considered

2. OpenAI Embeddings (text-embedding-3-small)

3. Fine-tuned Model on iDempiere Docs

4. Pinecone / Weaviate (Dedicated Vector DB)

5. Simple Character-Based Chunking


Success Metrics

Phase 1 (MVP - 3 months):

Phase 2 (6 months):

Long-term (12 months):


Monitoring & Observability

Metrics to Track:

Logging:

@Logged
public String query(String question) {
    log.info("RAG query: {}", question);
    long start = System.currentTimeMillis();
    
    // ... query logic
    
    log.info("RAG response time: {}ms, chunks: {}, relevance: {}", 
        System.currentTimeMillis() - start, 
        retrievedChunks.size(),
        avgRelevanceScore
    );
}

References


Decision

APPROVED - Proceed with implementation

Signature:
Norbert (CEO) - [Date]


Appendix A: Maven Dependencies

<!-- LangChain4j Core -->
<dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j</artifactId>
    <version>0.35.0</version>
</dependency>

<!-- Quarkus LangChain4j Integration -->
<dependency>
    <groupId>io.quarkiverse.langchain4j</groupId>
    <artifactId>quarkus-langchain4j-core</artifactId>
    <version>0.20.0</version>
</dependency>

<!-- pgvector Store -->
<dependency>
    <groupId>io.quarkiverse.langchain4j</groupId>
    <artifactId>quarkus-langchain4j-pgvector</artifactId>
    <version>0.20.0</version>
</dependency>

<!-- Embeddings (Local) -->
<dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j-embeddings-all-minilm-l6-v2</artifactId>
    <version>0.35.0</version>
</dependency>

<!-- Anthropic Claude -->
<dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j-anthropic</artifactId>
    <version>0.35.0</version>
</dependency>

Appendix B: Database Schema

-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;

-- Embeddings table (managed by LangChain4j)
CREATE TABLE IF NOT EXISTS embeddings (
    id UUID PRIMARY KEY,
    content TEXT NOT NULL,
    content_vector VECTOR(384),
    metadata JSONB
);

-- Indexes for fast similarity search
CREATE INDEX embeddings_vector_idx 
ON embeddings 
USING ivfflat (content_vector vector_cosine_ops) 
WITH (lists = 100);

-- Index for metadata filtering
CREATE INDEX embeddings_metadata_idx 
ON embeddings 
USING gin (metadata);

-- Add indexing timestamp to existing documents table
ALTER TABLE documents 
ADD COLUMN IF NOT EXISTS last_indexed_at TIMESTAMP;

-- Index for efficient document lookup
CREATE INDEX IF NOT EXISTS documents_last_indexed_idx 
ON documents (last_indexed_at);

Appendix C: Service Interfaces

package sk.cloudempiere.rag;

import dev.langchain4j.data.segment.TextSegment;
import java.util.List;

/**
 * Main RAG service interface
 */
public interface RagService {
    /**
     * Query the RAG system with a natural language question
     * @param question User's question in Slovak or English
     * @return AI-generated answer based on documentation
     */
    String query(String question);
    
    /**
     * Ingest a single document into the vector store
     * @param documentId Database ID of the document
     */
    void ingestDocument(Long documentId);
    
    /**
     * Batch ingest all documents from the database
     */
    void ingestAllDocuments();
    
    /**
     * Re-index a document (after updates)
     * @param documentId Database ID of the document
     */
    void reindexDocument(Long documentId);
}

/**
 * Chunking service for Editor.js documents
 */
public interface EditorJsChunker {
    /**
     * Convert Editor.js JSON into semantic chunks
     * @param docId Document database ID
     * @param title Document title
     * @param editorJson Editor.js JSON string
     * @return List of text segments with metadata
     */
    List<TextSegment> chunkDocument(Long docId, String title, String editorJson);
}

Appendix D: Example REST API Usage

Ingest Document

curl -X POST http://localhost:8080/api/rag/ingest/123 \
  -H "Content-Type: application/json" \
  -d '{
    "title": "Warehouse Setup Guide",
    "editor_json": "{\"blocks\":[...]}"
  }'

Response:

{
  "status": "ingested",
  "document_id": "123",
  "chunks_created": 15
}

Query RAG System

curl -X POST http://localhost:8080/api/rag/query \
  -H "Content-Type: application/json" \
  -d '{
    "question": "Ako nastavím nový sklad v iDempiere?"
  }'

Response:

{
  "question": "Ako nastavím nový sklad v iDempiere?",
  "answer": "Pre vytvorenie nového skladu v iDempiere postupujte nasledovne:\n\n1. Otvorte okno Material Management > Warehouse and Locators\n2. Kliknite na tlačidlo New\n3. Vyplňte základné údaje skladu...",
  "sources": [
    {
      "document_id": 123,
      "title": "Warehouse Setup Guide",
      "relevance_score": 0.89
    }
  ],
  "response_time_ms": 1247
}

Appendix E: Phased Rollout Strategy

Week 1-2: Internal Testing

Week 3-4: Beta Testing

Week 5-6: Production Launch

Week 7-8: Optimization


End of ADR-001


Document Metadata:

Path: /docs/developers/architecture/idempiere-hub/ADR-001-RAG-Implementation-REVIEW