ADR-001 RAG Implementation Review: Gap Analysis & Recommendations

Date: 2025-12-11 Context: Analysis of current K_Entry/RAG implementation against ADR-001 best practices Status: In Progress


Executive Summary

This document analyzes our current K_Entry RAG implementation against the recommendations in ADR-001 (RAG Implementation for iDempiere Documentation). We identify gaps, validate current approaches, and provide actionable recommendations.

Key Findings:


Comparison Matrix

Component ADR-001 Recommendation Current Implementation Status Gap Analysis
Framework Quarkus ✅ Quarkus 3.27.1 ✅ PASS Aligned
RAG Library LangChain4j 0.35+ ✅ LangChain4j via Quarkus extension ✅ PASS Using Quarkus-managed version
Vector DB PostgreSQL + pgvector ✅ PostgreSQL + pgvector (separate DB) ✅ PASS Better: Using dedicated vector DB
Embedding Model all-MiniLM-L6-v2 (384 dims) ✅ mxbai-embed-large (1024 dims) ✅ PASS Better: Multilingual support (SK/HU/EN)
LLM Claude Sonnet 4.5 ✅ Claude Sonnet 4.5 (via Anthropic) ✅ PASS Aligned
Chunking Strategy Semantic block-based with parent retrieval ⚠️ Simple recursive splitter (500 chars) ❌ CRITICAL GAP Missing semantic structure
Chunk Size Max 1000 chars ⚠️ Max 500 chars ⚠️ SUBOPTIMAL Too small, loses context
Parent Document Store full section (2-3 pages) ❌ No parent reference ❌ CRITICAL GAP Missing hierarchical retrieval
Metadata Enrichment section_title, block_types, parent_content ✅ Partial (breadcrumb, format, language) ⚠️ PARTIAL Missing block_types, parent_content
Multi-language Slovak + English ✅ en_US, sk_SK, hu_HU ✅ PASS Better: 3 languages

Critical Gaps

1. Chunking Strategy ❌ CRITICAL

ADR-001 Recommendation:

Parent Document Retrieval with semantic block-based chunking:
- Split on H1/H2 headers (major sections)
- Keep related blocks together (paragraph + list + code)
- Small chunks (1000 chars) for retrieval
- Store full parent section (2-5 chunks → 1 parent)
- Preserve block type metadata

Current Implementation:

// KEntryIngestor.java:212
DocumentSplitter splitter = DocumentSplitters.recursive(
    ragConfig.getMaxSegmentSize(),  // 500 chars
    ragConfig.getMaxOverlapSize()   // 50 chars
);

Problems:

  1. Character-based, not semantic: Breaks mid-sentence, mid-paragraph, mid-list
  2. No structure awareness: Doesn't understand Editor.js block types (header, paragraph, list, code)
  3. Too small (500 chars): ADR-001 recommends 1000 chars minimum
  4. No parent reference: Can't retrieve full context after finding relevant chunk

Impact:

Recommendation: ADOPT - Implement semantic block-based chunking (see Section 4)


2. Parent Document Retrieval ❌ CRITICAL

ADR-001 Recommendation:

Chunk metadata = {
  text: "Step 1: Navigate...",
  vector: [...],
  metadata: {
    parent_content: "[FULL SECTION TEXT]",  // 2-3 pages
    section_id: "warehouse-setup",
    section_title: "Initial Warehouse Configuration"
  }
}

// Retrieval flow:
1. Find top 5 small chunks (precise)
2. Extract parent sections (deduplicated)
3. Send full sections to Claude (complete context)

Current Implementation:

// KEntryIngestor.java:289-310
// Splits document into segments
// Each segment is standalone - NO parent reference
List<TextSegment> segments = splitter.split(doc);

Problems:

  1. No parent section: Each chunk is isolated
  2. No hierarchical context: Can't reconstruct full procedure from fragment
  3. Deduplication missing: Multiple chunks from same section = redundant context to Claude

Impact:

Recommendation: ADOPT - Implement parent document retrieval (see Section 4)


What We're Doing Better

1. Multilingual Embedding Model ✅

ADR-001: all-MiniLM-L6-v2 (384 dims, good Slovak support) Our Implementation: mxbai-embed-large (1024 dims, better multilingual)

Why Better:

Verdict: KEEP - Our choice is superior


2. Dedicated Vector Database ✅

ADR-001: Use same PostgreSQL as application Our Implementation: Separate PostgreSQL database for vectors

Why Better:

Configuration:

# application.properties:175-183 (iDempiere DB)
quarkus.datasource.jdbc.url=jdbc:postgresql://localhost:5433/idempiere

# application.properties:194-199 (Vector DB)
quarkus.datasource.vector.jdbc.url=jdbc:postgresql://localhost:5432/vector

Verdict: KEEP - Our architecture is superior


3. Hierarchical Breadcrumb Metadata ✅

ADR-001: Store section_title metadata Our Implementation: Store full breadcrumb path via recursive CTE

Example:

-- KEntryIngestor.java:77-141 (Recursive CTE builds breadcrumb)
breadcrumb: "Cloudempiere ERP > Warehouse Management > Initial Setup > Create Location"

Metadata:

doc.metadata().put("breadcrumb", breadcrumb);  // Full path
doc.metadata().put("tree_level", treeLevel);   // Depth
doc.metadata().put("knowledge_base", typeName); // Root KB

Why Better:

Verdict: KEEP - Our metadata is richer


4. Content Format Normalization (ADR-040) ✅

ADR-001: No mention of format handling Our Implementation: Convert all formats (BLK/HTM/GFM) to Markdown before embedding

Implementation:

// KEntryIngestor.java:422-460
private String convertToMarkdown(String textMsg, String editMode, int kEntryId) {
    return switch (editMode) {
        case "BLK" -> editorJsConverter.toMarkdown(textMsg);
        case "HTM" -> htmlConverter.toMarkdown(textMsg);
        case "GFM" -> textMsg; // Already Markdown
        default -> textMsg;
    };
}

Why Better:

Verdict: KEEP - Critical enhancement not in ADR-001


Recommendations

ADOPT ✅ (Implement from ADR-001)

1. Semantic Block-Based Chunking

Rationale: Editor.js stores content as structured blocks. We should respect this structure instead of character-based splitting.

Implementation Strategy:

/**
 * Semantic chunker for K_Entry Editor.js/Markdown content.
 * Splits on headers, preserves block structure, stores parent sections.
 */
public class KEntrySemanticChunker {

    /**
     * Chunk K_Entry content respecting semantic structure.
     *
     * @param content Markdown content (after format conversion)
     * @param metadata Entry metadata (breadcrumb, tree_level, etc.)
     * @return List of semantic chunks with parent references
     */
    public List<TextSegment> chunkSemantically(String content, Map<String, String> metadata) {
        List<TextSegment> chunks = new ArrayList<>();

        // Parse Markdown into sections (H1/H2 boundaries)
        List<Section> sections = parseMarkdownSections(content);

        for (Section section : sections) {
            // Each section becomes a "parent document"
            String sectionContent = section.getFullContent();
            String sectionTitle = section.getTitle();

            // Split section into smaller chunks (1000 chars) if needed
            if (sectionContent.length() <= 1000) {
                // Small section - store as single chunk
                chunks.add(createChunk(sectionContent, sectionContent, sectionTitle, 0, metadata));
            } else {
                // Large section - split into chunks, but keep parent reference
                List<String> subChunks = splitByParagraphs(sectionContent, 1000);
                for (int i = 0; i < subChunks.size(); i++) {
                    chunks.add(createChunk(
                        subChunks.get(i),      // Small chunk for retrieval
                        sectionContent,         // Full parent section
                        sectionTitle,           // Section title
                        i,                      // Chunk index
                        metadata
                    ));
                }
            }
        }

        return chunks;
    }

    private TextSegment createChunk(String chunkText, String parentContent,
                                     String sectionTitle, int chunkIndex,
                                     Map<String, String> baseMetadata) {
        TextSegment segment = TextSegment.from(chunkText);
        Metadata meta = segment.metadata();

        // Copy base metadata
        baseMetadata.forEach(meta::put);

        // Add parent document retrieval metadata
        meta.put("parent_content", parentContent);      // FULL SECTION (for Claude)
        meta.put("section_title", sectionTitle);        // Section heading
        meta.put("chunk_index", String.valueOf(chunkIndex));  // Position in section
        meta.put("chunk_size", String.valueOf(chunkText.length()));

        return segment;
    }

    /**
     * Parse Markdown into semantic sections (H1/H2 boundaries).
     */
    private List<Section> parseMarkdownSections(String markdown) {
        // Use flexmark-java or commonmark-java for proper Markdown parsing
        // Split on H1 (# ) and H2 (## )
        // Keep paragraphs, lists, code blocks together under their header
        // ...implementation...
    }

    /**
     * Split long section by paragraphs, respecting 1000 char limit.
     * Keeps related blocks together (paragraph + list stays together).
     */
    private List<String> splitByParagraphs(String sectionContent, int maxSize) {
        // Split on double newlines (paragraph boundaries)
        // Group consecutive blocks until reaching maxSize
        // Never break mid-paragraph, mid-list, mid-code-block
        // ...implementation...
    }
}

Changes Required:

  1. Add dependency:
<!-- Markdown parser for semantic chunking -->
<dependency>
    <groupId>com.vladsch.flexmark</groupId>
    <artifactId>flexmark-all</artifactId>
    <version>0.64.8</version>
</dependency>
  1. Update KEntryIngestor.java:
// Replace DocumentSplitters.recursive() with semantic chunker
@Inject
KEntrySemanticChunker semanticChunker;

// In ingest() method:
List<TextSegment> segments = semanticChunker.chunkSemantically(
    markdownContent,
    Map.of(
        "source_type", SOURCE_TYPE,
        "k_entry_id", String.valueOf(kEntryId),
        "breadcrumb", breadcrumb,
        "knowledge_base", typeName,
        "ad_language", language
    )
);
  1. Update configuration:
# Increase chunk size to ADR-001 recommendation
idempiere.cli.rag.splitter.max-segment-size=1000
idempiere.cli.rag.splitter.max-overlap-size=100

Benefits:


2. Parent Document Retrieval in Query Flow

Rationale: Small chunks for retrieval, full sections for Claude.

Implementation Strategy:

/**
 * Enhanced RAG retrieval with parent document strategy.
 */
public class RagQueryService {

    @Inject
    EmbeddingStore<TextSegment> embeddingStore;

    @Inject
    EmbeddingModel embeddingModel;

    /**
     * Query RAG with parent document retrieval.
     *
     * 1. Find top K small chunks (precise retrieval)
     * 2. Extract parent sections (deduplicated)
     * 3. Send full sections to Claude
     */
    public String query(String question, int maxResults) {
        // 1. Embed question
        Embedding queryEmbedding = embeddingModel.embed(question).content();

        // 2. Find top K relevant chunks (small, precise)
        EmbeddingSearchRequest searchRequest = EmbeddingSearchRequest.builder()
            .queryEmbedding(queryEmbedding)
            .maxResults(maxResults * 2)  // Get more chunks to dedupe parents
            .minScore(0.7)
            .build();

        EmbeddingSearchResult<TextSegment> searchResult = embeddingStore.search(searchRequest);

        // 3. Extract parent sections (deduplicated)
        List<String> parentSections = extractParentSections(searchResult.matches());

        // 4. Build prompt with full parent sections
        String prompt = buildPrompt(question, parentSections);

        // 5. Send to Claude
        return aiService.chat(prompt);
    }

    /**
     * Extract parent sections from chunks, deduplicated.
     */
    private List<String> extractParentSections(List<EmbeddingMatch<TextSegment>> matches) {
        Set<String> seenSections = new HashSet<>();
        List<String> parentSections = new ArrayList<>();

        for (EmbeddingMatch<TextSegment> match : matches) {
            Metadata meta = match.embedded().metadata();
            String parentContent = meta.get("parent_content");
            String sectionTitle = meta.get("section_title");

            // Use section title for deduplication
            if (parentContent != null && seenSections.add(sectionTitle)) {
                parentSections.add(formatSection(parentContent, meta));
            }

            // Limit to top 3-5 parent sections (ADR-001 recommendation)
            if (parentSections.size() >= 5) break;
        }

        return parentSections;
    }

    private String formatSection(String content, Metadata meta) {
        StringBuilder formatted = new StringBuilder();

        // Add breadcrumb for context
        formatted.append("Source: ").append(meta.get("breadcrumb")).append("\n");
        formatted.append("Section: ").append(meta.get("section_title")).append("\n\n");
        formatted.append(content).append("\n\n");
        formatted.append("---\n\n");

        return formatted.toString();
    }
}

Benefits:


REJECT ❌ (ADR-001 recommendations we should not adopt)

1. Smaller Embedding Model (all-MiniLM-L6-v2, 384 dims)

ADR-001 Recommendation: all-MiniLM-L6-v2 (384 dimensions) Our Implementation: mxbai-embed-large (1024 dimensions)

Rejection Rationale:

Verdict: REJECT - Keep mxbai-embed-large


2. Single Database for Application + Vectors

ADR-001 Recommendation: Use same PostgreSQL database for app + vectors Our Implementation: Separate databases (iDempiere DB + Vector DB)

Rejection Rationale:

Verdict: REJECT - Keep separate databases


3. IVFFLAT Index with Fixed List Count

ADR-001 Recommendation:

CREATE INDEX embeddings_vector_idx
ON embeddings
USING ivfflat (content_vector vector_cosine_ops)
WITH (lists = 100);

Our Implementation: Let Quarkus LangChain4j manage index creation

Rejection Rationale:

Verdict: REJECT - Let Quarkus LangChain4j handle indexing, revisit when we reach 10k+ embeddings


Implementation Priority

Priority Task Effort Impact Timeline
P0 Implement semantic block-based chunking 2-3 days High Week 1
P0 Add parent document metadata to chunks 1 day High Week 1
P1 Update query service for parent retrieval 1-2 days High Week 1
P1 Update chunk size config (500→1000) 10 min Medium Week 1
P2 Add flexmark-java dependency 10 min Medium Week 1
P2 Re-ingest K_Entry with new chunking 30 min Medium Week 2
P3 Add monitoring for chunk size distribution 1 day Low Week 2
P3 Document new chunking strategy in ADR 2 hours Low Week 2

Testing Strategy

1. Unit Tests

@Test
void chunkSemantically_preservesStructure() {
    String markdown = """
        # Warehouse Setup

        Follow these steps:

        1. Open Warehouse window
        2. Click New
        3. Fill in details:
           - Name: Main Warehouse
           - Locator: A-01-01

        ## Important Notes

        Always validate warehouse before saving.
        """;

    List<TextSegment> chunks = chunker.chunkSemantically(markdown, Map.of());

    // Should create 2 chunks (one per H1/H2 section)
    assertThat(chunks).hasSize(2);

    // First chunk should contain full list (not break mid-list)
    assertThat(chunks.get(0).text()).contains("1. Open Warehouse window");
    assertThat(chunks.get(0).text()).contains("3. Fill in details:");
    assertThat(chunks.get(0).text()).contains("- Locator: A-01-01");

    // Parent content should be full section
    assertThat(chunks.get(0).metadata().get("parent_content")).contains("# Warehouse Setup");
    assertThat(chunks.get(0).metadata().get("section_title")).isEqualTo("Warehouse Setup");
}

2. Integration Tests

@Test
void query_withParentRetrieval_returnsCoherentAnswer() {
    // Ingest test document with multi-step procedure
    ingestTestDocument("warehouse-setup-guide.md");

    // Query for specific step
    String answer = ragService.query("How do I create a warehouse location?");

    // Answer should include full procedure context (from parent section)
    assertThat(answer).contains("Step 1:");
    assertThat(answer).contains("Step 2:");
    assertThat(answer).contains("Step 3:");
}

3. Quality Metrics

Monitor before/after semantic chunking:


Migration Plan

Phase 1: Implement Semantic Chunking (Week 1)

  1. Add flexmark-java dependency to pom.xml
  2. Create KEntrySemanticChunker.java
  3. Add unit tests for chunker
  4. Update KEntryIngestor to use semantic chunker
  5. Test with sample K_Entry document

Phase 2: Update Query Service (Week 1)

  1. Modify RagQueryService to extract parent sections
  2. Update prompt building to use full sections
  3. Add deduplication logic
  4. Test query flow end-to-end

Phase 3: Re-index Knowledge Base (Week 2)

  1. Clear existing K_Entry embeddings:

    java -jar target/idempiere-hub-runner.jar knowledge clear --source k_entry
    
  2. Re-ingest with new chunking:

    java -jar target/idempiere-hub-runner.jar knowledge ingest --source k_entry --force
    
  3. Validate chunk quality:

    java -jar target/idempiere-hub-runner.jar knowledge status
    
  4. Test queries and compare answer quality

Phase 4: Document & Monitor (Week 2)

  1. Create ADR documenting semantic chunking strategy
  2. Update USER_GUIDE.md with new chunking behavior
  3. Add monitoring for chunk size distribution
  4. Collect user feedback on answer quality

Conclusion

Overall Assessment: Our implementation is 80% aligned with ADR-001 best practices.

Critical Gaps (Must Fix):

  1. ❌ Semantic block-based chunking (currently character-based)
  2. ❌ Parent document retrieval (currently isolated chunks)

What We're Doing Better:

  1. ✅ Multilingual embedding model (mxbai-embed-large > all-MiniLM-L6-v2)
  2. ✅ Separate vector database (better architecture)
  3. ✅ Hierarchical breadcrumb metadata (richer than section_title alone)
  4. ✅ Content format normalization (ADR-040)

Recommended Action: Adopt semantic chunking and parent retrieval from ADR-001. Reject recommendations to downgrade embedding model or merge databases.

Estimated Effort: 4-5 days development + 1 day re-indexing/testing

Expected Impact: Significant improvement in answer quality for multi-step procedures and complex documentation queries.


Next Steps:

  1. Get approval for implementation plan
  2. Create JIRA tickets for P0/P1 tasks
  3. Begin semantic chunker implementation
  4. Schedule re-indexing window (low-traffic period)

Document Metadata:

Path: /docs/developers/architecture/idempiere-hub/ADR-001-REVIEW-ANALYSIS