Architecture Decision Record: RAG Implementation for iDempiere Documentation
ADR-001: Retrieval-Augmented Generation for Editor.js Documentation
Status: Proposed
Date: 2024-12-11
Authors: CloudEmpiere Technical Team
Deciders: Norbert (CEO), Development Team
Context
CloudEmpiere maintains extensive iDempiere ERP documentation stored as Editor.js JSON format in PostgreSQL. Current challenges:
- Documentation Volume: 50-200+ documents, each with 20-500 Editor.js blocks
- User Pain Points:
- Manual search through documentation is time-consuming
- Knowledge scattered across multiple documents
- Support team answers repetitive questions
- New clients need quick access to setup procedures
- Business Impact:
- Support specialists spend 40% time answering documentation questions
- Slower client onboarding
- Missed cross-document insights (e.g., "How does warehouse integrate with invoicing?")
Goal: Enable natural language Q&A over documentation corpus using AI-powered semantic search.
Decision
Implement Retrieval-Augmented Generation (RAG) system using:
Technology Stack
| Component | Technology | Justification |
|---|---|---|
| Framework | Quarkus | Existing stack, native performance |
| RAG Library | LangChain4j 0.35+ | Java-native, Quarkus integration, active development |
| Vector Database | PostgreSQL + pgvector | Existing infrastructure, no new DB needed |
| Embedding Model | all-MiniLM-L6-v2 (local) | 384 dims, fast, free, good Slovak support |
| LLM | Claude Sonnet 4.5 | Best reasoning, handles Slovak/English, structured output |
| Chunking Strategy | Semantic block-based | Preserves Editor.js structure |
Architecture
┌─────────────────────────────────────────────────────────────┐
│ USER INTERFACE │
│ • Admin Panel: Document ingestion UI │
│ • Client Portal: Q&A Chat Interface │
│ • API: POST /api/rag/query │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ QUARKUS APPLICATION LAYER │
│ │
│ ┌──────────────────┐ ┌──────────────────┐ │
│ │ RagQueryService │ │ IngestionService │ │
│ └──────────────────┘ └──────────────────┘ │
│ │ │ │
│ │ ▼ │
│ │ ┌──────────────────┐ │
│ │ │ EditorJsChunker │ │
│ │ └──────────────────┘ │
│ │ │ │
│ └──────────────────────┴────────┐ │
│ ▼ │
│ ┌────────────────────────┐ │
│ │ LangChain4j Services │ │
│ │ • EmbeddingStore │ │
│ │ • ContentRetriever │ │
│ │ • AiServices │ │
│ └────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ DATA LAYER │
│ │
│ ┌────────────────────────┐ ┌────────────────────────┐ │
│ │ documents │ │ embeddings │ │
│ │ • id (PK) │ │ • id (UUID, PK) │ │
│ │ • title │ │ • content (TEXT) │ │
│ │ • editor_json (JSONB) │ │ • vector (VECTOR) │ │
│ │ • created_at │ │ • metadata (JSONB) │ │
│ │ • updated_at │ │ - document_id │ │
│ └────────────────────────┘ │ - chunk_index │ │
│ │ - section_title │ │
│ │ - block_types │ │
│ └────────────────────────┘ │
│ │
│ Indexes: │
│ • IVFFLAT on embeddings.vector (cosine similarity) │
│ • GIN on documents.editor_json │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ EXTERNAL SERVICES │
│ • Anthropic API (Claude) │
└─────────────────────────────────────────────────────────────┘
Chunking Strategy
Parent Document Retrieval with semantic block-based chunking:
// Small chunks for precise retrieval
Chunk = {
text: "Step 1: Navigate to Warehouse window...",
vector: [0.234, -0.891, ...],
metadata: {
document_id: 123,
section_id: "warehouse-setup",
section_title: "Initial Warehouse Configuration",
parent_content: "[FULL SECTION TEXT]", // 2-3 pages
chunk_index: 1,
block_types: "header,paragraph,list"
}
}
Chunking Rules:
- Split on H1/H2 headers (new major sections)
- Keep related blocks together (paragraph + list + code)
- Max chunk size: 1000 characters
- Overlap: Store parent section reference (2-5 chunks → 1 parent section)
- Preserve block type metadata for filtering
Retrieval Flow:
- User asks question → embed query
- Find top 5 relevant small chunks (fast, precise)
- Extract parent sections (deduplicated)
- Send full sections to Claude (complete context)
Implementation Plan
Phase 1: Infrastructure (Week 1)
- [ ] Add LangChain4j dependencies to
pom.xml - [ ] Enable pgvector extension in PostgreSQL
- [ ] Create
embeddingstable with IVFFLAT index - [ ] Configure embedding model (all-MiniLM-L6-v2)
- [ ] Configure Anthropic API connection
Phase 2: Core Services (Week 2)
- [ ] Implement
EditorJsChunkerwith parent section logic - [ ] Create
DocumentIngestionService - [ ] Build
RagQueryServicewith LangChain4j integration - [ ] Add question type classification (step-by-step vs. factual)
Phase 3: API Layer (Week 3)
- [ ] REST endpoints:
/api/rag/ingest,/api/rag/query - [ ] Batch ingestion for existing documents
- [ ] Error handling and validation
Phase 4: UI Integration (Week 4)
- [ ] Admin panel: Trigger document re-indexing
- [ ] Client portal: Chat interface for Q&A
- [ ] Support dashboard: View common questions
Phase 5: Optimization (Week 5-6)
- [ ] Tune retrieval parameters (K, similarity threshold)
- [ ] Add query caching for common questions
- [ ] Implement feedback loop (thumbs up/down)
- [ ] Monitor embedding quality
Configuration
application.properties:
# Database
quarkus.datasource.jdbc.url=jdbc:postgresql://localhost:5432/cloudempiere
quarkus.datasource.username=${DB_USER}
quarkus.datasource.password=${DB_PASSWORD}
# LangChain4j - pgvector
quarkus.langchain4j.pgvector.datasource=<default>
quarkus.langchain4j.pgvector.table-name=embeddings
quarkus.langchain4j.pgvector.dimension=384
quarkus.langchain4j.pgvector.create-table=true
quarkus.langchain4j.pgvector.drop-table-first=false
# Embedding Model (Local)
quarkus.langchain4j.embedding-model.provider=all-minilm-l6-v2
# LLM - Claude
quarkus.langchain4j.anthropic.api-key=${ANTHROPIC_API_KEY}
quarkus.langchain4j.anthropic.chat-model.model-name=claude-sonnet-4-20250514
quarkus.langchain4j.anthropic.chat-model.max-tokens=4096
quarkus.langchain4j.anthropic.timeout=60s
# RAG Configuration
rag.retrieval.max-results=5
rag.retrieval.min-score=0.6
rag.chunk.max-size=1000
rag.chunk.overlap=200
Consequences
Positive
✅ User Experience
- Instant answers to documentation questions
- Natural language queries (Slovak + English)
- Cross-document insights automatically
✅ Operational Efficiency
- Reduce support team workload by 30-40%
- Faster client onboarding
- 24/7 availability
✅ Technical Benefits
- Leverages existing PostgreSQL infrastructure
- No additional database maintenance
- Local embedding model = low latency, no external API costs
- Scales to 1000+ documents
✅ Business Value
- Competitive differentiator for CloudEmpiere SaaS
- Reduces time-to-value for new clients
- Enables self-service support tier
Negative
⚠️ Costs
- Anthropic API: ~$0.003 per query (Sonnet 4.5)
- Estimated monthly: €50-150 (500-5000 queries)
- Embedding model: Free (local)
⚠️ Complexity
- New technology stack (LangChain4j)
- Vector search tuning required
- Document re-indexing on updates
⚠️ Limitations
- Answers only as good as documentation quality
- May miss implicit knowledge not in docs
- Requires periodic re-indexing (weekly/monthly)
Risks & Mitigations
| Risk | Impact | Mitigation |
|---|---|---|
| Poor answer quality | High | Implement feedback loop, tune retrieval params, improve chunking |
| High API costs | Medium | Set rate limits, cache common queries, monitor usage |
| Outdated embeddings | Medium | Automated re-indexing on doc updates (webhook) |
| Slovak language quality | Medium | Test extensively, provide fallback to English docs |
| Vector index performance | Low | IVFFLAT index, query optimization, consider disk-ANN if needed |
Alternatives Considered
1. Elasticsearch Full-Text Search
- ❌ Rejected: Keyword-based, misses semantic similarity
- ❌ No understanding of "warehouse" ≈ "storage location"
- ✅ Could be added as hybrid search later
2. OpenAI Embeddings (text-embedding-3-small)
- ❌ Rejected: External API dependency, costs add up
- ❌ 1536 dimensions = larger storage, slower search
- ✅ Better for English, but all-MiniLM sufficient
3. Fine-tuned Model on iDempiere Docs
- ❌ Rejected: High cost, expertise required, maintenance burden
- ❌ Overkill for current scale (50-200 docs)
- ✅ Consider if we reach 10,000+ docs
4. Pinecone / Weaviate (Dedicated Vector DB)
- ❌ Rejected: Additional infrastructure, costs (€70+/month)
- ❌ Operational complexity (another service to monitor)
- ✅ pgvector sufficient for our scale (<100k vectors)
5. Simple Character-Based Chunking
- ❌ Rejected: Breaks semantic units (tables, code blocks)
- ❌ Loses structure = worse context for Claude
- ✅ Our block-based chunking preserves meaning
Success Metrics
Phase 1 (MVP - 3 months):
- [ ] 95% of documentation ingested successfully
- [ ] <2 second query response time (p95)
- [ ] 70% user satisfaction (thumbs up rate)
- [ ] 20% reduction in support tickets related to "how-to" questions
Phase 2 (6 months):
- [ ] 80% user satisfaction
- [ ] 40% reduction in documentation-related support tickets
- [ ] <€100/month Anthropic API costs
- [ ] Support team uses RAG for 50% of responses
Long-term (12 months):
- [ ] Self-service support tier for common questions
- [ ] RAG integrated into iDempiere UI (in-app help)
- [ ] Multi-language support (Slovak, Czech, Hungarian, English)
Monitoring & Observability
Metrics to Track:
- Query latency (p50, p95, p99)
- Retrieval relevance scores
- Claude API errors and retries
- User feedback (thumbs up/down)
- Common query patterns
- Cache hit rate
Logging:
@Logged
public String query(String question) {
log.info("RAG query: {}", question);
long start = System.currentTimeMillis();
// ... query logic
log.info("RAG response time: {}ms, chunks: {}, relevance: {}",
System.currentTimeMillis() - start,
retrievedChunks.size(),
avgRelevanceScore
);
}
References
- LangChain4j Documentation
- pgvector Performance Guide
- Anthropic Claude API
- RAG Best Practices (Anthropic)
Decision
APPROVED - Proceed with implementation
Signature:
Norbert (CEO) - [Date]
Appendix A: Maven Dependencies
<!-- LangChain4j Core -->
<dependency>
<groupId>dev.langchain4j</groupId>
<artifactId>langchain4j</artifactId>
<version>0.35.0</version>
</dependency>
<!-- Quarkus LangChain4j Integration -->
<dependency>
<groupId>io.quarkiverse.langchain4j</groupId>
<artifactId>quarkus-langchain4j-core</artifactId>
<version>0.20.0</version>
</dependency>
<!-- pgvector Store -->
<dependency>
<groupId>io.quarkiverse.langchain4j</groupId>
<artifactId>quarkus-langchain4j-pgvector</artifactId>
<version>0.20.0</version>
</dependency>
<!-- Embeddings (Local) -->
<dependency>
<groupId>dev.langchain4j</groupId>
<artifactId>langchain4j-embeddings-all-minilm-l6-v2</artifactId>
<version>0.35.0</version>
</dependency>
<!-- Anthropic Claude -->
<dependency>
<groupId>dev.langchain4j</groupId>
<artifactId>langchain4j-anthropic</artifactId>
<version>0.35.0</version>
</dependency>
Appendix B: Database Schema
-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;
-- Embeddings table (managed by LangChain4j)
CREATE TABLE IF NOT EXISTS embeddings (
id UUID PRIMARY KEY,
content TEXT NOT NULL,
content_vector VECTOR(384),
metadata JSONB
);
-- Indexes for fast similarity search
CREATE INDEX embeddings_vector_idx
ON embeddings
USING ivfflat (content_vector vector_cosine_ops)
WITH (lists = 100);
-- Index for metadata filtering
CREATE INDEX embeddings_metadata_idx
ON embeddings
USING gin (metadata);
-- Add indexing timestamp to existing documents table
ALTER TABLE documents
ADD COLUMN IF NOT EXISTS last_indexed_at TIMESTAMP;
-- Index for efficient document lookup
CREATE INDEX IF NOT EXISTS documents_last_indexed_idx
ON documents (last_indexed_at);
Appendix C: Service Interfaces
package sk.cloudempiere.rag;
import dev.langchain4j.data.segment.TextSegment;
import java.util.List;
/**
* Main RAG service interface
*/
public interface RagService {
/**
* Query the RAG system with a natural language question
* @param question User's question in Slovak or English
* @return AI-generated answer based on documentation
*/
String query(String question);
/**
* Ingest a single document into the vector store
* @param documentId Database ID of the document
*/
void ingestDocument(Long documentId);
/**
* Batch ingest all documents from the database
*/
void ingestAllDocuments();
/**
* Re-index a document (after updates)
* @param documentId Database ID of the document
*/
void reindexDocument(Long documentId);
}
/**
* Chunking service for Editor.js documents
*/
public interface EditorJsChunker {
/**
* Convert Editor.js JSON into semantic chunks
* @param docId Document database ID
* @param title Document title
* @param editorJson Editor.js JSON string
* @return List of text segments with metadata
*/
List<TextSegment> chunkDocument(Long docId, String title, String editorJson);
}
Appendix D: Example REST API Usage
Ingest Document
curl -X POST http://localhost:8080/api/rag/ingest/123 \
-H "Content-Type: application/json" \
-d '{
"title": "Warehouse Setup Guide",
"editor_json": "{\"blocks\":[...]}"
}'
Response:
{
"status": "ingested",
"document_id": "123",
"chunks_created": 15
}
Query RAG System
curl -X POST http://localhost:8080/api/rag/query \
-H "Content-Type: application/json" \
-d '{
"question": "Ako nastavím nový sklad v iDempiere?"
}'
Response:
{
"question": "Ako nastavím nový sklad v iDempiere?",
"answer": "Pre vytvorenie nového skladu v iDempiere postupujte nasledovne:\n\n1. Otvorte okno Material Management > Warehouse and Locators\n2. Kliknite na tlačidlo New\n3. Vyplňte základné údaje skladu...",
"sources": [
{
"document_id": 123,
"title": "Warehouse Setup Guide",
"relevance_score": 0.89
}
],
"response_time_ms": 1247
}
Appendix E: Phased Rollout Strategy
Week 1-2: Internal Testing
- Deploy to staging environment
- Ingest 10-20 key documents
- Test with support team
- Collect feedback on answer quality
Week 3-4: Beta Testing
- Invite 5 pilot clients
- Monitor usage patterns
- Tune retrieval parameters
- Document common issues
Week 5-6: Production Launch
- Deploy to production
- Ingest full documentation corpus
- Enable for all clients
- Set up monitoring dashboards
Week 7-8: Optimization
- Analyze metrics
- Implement caching for common queries
- Add feedback collection UI
- Create internal knowledge base articles
End of ADR-001
Document Metadata:
- Version: 1.0
- Last Updated: 2024-12-11
- Maintained By: CloudEmpiere Technical Team
- Review Cycle: Quarterly