ADR-040: K_Entry Content Format Normalization for RAG Ingestion
<!-- MADR 3.0 Template - Markdown Any Decision Records --> <!-- Reference: https://adr.github.io/madr/ -->
Status
Proposed
Date
2025-12-11
Deciders
- Norbert Bede
- Development Team
Context and Problem Statement
The K_Entry table stores knowledge base articles in three different formats (K_EntryEditMode):
- BLK - Editor.js Block (JSON format)
- GFM - GitHub Flavored Markdown
- HTM - HTML
Currently, the KEntryIngestor extracts the TextMsg field as-is without format conversion. This creates problems for RAG (Retrieval-Augmented Generation) because:
- Editor.js JSON is not human-readable - AI embeddings work best with natural language text
- HTML contains markup noise - Tags and attributes reduce embedding quality
- Inconsistent format - Same semantic content has different representations
- Poor search results - JSON structure words (e.g., "blocks", "data", "type") pollute the search space
We need to normalize all content formats to a single, AI-friendly representation before embedding.
Decision Drivers
- Embedding Quality - Embeddings should represent semantic meaning, not format syntax
- Search Accuracy - Queries should match content, not JSON/HTML structure
- Maintainability - Single conversion pipeline easier than format-specific handling
- Lossless Conversion - Preserve all semantic information (headings, lists, emphasis)
- Performance - Conversion should not significantly slow down ingestion
- Existing Tools - Prefer using LangChain4j or standard libraries over custom parsers
Considered Options
- Convert all formats to Markdown using LangChain4j DocumentParser + custom Editor.js converter
- Convert all formats to plain text (strip all formatting)
- Keep formats as-is, add format-specific metadata for filtering
- Use AI to convert formats on-the-fly during ingestion
Decision Outcome
Chosen option: "Convert all formats to Markdown using LangChain4j DocumentParser + custom Editor.js converter", because it:
- Preserves semantic structure (headings, lists, emphasis) for better embeddings
- Uses standard Markdown as the universal format
- Leverages LangChain4j's Apache Tika for HTML parsing
- Only requires custom converter for Editor.js (well-defined JSON schema)
- Markdown is readable by both humans and AI
Confirmation
The decision is confirmed when:
- [ ] All three formats (BLK, GFM, HTM) are converted to Markdown before embedding
- [ ] Search for "order process" returns relevant K_Entry regardless of original format
- [ ] No JSON structure words ("blocks", "data") appear in search results for Editor.js entries
- [ ] Markdown output preserves headings, lists, bold/italic formatting
- [ ] Ingestion performance remains acceptable (< 10% slowdown)
Pros and Cons of the Options
Option 1: Convert to Markdown (chosen)
Use LangChain4j DocumentParser for HTML, custom converter for Editor.js, pass-through for GFM.
- Good, because Markdown preserves semantic structure (headings, lists, emphasis)
- Good, because LangChain4j provides Apache Tika integration for HTML→Markdown
- Good, because Editor.js has well-defined block types
- Good, because Markdown is human-readable and AI-friendly
- Good, because single output format simplifies downstream processing
- Neutral, because requires custom Editor.js→Markdown converter
- Bad, because adds conversion overhead to ingestion
Option 2: Convert to Plain Text
Strip all formatting and extract only text content.
- Good, because simplest implementation
- Good, because fastest conversion
- Good, because no format-specific handling
- Bad, because loses semantic structure (headings become indistinguishable from body text)
- Bad, because loses emphasis (bold/italic)
- Bad, because lists lose structure
- Bad, because lower quality embeddings (structure helps meaning)
Option 3: Keep Formats As-Is
Store content in original format, add metadata for filtering.
- Good, because no conversion needed
- Good, because preserves original format exactly
- Bad, because AI must handle three different formats
- Bad, because JSON structure pollutes embeddings
- Bad, because HTML markup pollutes embeddings
- Bad, because search results inconsistent across formats
- Bad, because complex query logic (format-aware filtering)
Option 4: AI-Powered Conversion
Use LLM to convert formats during ingestion.
- Good, because handles edge cases intelligently
- Good, because can improve formatting
- Bad, because slow (LLM call per entry)
- Bad, because expensive (API costs or local model overhead)
- Bad, because non-deterministic (same input may produce different output)
- Bad, because requires LLM availability during ingestion
Architecture
┌─────────────────────────────────────────────────────────────────────────────┐
│ K_ENTRY CONTENT NORMALIZATION PIPELINE │
├─────────────────────────────────────────────────────────────────────────────┤
│ │
│ ┌─────────────────────────────────────────────────────────────────────────┐ │
│ │ 1. QUERY K_Entry with K_EntryEditMode │ │
│ │ │ │
│ │ SELECT k_entry_id, name, keywords, textmsg, │ │
│ │ k_entryeditmode, -- BLK | GFM | HTM │ │
│ │ descriptionurl │ │
│ │ FROM k_entry WHERE ... │ │
│ └─────────────────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────────────────────────────────┐ │
│ │ 2. CONTENT FORMAT ROUTER │ │
│ │ │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │
│ │ │ BLK │ │ GFM │ │ HTM │ │ │
│ │ │Editor.js │ │Markdown │ │ HTML │ │ │
│ │ └────┬─────┘ └────┬─────┘ └────┬─────┘ │ │
│ │ │ │ │ │ │
│ │ ▼ ▼ ▼ │ │
│ └─────────────────────────────────────────────────────────────────────────┘ │
│ │ │ │ │
│ │ │ │ │
│ ▼ │ ▼ │
│ ┌──────────────────────┐ │ ┌──────────────────────┐ │
│ │ EditorJsConverter │ │ │ HtmlToMarkdown │ │
│ │ │ │ │ (Apache Tika) │ │
│ │ Parse JSON: │ │ │ │ │
│ │ { │ │ │ <h2>Title</h2> │ │
│ │ "blocks": [...] │ │ │ <p>Text</p> │ │
│ │ } │ │ │ <ul><li>Item</li> │ │
│ │ │ │ │ │ │
│ │ Output Markdown: │ │ │ Output Markdown: │ │
│ │ ## Header │ │ │ ## Title │ │
│ │ Paragraph text │ │ │ Text │ │
│ │ - List item │ │ │ - Item │ │
│ └────────┬─────────────┘ │ └────────┬─────────────┘ │
│ │ │ │ │
│ │ │ (pass-through) │
│ │ │ │ │
│ └────────────────┼─────────────┘ │
│ ▼ │
│ ┌─────────────────────────────────────────────────────────────────────────┐ │
│ │ 3. UNIFIED MARKDOWN CONTENT │ │
│ │ │ │
│ │ Path: Knowledge Base > Topic > Entry │ │
│ │ │ │
│ │ # Entry Title │ │
│ │ │ │
│ │ Keywords: order, process, sales │ │
│ │ │ │
│ │ ## Section Heading │ │
│ │ │ │
│ │ Paragraph text with **bold** and *italic* formatting. │ │
│ │ │ │
│ │ - List item 1 │ │
│ │ - List item 2 │ │
│ │ │ │
│ │ Reference: https://example.com/doc │ │
│ └─────────────────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────────────────────────────────┐ │
│ │ 4. DOCUMENT SPLITTING │ │
│ │ (LangChain4j DocumentSplitters.recursive) │ │
│ └─────────────────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────────────────────────────────┐ │
│ │ 5. EMBEDDING + STORAGE │ │
│ │ (BGE-small-en-q + PGVector) │ │
│ │ │ │
│ │ Metadata: │ │
│ │ - source_type: "k_entry" │ │
│ │ - source_id: "12345_en_US" │ │
│ │ - original_format: "BLK" | "GFM" | "HTM" <-- New! │ │
│ │ - name: "Entry Title" │ │
│ │ - knowledge_base: "CloudEmpiere ERP" │ │
│ │ - breadcrumb: "KB > Topic > Entry" │ │
│ │ - ad_language: "en_US" │ │
│ └─────────────────────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────────────┘
Implementation Plan
Phase 1: Editor.js Converter (Core)
File: src/main/java/org/idempiere/cli/rag/converter/EditorJsConverter.java
@ApplicationScoped
public class EditorJsConverter {
/**
* Convert Editor.js JSON to Markdown.
*
* @param editorJsJson JSON string in Editor.js format
* @return Markdown text
*/
public String toMarkdown(String editorJsJson) {
// Parse JSON
JsonObject root = Json.createReader(new StringReader(editorJsJson)).readObject();
JsonArray blocks = root.getJsonArray("blocks");
StringBuilder markdown = new StringBuilder();
for (JsonObject block : blocks.getValuesAs(JsonObject.class)) {
String type = block.getString("type");
JsonObject data = block.getJsonObject("data");
switch (type) {
case "header" -> {
int level = data.getInt("level", 1);
String text = data.getString("text");
markdown.append("#".repeat(level)).append(" ")
.append(stripHtml(text)).append("\n\n");
}
case "paragraph" -> {
String text = data.getString("text");
markdown.append(convertHtmlToMarkdown(text)).append("\n\n");
}
case "list" -> {
JsonArray items = data.getJsonArray("items");
String style = data.getString("style", "unordered");
for (int i = 0; i < items.size(); i++) {
if ("ordered".equals(style)) {
markdown.append(i + 1).append(". ");
} else {
markdown.append("- ");
}
markdown.append(stripHtml(items.getString(i))).append("\n");
}
markdown.append("\n");
}
case "quote" -> {
String text = data.getString("text");
markdown.append("> ").append(stripHtml(text)).append("\n\n");
}
case "code" -> {
String code = data.getString("code");
markdown.append("```\n").append(code).append("\n```\n\n");
}
case "delimiter" -> {
markdown.append("---\n\n");
}
case "warning" -> {
String title = data.getString("title", "");
String message = data.getString("message", "");
markdown.append("> **⚠️ ").append(title).append("**\n> ")
.append(message).append("\n\n");
}
case "table" -> {
// Simple table conversion (Editor.js tables are complex)
markdown.append(convertTableToMarkdown(data)).append("\n\n");
}
default -> {
// Unknown block type - extract text if available
if (data.containsKey("text")) {
markdown.append(stripHtml(data.getString("text"))).append("\n\n");
}
}
}
}
return markdown.toString().trim();
}
/**
* Convert inline HTML to Markdown (bold, italic, links).
*/
private String convertHtmlToMarkdown(String html) {
return html
.replaceAll("<b>(.*?)</b>", "**$1**")
.replaceAll("<strong>(.*?)</strong>", "**$1**")
.replaceAll("<i>(.*?)</i>", "*$1*")
.replaceAll("<em>(.*?)</em>", "*$1*")
.replaceAll("<a href=\"(.*?)\">(.*?)</a>", "[$2]($1)")
.replaceAll("<code>(.*?)</code>", "`$1`")
.replaceAll("<[^>]+>", ""); // Strip remaining tags
}
private String stripHtml(String html) {
return convertHtmlToMarkdown(html);
}
private String convertTableToMarkdown(JsonObject data) {
// Simplified table conversion
// Editor.js table format: { "content": [[cell, cell], [cell, cell]] }
JsonArray content = data.getJsonArray("content");
StringBuilder table = new StringBuilder();
// Header row
JsonArray headerRow = content.getJsonArray(0);
table.append("| ");
for (int i = 0; i < headerRow.size(); i++) {
table.append(stripHtml(headerRow.getString(i))).append(" | ");
}
table.append("\n");
// Separator
table.append("| ");
for (int i = 0; i < headerRow.size(); i++) {
table.append("--- | ");
}
table.append("\n");
// Data rows
for (int r = 1; r < content.size(); r++) {
JsonArray row = content.getJsonArray(r);
table.append("| ");
for (int c = 0; c < row.size(); c++) {
table.append(stripHtml(row.getString(c))).append(" | ");
}
table.append("\n");
}
return table.toString();
}
}
Phase 2: HTML Converter (LangChain4j)
File: src/main/java/org/idempiere/cli/rag/converter/HtmlToMarkdownConverter.java
@ApplicationScoped
public class HtmlToMarkdownConverter {
/**
* Convert HTML to Markdown using Apache Tika.
* LangChain4j provides Tika integration via DocumentParser.
*/
public String toMarkdown(String html) {
// Option 1: Use Tika directly
try {
InputStream stream = new ByteArrayInputStream(html.getBytes(StandardCharsets.UTF_8));
BodyContentHandler handler = new BodyContentHandler();
Metadata metadata = new Metadata();
metadata.set(Metadata.CONTENT_TYPE, "text/html");
HtmlParser parser = new HtmlParser();
parser.parse(stream, handler, metadata, new ParseContext());
// Tika outputs plain text, need Markdown conversion
// Use external library like flexmark or custom converter
return convertPlainTextToMarkdown(handler.toString(), html);
} catch (Exception e) {
LOG.warn("Failed to parse HTML, falling back to plain text", e);
return html.replaceAll("<[^>]+>", ""); // Strip tags
}
}
/**
* Convert Tika plain text output to Markdown by analyzing original HTML.
*/
private String convertPlainTextToMarkdown(String plainText, String originalHtml) {
// Use jsoup for HTML parsing
org.jsoup.nodes.Document doc = Jsoup.parse(originalHtml);
StringBuilder markdown = new StringBuilder();
// Convert heading tags
for (org.jsoup.nodes.Element h : doc.select("h1, h2, h3, h4, h5, h6")) {
int level = Integer.parseInt(h.tagName().substring(1));
markdown.append("#".repeat(level)).append(" ").append(h.text()).append("\n\n");
}
// Convert paragraphs
for (org.jsoup.nodes.Element p : doc.select("p")) {
markdown.append(convertInlineElements(p)).append("\n\n");
}
// Convert lists
for (org.jsoup.nodes.Element ul : doc.select("ul")) {
for (org.jsoup.nodes.Element li : ul.select("li")) {
markdown.append("- ").append(li.text()).append("\n");
}
markdown.append("\n");
}
for (org.jsoup.nodes.Element ol : doc.select("ol")) {
int index = 1;
for (org.jsoup.nodes.Element li : ol.select("li")) {
markdown.append(index++).append(". ").append(li.text()).append("\n");
}
markdown.append("\n");
}
return markdown.toString();
}
private String convertInlineElements(org.jsoup.nodes.Element element) {
String html = element.html();
return html
.replaceAll("<strong>(.*?)</strong>", "**$1**")
.replaceAll("<b>(.*?)</b>", "**$1**")
.replaceAll("<em>(.*?)</em>", "*$1*")
.replaceAll("<i>(.*?)</i>", "*$1*")
.replaceAll("<a href=\"(.*?)\">(.*?)</a>", "[$2]($1)")
.replaceAll("<code>(.*?)</code>", "`$1`")
.replaceAll("<[^>]+>", "");
}
}
Phase 3: Update KEntryIngestor
File: src/main/java/org/idempiere/cli/rag/ingest/KEntryIngestor.java
Update SQL query to include K_EntryEditMode:
private static final String K_ENTRY_SQL = """
WITH RECURSIVE entry_tree AS (
-- ... existing CTE ...
)
SELECT
et.k_entry_id,
COALESCE(NULLIF(trl.name, ''), et.base_name) as name,
et.issummary,
COALESCE(NULLIF(trl.keywords, ''), et.base_keywords) as keywords,
COALESCE(NULLIF(trl.textmsg, ''), et.base_textmsg) as textmsg,
et.k_entryeditmode, -- NEW: Get format type
et.descriptionurl,
et.created,
et.updated,
et.k_type_id,
et.type_name,
et.tree_level,
et.breadcrumb
FROM entry_tree et
LEFT JOIN k_entry_trl trl ON et.k_entry_id = trl.k_entry_id
AND trl.ad_language = ?
AND trl.isactive = 'Y'
ORDER BY et.breadcrumb
""";
Update ingestion logic:
@Inject
EditorJsConverter editorJsConverter;
@Inject
HtmlToMarkdownConverter htmlConverter;
// Inside ResultSet loop:
while (rs.next()) {
// ... existing field extraction ...
String textMsg = rs.getString("textmsg");
String editMode = rs.getString("k_entryeditmode"); // BLK | GFM | HTM
// Convert to Markdown based on format
String markdownContent = switch (editMode) {
case "BLK" -> {
// Editor.js JSON to Markdown
try {
yield editorJsConverter.toMarkdown(textMsg);
} catch (Exception e) {
LOG.warnf("Failed to convert Editor.js for K_Entry %d: %s",
kEntryId, e.getMessage());
yield textMsg; // Fallback to raw content
}
}
case "HTM" -> {
// HTML to Markdown
try {
yield htmlConverter.toMarkdown(textMsg);
} catch (Exception e) {
LOG.warnf("Failed to convert HTML for K_Entry %d: %s",
kEntryId, e.getMessage());
yield textMsg; // Fallback to raw content
}
}
case "GFM" -> textMsg; // Already Markdown, pass-through
default -> {
LOG.warnf("Unknown K_EntryEditMode '%s' for K_Entry %d, treating as plain text",
editMode, kEntryId);
yield textMsg;
}
};
// Build document content with hierarchical context
StringBuilder content = new StringBuilder();
// Add breadcrumb as context header
if (breadcrumb != null && !breadcrumb.isBlank()) {
content.append("Path: ").append(breadcrumb).append("\n\n");
}
if (name != null && !name.isBlank()) {
content.append("# ").append(name).append("\n\n");
}
if (keywords != null && !keywords.isBlank()) {
content.append("Keywords: ").append(keywords).append("\n\n");
}
if (markdownContent != null && !markdownContent.isBlank()) {
content.append(markdownContent); // Use converted Markdown
}
if (descriptionUrl != null && !descriptionUrl.isBlank()) {
content.append("\n\nReference: ").append(descriptionUrl);
}
// ... existing document creation and embedding ...
// Add original_format to metadata
doc.metadata().put("original_format", editMode);
}
Phase 4: Dependencies
Add to pom.xml:
<!-- JSON Processing (Jakarta EE) -->
<dependency>
<groupId>jakarta.json</groupId>
<artifactId>jakarta.json-api</artifactId>
<version>2.1.3</version>
</dependency>
<dependency>
<groupId>org.glassfish</groupId>
<artifactId>jakarta.json</artifactId>
<version>2.0.1</version>
</dependency>
<!-- HTML Parsing -->
<dependency>
<groupId>org.jsoup</groupId>
<artifactId>jsoup</artifactId>
<version>1.17.2</version>
</dependency>
<!-- Apache Tika (already included via langchain4j-document-parser-apache-tika) -->
Phase 5: Testing
File: src/test/java/org/idempiere/cli/rag/converter/EditorJsConverterTest.java
@QuarkusTest
class EditorJsConverterTest {
@Inject
EditorJsConverter converter;
@Test
void testHeaderConversion() {
String json = """
{
"time": 1550476186479,
"blocks": [
{
"type": "header",
"data": {
"text": "Sample Heading",
"level": 2
}
}
],
"version": "2.8.1"
}
""";
String result = converter.toMarkdown(json);
assertTrue(result.contains("## Sample Heading"));
}
@Test
void testParagraphWithHtml() {
String json = """
{
"blocks": [
{
"type": "paragraph",
"data": {
"text": "Text with <b>bold</b> and <i>italic</i>"
}
}
]
}
""";
String result = converter.toMarkdown(json);
assertTrue(result.contains("**bold**"));
assertTrue(result.contains("*italic*"));
}
@Test
void testListConversion() {
String json = """
{
"blocks": [
{
"type": "list",
"data": {
"style": "unordered",
"items": ["Item 1", "Item 2", "Item 3"]
}
}
]
}
""";
String result = converter.toMarkdown(json);
assertTrue(result.contains("- Item 1"));
assertTrue(result.contains("- Item 2"));
assertTrue(result.contains("- Item 3"));
}
}
Configuration
Add to application.properties:
# K_Entry Content Format Conversion
idempiere.cli.rag.k-entry.convert-to-markdown=true
idempiere.cli.rag.k-entry.fallback-on-error=true # Use raw content if conversion fails
Migration Strategy
Backward Compatibility
- No breaking changes - existing ingested data remains valid
- New metadata field
original_formatadded (non-breaking) - Re-ingestion required to get converted content
Migration Steps
- Deploy new version with converters
- Run
knowledge clear --source k_entry -y - Run
knowledge ingest --source k_entry - Verify search quality improved
Rollback Plan
If conversion causes issues:
- Set
idempiere.cli.rag.k-entry.convert-to-markdown=false - Re-ingest using raw content
Performance Considerations
Expected Impact
| Format | Conversion Time | Impact |
|---|---|---|
| GFM (Markdown) | 0ms (pass-through) | None |
| HTM (HTML) | ~5ms per entry | Low |
| BLK (Editor.js) | ~2ms per entry | Low |
Total ingestion slowdown: ~5-10% (acceptable for batch operation)
Optimization Opportunities
- Parallel conversion - Convert multiple entries concurrently
- Caching - Cache converted content (keyed by source_id + hash)
- Lazy conversion - Convert on-demand instead of during ingestion
More Information
Editor.js Block Types
Common block types in CloudEmpiere:
| Type | Markdown Equivalent | Notes |
|---|---|---|
header |
## Heading |
Level 1-6 |
paragraph |
Plain text | May contain HTML |
list |
- Item or 1. Item |
Ordered/unordered |
quote |
> Quote |
Blockquote |
code |
```code``` |
Code block |
delimiter |
--- |
Horizontal rule |
warning |
> ⚠️ Warning |
Custom block |
table |
Markdown table | Complex format |
Related ADRs
- ADR-021 - RAG Architecture
- ADR-039 - Code-to-Knowledge Extraction
- ADR-022 - Shared Embedding Infrastructure