ADR-078: Mattermost Batch Ingestion for RAG Knowledge Base
Status
Proposed
Date
2025-12-31
Deciders
- Norbert Bede
- Development Team
Context and Problem Statement
ADR-018 proposed live Mattermost API integration for real-time knowledge access. However, real-time API access has limitations:
- Rate limiting affects heavy usage
- Requires persistent API token management
- Network latency on each query
- No preprocessing or quality filtering
A complementary approach is batch ingestion of Mattermost exports, which enables:
- Pre-processing and quality scoring of discussions
- Filtering for high-value, resolved threads
- Integration with existing RAG pipeline (ADR-021)
- Offline operation after initial ingestion
The iDempiere Mattermost community has 10,000+ messages in the ~support channel alone, containing valuable troubleshooting patterns, solutions, and best practices.
Decision Drivers
- Quality over Quantity: Not all chat messages are valuable; need to filter for resolution-bearing threads
- Trusted Authors: Core developers (hengsin, carlosruiz, druiz, hieplq) provide authoritative answers
- Code Examples: Threads with code blocks contain actionable solutions
- RAG Integration: Must integrate with existing PGVector + LangChain4j pipeline (ADR-021)
- Documentation Pipeline: Should also feed the Docusaurus documentation generation
Considered Options
- Batch JSON Ingestion - Process exported Mattermost JSON with quality filtering
- Live API Only (ADR-018) - Real-time API access for each query
- Hybrid - Batch for historical data, live API for recent discussions
Decision Outcome
Chosen option: "Batch JSON Ingestion" with future hybrid capability, because:
- Enables quality filtering and author trust scoring
- Integrates with existing RAG architecture (ADR-021)
- No runtime API dependency
- Supports dual output: RAG embeddings + Docusaurus documentation
Confirmation
The decision is confirmed when:
knowledge ingest --source mattermostsuccessfully processes exported JSON- Thread quality scoring identifies high-value discussions
knowledge searchreturns relevant Mattermost-sourced results- Generated documentation passes Docusaurus build
Architecture
┌─────────────────────────────────────────────────────────────────────────────┐
│ Mattermost Export Pipeline │
└──────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ IngestedDocument JSON (from idempiere-content-ingestor) │
│ ┌────────────────────────────────────────────────────────────────────────┐ │
│ │ { │ │
│ │ "id": "mattermost-abc123", │ │
│ │ "source_type": "mattermost", │ │
│ │ "content": "**hengsin**: The key is to use MStorageOnHand...", │ │
│ │ "metadata": { │ │
│ │ "author_trust": 1.0, │ │
│ │ "topics": ["warehouse", "core"], │ │
│ │ "is_thread_reply": true, │ │
│ │ "root_id": "xyz789" │ │
│ │ }, │ │
│ │ "code_blocks": [{"language": "java", "code": "..."}] │ │
│ │ } │ │
│ └────────────────────────────────────────────────────────────────────────┘ │
└──────────────────────────────────┬──────────────────────────────────────────┘
│
┌──────────────┴──────────────┐
▼ ▼
┌──────────────────────────────────┐ ┌──────────────────────────────────┐
│ RAG Ingestion Pipeline │ │ Documentation Pipeline │
│ ┌────────────────────────────┐ │ │ ┌────────────────────────────┐ │
│ │ MattermostIngestor.java │ │ │ │ mattermost_to_docusaurus.py│ │
│ │ - Thread reconstruction │ │ │ │ - Thread reconstruction │ │
│ │ - Quality scoring │ │ │ │ - Quality scoring │ │
│ │ - Embedding generation │ │ │ │ - MDX generation │ │
│ └─────────────┬──────────────┘ │ │ └─────────────┬──────────────┘ │
│ │ │ │ │ │
│ ▼ │ │ ▼ │
│ ┌────────────────────────────┐ │ │ ┌────────────────────────────┐ │
│ │ PGVector Embedding Store │ │ │ │ Docusaurus MDX Files │ │
│ │ (cli_embeddings table) │ │ │ │ docs/community-discussions│ │
│ └────────────────────────────┘ │ │ └────────────────────────────┘ │
└──────────────────────────────────┘ └──────────────────────────────────┘
Thread Quality Scoring
Threads are scored to prioritize high-value content:
| Factor | Points | Description |
|---|---|---|
| Trusted author (≥0.9) | +3 | Core developers: hengsin, carlosruiz, druiz, hieplq |
| Has code blocks | +2 | Contains actionable code examples |
| Has resolution | +2 | Thread ends with solved/fixed/thanks indicators |
| Multi-participant (≥2) | +1 | Collaborative problem-solving |
| Substantive (≥5 messages) | +1 | Not a quick question/answer |
Quality thresholds:
- High (≥6): Ingest for RAG + generate documentation
- Medium (4-5): Ingest for RAG only
- Low (<4): Skip ingestion
IngestedDocument Schema
The ingestion pipeline expects documents from idempiere-content-ingestor:
public class IngestedDocument {
String id; // "mattermost-{post_id}"
String sourceType; // "mattermost"
String sourceUrl; // Link to original post
String contentType; // "chat_thread" | "chat_message"
String title; // First line or extracted title
String content; // Full message content
String qualified; // "partial" | "full" | "none"
IngestMetadata metadata;
List<CodeBlock> codeBlocks;
}
public class IngestMetadata {
String author; // Mattermost username
double authorTrust; // 0.0-1.0 trust score
Instant timestamp;
List<String> topics; // ["core", "window", "warehouse"]
List<String> classesMentioned;// ["MStorage", "MOrder"]
List<String> jiraTickets; // ["IDEMPIERE-1234"]
String era; // "v10+" | "v9" | "legacy"
String channel; // "support" | "developers"
String rootId; // Thread root post ID
boolean isThreadReply; // true if reply to thread
}
Author Trust Scores
| Trust Level | Score | Authors |
|---|---|---|
| Core Developer | 1.0 | hengsin |
| Core Developer | 0.95 | carlosruiz |
| Core Developer | 0.9 | druiz, hieplq |
| Active Contributor | 0.8 | norbertbede, nmicoud |
| Regular Contributor | 0.7 | chuck, muriloht |
| Community Member | 0.5 | (default) |
CLI Commands
# Ingest Mattermost export
idempiere-cli knowledge ingest --source mattermost \
--file mattermost_support.json
# Ingest with quality threshold
idempiere-cli knowledge ingest --source mattermost \
--file mattermost_support.json \
--min-score 6
# Ingest specific channel
idempiere-cli knowledge ingest --source mattermost \
--file mattermost_developers.json \
--channel developers
# Search Mattermost knowledge
idempiere-cli knowledge search "MStorage lazy loading" --source mattermost
# Show Mattermost statistics
idempiere-cli knowledge status --source mattermost
Configuration
# application.properties
# Mattermost Ingestion
idempiere.cli.rag.mattermost.enabled=true
idempiere.cli.rag.mattermost.min-quality-score=4
idempiere.cli.rag.mattermost.min-thread-length=3
# Author trust configuration
idempiere.cli.rag.mattermost.trusted-authors=hengsin,carlosruiz,druiz,hieplq
# Topic extraction
idempiere.cli.rag.mattermost.topic-keywords=core,window,report,warehouse,integration,plugins
Implementation Plan
Phase 1: Ingestor Implementation (Days 1-3)
- [ ] Create
MattermostIngestor.javaimplementingKnowledgeIngestor - [ ] Implement thread reconstruction from individual messages
- [ ] Implement quality scoring algorithm
- [ ] Add metadata extraction (topics, JIRA tickets, classes)
Phase 2: Embedding Generation (Days 4-5)
- [ ] Generate document format:
{title}\n{problem}\n{solution}\n{code} - [ ] Add source metadata for filtering
- [ ] Test search relevance with sample queries
Phase 3: CLI Integration (Day 6)
- [ ] Add
--source mattermostto knowledge ingest command - [ ] Add file input option for JSON export
- [ ] Add quality threshold option
Phase 4: Documentation (Day 7)
- [ ] Update USER_GUIDE.md with Mattermost ingestion
- [ ] Document export process from Mattermost
- [ ] Add example queries
Export Process
Mattermost data is exported using idempiere-content-ingestor:
# Export support channel
python ingest_mattermost.py \
--server https://mattermost.idempiere.org \
--token $MATTERMOST_TOKEN \
--channel support \
--output mattermost_support.json
# Export with date range
python ingest_mattermost.py \
--channel support \
--since 2024-01-01 \
--output mattermost_support_2024.json
Dual Output Strategy
The same ingested data serves two purposes:
- RAG Knowledge Base: Embeddings in PGVector for semantic search
- Docusaurus Documentation: Generated MDX files for browsable docs
This ensures knowledge is accessible both:
- Programmatically (AI-assisted queries)
- Human-readable (documentation site)
Consequences
Positive
- High-quality knowledge extraction (not all chat noise)
- Integrates with existing RAG pipeline
- Offline operation after ingestion
- Dual output: RAG + Documentation
- Author trust scoring improves answer quality
Negative
- Requires periodic re-export for fresh data
- Export process needs Mattermost API token
- Storage overhead for embeddings
Neutral
- Complements live API integration (ADR-018)
- Can be combined with real-time updates in future
Related ADRs
- ADR-018: Mattermost Knowledge Integration - Live API approach
- ADR-021: RAG Architecture - Base RAG infrastructure
- ADR-039: Code to Knowledge Extraction - Similar ingestion pattern
References
- iDempiere Mattermost
- idempiere-content-ingestor - Export scripts
- LangChain4j PGVector
ADR-078 | Version 1.0 | 2025-12-31 Status: Proposed Decision: Batch Mattermost ingestion with quality scoring for RAG knowledge base