Skip to main content

Overview

Port: 4005 The Knowledge Service handles the RAG pipeline including document processing, content connectors, multimodal vector embeddings, knowledge graph extraction, and unified search across multiple knowledge bases.

Endpoints

Knowledge Base CRUD

A request with no x-org-role header is decided as a viewer, never a member (C-191).

Documents

Knowledge Graph

Graph Optimization

These endpoints are unchanged, and still back the Issues queue’s Generate / AI review / Approve / Reject / Apply actions. For reading what the graph flagged, prefer GET /knowledge/kb/:id/graph/issues above — the suggestions list knows nothing about orphaned or low-confidence entities. Link and unlink gate edit on the knowledge base (knowledge_base) and edit on the agent (bot), branching on each verdict. An agent or KB in another organization answers 404, and a cross-organization link is refused with 409 even for a superadmin; a superadmin may unlink one. Both write an audit entry under the KB’s organization. See Link Agent to KB.

Utilities

RAG Pipeline

Indexing Flow

Pipeline steps:
  1. Upload — Document, URL, or social media source
  2. Load — LlamaParse (PDF/docx) or LangChain fallback loaders
  3. Chunk — 1000 characters, 200 overlap
  4. Extract Media — Vision text extraction for images/PDFs, transcription for audio/video
  5. Embed — Gemini Embedding 2 (3072 dimensions, multimodal)
  6. Store — Pinecone (production) or ChromaDB (local)
  7. Graph Extract — Entity seeding from existing KB entities + LLM extraction
  8. Summarize — Per-document summaries for KB map

Graph Enhancement

After document-level extraction, an optional enhance-graph pass runs:
1

Entity Resolution

LLM-confirmed duplicate detection via name similarity (Levenshtein, substring, abbreviation matching). Merged entities consolidate relationships, aliases, and mention counts.
2

Cross-Document Inference

Discovers relationships between entities appearing in different documents but never explicitly connected in any single document.

Unified Search Flow

Content Connectors

All KB source ingestion flows through a unified pipeline built on the ContentConnector pattern.

NormalizedContent

All connectors produce NormalizedContent[] with:
  • externalId (dedup key), contentHash (change detection via SHA-256)
  • title, textContent, media[] (image/video/audio attachments)
  • source metadata (connector name, URL, platform, author, publishedAt)
  • Optional: thumbnailUrl, engagementMetrics, mediaType

Core Services

Database Tables

External Integrations

Critical Patterns

All external calls MUST log costs via record_external_cost_event().
Use vectorStoreService abstraction — never call Pinecone or Chroma directly.
BullMQ for async processing: max 5 concurrent jobs, 3 retries.
Chunks are tagged with embeddingType: "text" | "image" | "video" | "audio" | "transcript".
request.user.organizationId from JWT middleware, never raw headers.
checkResourceAccess() before all KB operations.