Financial Services & FintechPrivate RAG Architecture: Deploying Retrieval Systems Inside Your VPC Without Data Leakage
Strategic White PaperIndustry: Financial Services & FintechPractice: Artificial Intelligence & Data

Private RAG Architecture: Deploying Retrieval Systems Inside Your VPC Without Data Leakage

An end-to-end engineering blueprint for building air-gapped Retrieval-Augmented Generation inside private VPC subnets with zero external model egress: deterministic PDF layout parsing, pgvector hybrid search (BM25 + HNSW), and sovereign quantized LLM inference.

D

Danisur Rahman

Verified Practice Lead
Lead Systems Architect•Sep 23, 2026•14 min read
Private RAG Architecture: Deploying Retrieval Systems Inside Your VPC Without Data Leakage

For regulated enterprises—spanning financial services, healthcare systems, defense, and multinational commerce—the commercial appeal of Retrieval-Augmented Generation (RAG) is frequently halted by the reality of corporate data governance. While product teams rush to query internal documents using frontier cloud models, Chief Information Security Officers (CISOs) and enterprise compliance boards are confronted with an unacceptable trade-off: transmitting confidential customer ledgers, patient records, unreleased financial statements, and core IP to third-party public API endpoints.

Public API RAG introduces catastrophic compliance vulnerabilities. Under GDPR Article 44, HIPAA Security Rule §164.312, and SEC/FINRA Rule 4511, routing proprietary data through multi-tenant external endpoints creates uncontrolled data residency risk, shadow model training exposure, and unverified data retention vectors.

The enterprise answer is neither banning generative AI nor accepting data leakage. It is the deployment of a fully Sovereign, Air-Gapped Private RAG Architecture inside your isolated Virtual Private Cloud (VPC).

In this comprehensive reference architecture, we examine the production blueprint for an enterprise-grade, air-gapped retrieval and synthesis pipeline operating entirely within private subnets: from deterministic document parsing and layout-aware chunking to PostgreSQL 16 pgvector hybrid search (BM25 + HNSW) and high-throughput local quantized LLM inference via vLLM.

[Visual Asset: Air-Gapped VPC Security Perimeter - Zero-Egress Ingress & Compute Isolation]

Exact Visual Specification: A comprehensive cloud infrastructure diagram illustrating an enterprise Virtual Private Cloud (VPC) segmented into three isolated network tiers: Public DMZ (TLS termination, WAF, Internal API Gateway), Isolated Compute Subnet (Layout Parser Pods, TEI Local Embedding Engine, vLLM Local Inference Cluster on NVIDIA GPUs), and Isolated Data Subnet (PostgreSQL 16 + pgvector Primary/Replica cluster, encrypted S3 buckets via Gateway Endpoints). Highlights the air-gapped security perimeter with all public internet egress (0.0.0.0/0) strictly severed via NACLs and AWS PrivateLink.

mermaid
flowchart TD
    subgraph PublicInternet [400 font-semibold">class="text-emerald-300">"External Network / Corporate VPN"]
        ClientApp[400 font-semibold">class="text-emerald-300">"Authenticated Enterprise Client<br/>(Internal Web / ERP / CRM)"]
    end

    subgraph AWS_VPC [400 font-semibold">class="text-emerald-300">"Air-Gapped Cloud VPC (10.100.0.0/16)"]
        subgraph DMZ_Subnet [400 font-semibold">class="text-emerald-300">"DMZ Subnet (10.100.1.0/24)"]
            WAF[400 font-semibold">class="text-emerald-300">"Cloud WAF & DDoS Shield"]
            ALB[400 font-semibold">class="text-emerald-300">"Internal ALB / Reverse Proxy<br/>(mTLS + JWT Validation)"]
        end

        subgraph Compute_Subnet [400 font-semibold">class="text-emerald-300">"Private Isolated Compute Subnet (10.100.10.0/24) - ZERO EGRESS"]
            IngestService[400 font-semibold">class="text-emerald-300">"Document Ingestion Engine<br/>(Docling / Layout Parser Pods)"]
            TEI[400 font-semibold">class="text-emerald-300">"Local Embedding Service<br/>(BAAI/bge-m3 on NVIDIA L4)"]
            Reranker[400 font-semibold">class="text-emerald-300">"Local Cross-Encoder<br/>(bge-reranker-large)"]
            vLLM[400 font-semibold">class="text-emerald-300">"Air-Gapped Local LLM Engine<br/>(Llama-3.1-70B-AWQ on 2x A100 80GB)"]
            RAGOrchestrator[400 font-semibold">class="text-emerald-300">"RAG Query & Synthesis Gateway<br/>(Context Assembly & Guardrails)"]
        end

        subgraph Data_Subnet [400 font-semibold">class="text-emerald-300">"Private Isolated Data Subnet (10.100.20.0/24)"]
            PGVector[400 font-semibold">class="text-emerald-300">"PostgreSQL 16 Cluster + pgvector<br/>(HNSW Vector Index + BM25 tsvector)"]
            S3_VPCE[400 font-semibold">class="text-emerald-300">"S3 VPC Gateway Endpoint<br/>(Raw Encrypted PDFs & Document Store)"]
            KMS_VPCE[400 font-semibold">class="text-emerald-300">"AWS KMS Interface Endpoint<br/>(Customer-Managed Keys - CMK)"]
        end
    end

    subgraph Blocked_Perimeter [400 font-semibold">class="text-emerald-300">"Perimeter Enforcement"]
        IGW[400 font-semibold">class="text-emerald-300">"Internet Gateway / NAT Egress<br/>[ROUTE 0.0.0.0/0: 400 font-semibold">DROP / DENY]"]
    end

    ClientApp -->|Mutual TLS 1.3 + OIDC| WAF
    WAF --> ALB
    ALB --> RAGOrchestrator
    RAGOrchestrator -->|Internal gRPC| TEI
    RAGOrchestrator -->|SQL over TLS + RLS| PGVector
    RAGOrchestrator -->|Rank Pool| Reranker
    RAGOrchestrator -->|Local HTTP Streaming| vLLM
    IngestService -->|Read Raw Docs| S3_VPCE
    IngestService -->|Batch Embeddings| TEI
    IngestService -->|ACID Multi-Row Upsert| PGVector
    PGVector -.->|EBS Encryption| KMS_VPCE
    Compute_Subnet -.->|BLOCKED| IGW

sh
+---------------------------------------------------------------------------------------------------+
|                            ENTERPRISE AIR-GAPPED VPC PERIMETER (10.100.0.0/16)                    |
+---------------------------------------------------------------------------------------------------+
|                                                                                                   |
|  [Corporate Network / VPN]                                                                        |
|            │ (Mutual TLS 1.3 + RBAC Token)                                                        |
|            ▼                                                                                      |
|  +───────────────────────────────────────────+                                                    |
|  | DMZ Subnet: Internal WAF + Envoy Gateway  |                                                    |
|  +───────────────────────────────────────────+                                                    |
|            │                                                                                      |
|            ▼ (Private VPC Network Interface)                                                      |
|  +─────────────────────────────────────────────────────────────────────────────────────────────+  |
|  | PRIVATE COMPUTE SUBNET (Zero Public Egress Route Table: 0.0.0.0/0 -> DENY)                  |  |
|  |                                                                                             |  |
|  |   ┌───────────────────────────┐      ┌───────────────────────────┐                          |  |
|  |   │ Document Layout Parser    │ ───► │ Local Embedding Engine    │                          |  |
|  |   │ (Docling Table Extraction)│      │ (BGE-M3 / TEI on GPU)     │                          |  |
|  |   └─────────────┬─────────────┘      └─────────────┬─────────────┘                          |  |
|  |                 │                                  │ (1024-dim vectors)                     |  |
|  |                 ▼                                  ▼                                        |  |
|  |   ┌──────────────────────────────────────────────────────────────┐                          |  |
|  |   │ RAG Orchestrator (Guardrails, RRF Fusion, Citation Tracking) │                          |  |
|  |   └─────────────────────────────┬────────────────────────────────┘                          |  |
|  |                                 │                                                           |  |
|  |                                 ▼                                                           |  |
|  |                  ┌──────────────────────────────┐                                           |  |
|  |                  │ Air-Gapped vLLM Engine       │                                           |  |
|  |                  │ (Llama-3.1-70B-AWQ / 4-bit)  │                                           |  |
|  |                  └──────────────────────────────┘                                           |  |
|  +─────────────────────────────────────────────────────────────────────────────────────────────+  |
|            │                                                                                      |
|            ▼ (PrivateLink / VPC Peering)                                                          |
|  +─────────────────────────────────────────────────────────────────────────────────────────────+  |
|  | PRIVATE DATA SUBNET                                                                         |  |
|  |   ┌─────────────────────────────────────────┐   ┌────────────────────────────────────────┐  |  |
|  |   │ PostgreSQL 16 + pgvector                │   │ Encrypted S3 Document Vault            │  |  |
|  |   │ (HNSW Cosine Index + BM25 Full-Text)    │   │ (Via S3 VPC Gateway Endpoint)          │  |  |
|  |   └─────────────────────────────────────────┘   └────────────────────────────────────────┘  |  |
|  +─────────────────────────────────────────────────────────────────────────────────────────────+  |
|                                                                                                   |
|  [FIREWALL EGRESS AUDIT]: Outbound Internet Gateway (IGW) = NON-EXISTENT                          |
|                           Public NAT Gateway = DISABLED                                           |
|                           DNS Resolver = Private Route 53 (External Domains Blackholed)           |
+---------------------------------------------------------------------------------------------------+

Figure 1: Architectural topology of an air-gapped enterprise Private RAG environment. The compute and data tiers run entirely on isolated subnets with zero route to the public internet.

1. The Zero-Egress Network Architecture#

Enterprise data leakage rarely happens because of malicious penetration. It happens because of lazy infrastructure routing: an engineer stands up a prototype cluster, points a retrieval script to an external API endpoint, and leaves a default route 0.0.0.0/0 -> nat-xxxxxx active in the subnet's route table.

To achieve mathematically verifiable data sovereignty, the VPC must be designed around Zero-Egress Isolation:

  1. Routing Topology: Compute and database subnets have no default route (0.0.0.0/0) configured in their route tables. Any packet destined for a non-VPC IP address is discarded at the hypervisor level.
  2. AWS PrivateLink & VPC Gateway Endpoints: Internal services communicate with necessary cloud platform primitives exclusively via VPC Endpoints:
  • com.amazonaws.[region].s3 (Gateway Endpoint for document object storage).
  • com.amazonaws.[region].kms (Interface Endpoint for Customer-Managed Keys encryption).
  • com.amazonaws.[region].ecr.api and ecr.dkr (Interface Endpoints for container image pulls).
  • com.amazonaws.[region].logs (Interface Endpoint for CloudWatch audit trails).
  1. Strict Ingress Filtering: Ingress is restricted to corporate VPN or Direct Connect (DX) peering via an internal Application Load Balancer terminating Mutual TLS (mTLS 1.3) with enterprise PKI certificates.
  2. DNS Blackholing: Amazon Route 53 Resolver is configured with a Private Hosted Zone and Route 53 Resolver DNS Firewall rules that block outbound resolution for all external domains, logging any unexpected DNS queries as immediate security anomalies.

2. Deterministic Ingestion: Eliminating Naive Chunking Breakdown#

The single most common point of failure in enterprise RAG systems is naive text chunking. When building a toy demo, splitting text every 500 characters using RecursiveCharacterTextSplitter appears to work. In an enterprise financial or legal environment, it causes catastrophic data corruption.

Consider a multi-column balance sheet or a commercial vendor invoice. A fixed character splitter cuts through the middle of an HTML or ASCII table, isolating numerical amounts on page 3 from their corresponding line-item descriptions on page 2. When the retrieval engine fetches that chunk, the LLM receives arbitrary numbers stripped of context, leading to immediate hallucinations.

sh
NAIVE CHUNKING FAILURE:
Chunk 1: 400 font-semibold">class="text-emerald-300">"...Total Operating Expenses 400 font-semibold">for Fiscal Year 2025 were detailed as follows: Salaries and Wages"
[CHUNK BOUNDARY BREAK]
Chunk 2: 400 font-semibold">class="text-emerald-300">"$14,820,000. Rent and Utilities: $2,140,000. Depreciation: $840,000..."
--> LLM Query: 400 font-semibold">class="text-emerald-300">"What were the salaries 400 font-semibold">for 2025?" 
--> Result: Retrieval mismatch or inverted column mapping.

Layout-Aware Deterministic Parsing#

Production private RAG requires structural, layout-aware parsing using libraries such as Docling, PyMuPDF, or pdfminer.six coupled with layout analysis models. The ingestion pipeline must preserve three structural primitives:

  1. Table Reconstruction: Every tabular structure is converted into clean Markdown grid syntax (| Header | Header |) accompanied by a programmatic summary injected directly into the table's context block.
  2. Hierarchical Document Breadcrumbs: Every chunk retains its document hierarchy metadata: Document Name > Section 4: Operational Disclosures > Subsection B: Capital Expenditures.
  3. Spatial Bounding Box Tracking: Each extracted chunk stores its bounding coordinates [page_number, x0, y0, x1, y1] as JSONB attributes in PostgreSQL, enabling front-end client applications to highlight the exact visual text snippet on the original PDF.

Production Document Ingestion Worker#

The following Python script illustrates a production layout-aware document ingestion worker running inside the compute subnet, extracting tabular structures and generating high-dimensional vectors:

python
400 font-semibold">import io
400 font-semibold">import json
400 font-semibold">import uuid
400 font-semibold">import psycopg2
400 font-semibold">from psycopg2.extras 400 font-semibold">import execute_values
400 font-semibold">import requests
400 font-semibold">from docling.document_converter 400 font-semibold">import DocumentConverter

  400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Configuration pointing to VPC internal endpoints
PG_DSN = 400 font-semibold">class="text-emerald-300">"postgresql:400 font-semibold">class="text-slate-500 italic400 font-semibold">class="text-emerald-300">">//rag_app:StrongVpcPassword@10.100.20.15:5432/enterprise_rag?sslmode=verify-full"
TEI_EMBED_URL = 400 font-semibold">class="text-emerald-300">"http:400 font-semibold">class="text-slate-500 italic400 font-semibold">class="text-emerald-300">">//tei-bge-m3.compute.internal:8080/embed"

converter = DocumentConverter()

400 font-semibold">def extract_structured_chunks(pdf_bytes: bytes, doc_id: str, tenant_id: str):
    400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
    Parses complex multi-column PDFs into layout-preserved semantic chunks.
    Preserves tables as pristine markdown structures with spatial coordinates.
    "400 font-semibold">class="text-emerald-300">""
    result = converter.convert(io.BytesIO(pdf_bytes))
    doc = result.document
    
    chunks = []
    current_section = 400 font-semibold">class="text-emerald-300">"Introduction"
    
    400 font-semibold">for item in doc.iterate_items():
        400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Track hierarchical headings
        400 font-semibold">if item.label in [400 font-semibold">class="text-emerald-300">"section_header", 400 font-semibold">class="text-emerald-300">"heading"]:
            current_section = item.text.strip()
            continue
            
        400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Deterministic table preservation
        400 font-semibold">if item.label == 400 font-semibold">class="text-emerald-300">"table":
            table_md = item.export_to_markdown()
            chunk_text = f400 font-semibold">class="text-emerald-300">"Context: {current_section}\n\n[400 font-semibold">TABLE DATA]:\n{table_md}"
            bbox = item.prov[0].bbox 400 font-semibold">if item.prov 400 font-semibold">else None
            page_no = item.prov[0].page_no 400 font-semibold">if item.prov 400 font-semibold">else 1
            
            chunks.append({
                400 font-semibold">class="text-emerald-300">"chunk_id": str(uuid.uuid4()),
                400 font-semibold">class="text-emerald-300">"doc_id": doc_id,
                400 font-semibold">class="text-emerald-300">"tenant_id": tenant_id,
                400 font-semibold">class="text-emerald-300">"section": current_section,
                400 font-semibold">class="text-emerald-300">"content": chunk_text,
                400 font-semibold">class="text-emerald-300">"is_table": True,
                400 font-semibold">class="text-emerald-300">"page_number": page_no,
                400 font-semibold">class="text-emerald-300">"bbox": {400 font-semibold">class="text-emerald-300">"x0": bbox.l, 400 font-semibold">class="text-emerald-300">"y0": bbox.t, 400 font-semibold">class="text-emerald-300">"x1": bbox.r, 400 font-semibold">class="text-emerald-300">"y1": bbox.b} 400 font-semibold">if bbox 400 font-semibold">else {}
            })
            
        elif item.label in [400 font-semibold">class="text-emerald-300">"paragraph", 400 font-semibold">class="text-emerald-300">"text"] and len(item.text.strip()) > 40:
            chunk_text = f400 font-semibold">class="text-emerald-300">"Context: {current_section}\n\n{item.text.strip()}"
            bbox = item.prov[0].bbox 400 font-semibold">if item.prov 400 font-semibold">else None
            page_no = item.prov[0].page_no 400 font-semibold">if item.prov 400 font-semibold">else 1
            
            chunks.append({
                400 font-semibold">class="text-emerald-300">"chunk_id": str(uuid.uuid4()),
                400 font-semibold">class="text-emerald-300">"doc_id": doc_id,
                400 font-semibold">class="text-emerald-300">"tenant_id": tenant_id,
                400 font-semibold">class="text-emerald-300">"section": current_section,
                400 font-semibold">class="text-emerald-300">"content": chunk_text,
                400 font-semibold">class="text-emerald-300">"is_table": False,
                400 font-semibold">class="text-emerald-300">"page_number": page_no,
                400 font-semibold">class="text-emerald-300">"bbox": {400 font-semibold">class="text-emerald-300">"x0": bbox.l, 400 font-semibold">class="text-emerald-300">"y0": bbox.t, 400 font-semibold">class="text-emerald-300">"x1": bbox.r, 400 font-semibold">class="text-emerald-300">"y1": bbox.b} 400 font-semibold">if bbox 400 font-semibold">else {}
            })
            
    400 font-semibold">return chunks

400 font-semibold">def batch_embed_chunks(chunks: list[dict]) -> list[list[float]]:
    400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
    Calls local Hugging Face Text Embeddings Inference (TEI) service.
    Zero egress: traffic stays on the local 10.100.10.x subnet.
    "400 font-semibold">class="text-emerald-300">""
    texts = [c[400 font-semibold">class="text-emerald-300">"content"] 400 font-semibold">for c in chunks]
    response = requests.post(
        TEI_EMBED_URL,
        json={400 font-semibold">class="text-emerald-300">"inputs": texts, 400 font-semibold">class="text-emerald-300">"truncate": True},
        timeout=15
    )
    response.raise_for_status()
    400 font-semibold">return response.json()

400 font-semibold">def persist_to_pgvector(chunks: list[dict], embeddings: list[list[float]]):
    400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
    Atomic multi-row insertion into PostgreSQL 16 with pgvector.
    Updates full-text tsvector automatically via SQL trigger.
    "400 font-semibold">class="text-emerald-300">""
    records = []
    400 font-semibold">for c, emb in zip(chunks, embeddings):
        records.append((
            c[400 font-semibold">class="text-emerald-300">"chunk_id"],
            c[400 font-semibold">class="text-emerald-300">"doc_id"],
            c[400 font-semibold">class="text-emerald-300">"tenant_id"],
            c[400 font-semibold">class="text-emerald-300">"section"],
            c[400 font-semibold">class="text-emerald-300">"content"],
            c[400 font-semibold">class="text-emerald-300">"is_table"],
            c[400 font-semibold">class="text-emerald-300">"page_number"],
            json.dumps(c[400 font-semibold">class="text-emerald-300">"bbox"]),
            emb
        ))
        
    query = 400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
    400 font-semibold">INSERT INTO document_chunks (
        id, doc_id, tenant_id, section_name, content, 
        is_table, page_number, bbox_metadata, embedding
    ) VALUES %s
    ON CONFLICT (id) DO NOTHING;
    "400 font-semibold">class="text-emerald-300">""
    
    with psycopg2.connect(PG_DSN) as conn:
        with conn.cursor() as cur:
            execute_values(cur, query, records, template=400 font-semibold">class="text-emerald-300">"(%s, %s, %s, %s, %s, %s, %s, %s, %s::vector)")
        conn.commit()

[Visual Asset: Deterministic Ingestion & Hybrid Retrieval Pipeline - From Layout Extraction to RRF]

Exact Visual Specification: A detailed end-to-end dataflow diagram tracing an unstructured PDF invoice through parsing, dual-indexing, hybrid retrieval, reciprocal rank fusion (RRF), cross-encoder re-ranking, and local quantized LLM generation. Highlights the contrast between dense vector semantics and sparse BM25 exact matching.

mermaid
flowchart LR
    subgraph Ingestion_Stage [400 font-semibold">class="text-emerald-300">"1. Layout Ingestion Pipeline"]
        PDF[400 font-semibold">class="text-emerald-300">"Raw Financial PDF / Invoice"] --> Parser[400 font-semibold">class="text-emerald-300">"Layout-Aware Parser<br/>(Table Grid Extraction)"]
        Parser --> Chunker[400 font-semibold">class="text-emerald-300">"Semantic Hierarchy Chunking<br/>(Markdown Tables + Breadcrumbs)"]
        Chunker --> LocalEmbed[400 font-semibold">class="text-emerald-300">"Local BGE-M3 Service<br/>(Dense 1024-dim Vector)"]
        Chunker --> TextIndex[400 font-semibold">class="text-emerald-300">"Postgres tsvector<br/>(Sparse BM25 Tokens)"]
    end

    subgraph Storage_Stage [400 font-semibold">class="text-emerald-300">"2. Unified pgvector Storage"]
        LocalEmbed --> PG_HNSW[(400 font-semibold">class="text-emerald-300">"pgvector HNSW Index<br/>(Cosine Distance)")]
        TextIndex --> PG_GIN[(400 font-semibold">class="text-emerald-300">"PostgreSQL GIN Index<br/>(Lexical Matching)")]
    end

    subgraph Retrieval_Stage [400 font-semibold">class="text-emerald-300">"3. Hybrid Search & Reranking"]
        Query[400 font-semibold">class="text-emerald-300">"Enterprise User Query"] --> DenseSearch[400 font-semibold">class="text-emerald-300">"Dense HNSW Search<br/>(Top 50 Candidates)"]
        Query --> SparseSearch[400 font-semibold">class="text-emerald-300">"BM25 Lexical Search<br/>(Top 50 Candidates)"]
        PG_HNSW -.-> DenseSearch
        PG_GIN -.-> SparseSearch
        DenseSearch --> RRF[400 font-semibold">class="text-emerald-300">"Reciprocal Rank Fusion (RRF)<br/>Score = 1/(60 + Rank_Dense) + 1/(60 + Rank_Sparse)"]
        SparseSearch --> RRF
        RRF --> TopPool[400 font-semibold">class="text-emerald-300">"Top 30 Merged Candidates"]
        TopPool --> Reranker[400 font-semibold">class="text-emerald-300">"Local bge-reranker-large<br/>(Cross-Encoder Scoring)"]
        Reranker --> TopK[400 font-semibold">class="text-emerald-300">"Top 5 High-Precision Chunks"]
    end

    subgraph Generation_Stage [400 font-semibold">class="text-emerald-300">"4. Sovereign Synthesis"]
        TopK --> ContextAssembler[400 font-semibold">class="text-emerald-300">"Strict Prompt Assembler<br/>(Bounding Box Citations)"]
        ContextAssembler --> LocalLLM[400 font-semibold">class="text-emerald-300">"Air-Gapped vLLM<br/>(Llama-3.1-70B-AWQ)"]
        LocalLLM --> VerifiedOutput[400 font-semibold">class="text-emerald-300">"Grounded JSON Output<br/>+ Page Coordinate Highlights"]
    end

sh
+─────────────────────────────────────────────────────────────────────────────────────────────────+
|               DETERMINISTIC INGESTION & HYBRID RETRIEVAL PIPELINE (ZERO EGRESS)                 |
+─────────────────────────────────────────────────────────────────────────────────────────────────+
|                                                                                                 |
| [1. Ingestion]                                                                                  |
|   PDF Invoice ──► Layout-Aware Parser ──► Markdown Tables ──► Parent-Child Chunks               |
|                                                                 │                               |
|                                ┌────────────────────────────────┴──────────────────────────┐    |
|                                ▼ (1024-dim Embedding)                                      ▼    |
|                     [Local BGE-M3 Engine]                               [tsvector Generator]    |
|                                │                                                   │            |
| [2. Storage]                   ▼                                                   ▼            |
|                     PostgreSQL pgvector (HNSW)                          PostgreSQL GIN (BM25)   |
|                                │                                                   │            |
|                                └────────────────────────────────┬──────────────────┘            |
|                                                                 │                               |
| [3. Retrieval] Enterprise Query ────────────────────────────────┼────────────────────────┐      |
|                                                                 ▼                        ▼      |
|                                                     Dense Top-50 Search         Sparse Top-50   |
|                                                                 │                        │      |
|                                                                 └───────────┬────────────┘      |
|                                                                             ▼                   |
|                                                               Reciprocal Rank Fusion (RRF)      |
|                                                                             │                   |
|                                                                             ▼                   |
|                                                               Local Cross-Encoder Reranker      |
|                                                               (bge-reranker-large: Top 5)       |
|                                                                             │                   |
| [4. Generation]                                                             ▼                   |
|   Air-Gapped vLLM (Llama-3.1-70B-AWQ) ◄── Grounded Prompt ◄── Context Assembly & Citations     |
|             │                                                                                   |
|             ▼                                                                                   |
|   Deterministic Output + Exact Page Bounding Box Highlights                                     |
+─────────────────────────────────────────────────────────────────────────────────────────────────+

Figure 2: End-to-end dataflow from layout extraction through dual-indexing, hybrid RRF scoring, cross-encoder pruning, and sovereign LLM synthesis.

3. Dual-Indexing with pgvector: Unifying BM25 & Dense Semantics#

A common architectural miscalculation is deploying a vector database with dense embeddings alone.

Dense semantic vectors excel at finding high-level concepts: querying "How do we handle contract termination?" matches paragraphs discussing "cancellation of service" or "severance of agreement" with remarkable accuracy.

However, in real-world enterprise documents, queries are frequently exact and lexical. If a finance analyst searches for:

  • An exact invoice reference: INV-2026-092A
  • An IRS tax schedule code: Form 1120-S Line 14
  • A specific product model: SKU-8921-XRT

Dense embeddings fail. The high-dimensional embedding maps these specific alphanumeric strings to generic coordinate neighborhoods, frequently surfacing the wrong invoice with high cosine similarity.

By using PostgreSQL 16 with the pgvector extension, enterprises achieve true hybrid retrieval in a single database engine without managing separate Elasticsearch or Pinecone clusters.

  1. pgvector with HNSW Indexing: Hierarchical Navigable Small World (HNSW) graphs deliver sub-millisecond approximate nearest neighbor (ANN) retrieval with high recall. Unlike IVFFlat, HNSW requires no initial clustering training phase and does not degrade under continuous write operations.
  2. PostgreSQL Full-Text Search (tsvector): Lexical indexing via Generalized Inverted Indexes (GIN) provides exact keyword matching and token proximity scoring.

PostgreSQL Schema & DDL#

sql
-- Enable required extensions
400 font-semibold">CREATE EXTENSION IF NOT EXISTS 400 font-semibold">class="text-emerald-300">"uuid-ossp";
400 font-semibold">CREATE EXTENSION IF NOT EXISTS 400 font-semibold">class="text-emerald-300">"vector";

-- Production chunks storage table
400 font-semibold">CREATE 400 font-semibold">TABLE document_chunks (
    id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    doc_id UUID NOT NULL,
    tenant_id VARCHAR(64) NOT NULL,
    clearance_level VARCHAR(32) NOT NULL DEFAULT 400 font-semibold">class="text-emerald-300">'standard',
    section_name VARCHAR(255) NOT NULL,
    content TEXT NOT NULL,
    content_tsv TSVECTOR GENERATED ALWAYS AS (to_tsvector(400 font-semibold">class="text-emerald-300">'english', content)) STORED,
    is_table BOOLEAN NOT NULL DEFAULT FALSE,
    page_number INTEGER NOT NULL,
    bbox_metadata JSONB NOT NULL DEFAULT 400 font-semibold">class="text-emerald-300">'{}'::jsonb,
    embedding VECTOR(1024) NOT NULL,
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);

-- HNSW Vector Index 400 font-semibold">for dense cosine similarity
-- m = 16 (bi-directional links per node), ef_construction = 128 (graph build quality)
400 font-semibold">CREATE 400 font-semibold">INDEX idx_chunks_hnsw_embedding 
ON document_chunks 
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 128);

-- GIN Index 400 font-semibold">for BM25-equivalent sparse full-text search
400 font-semibold">CREATE 400 font-semibold">INDEX idx_chunks_content_tsv 
ON document_chunks 
USING gin (content_tsv);

-- B-Tree Index 400 font-semibold">for multi-tenant isolation and security filtering
400 font-semibold">CREATE 400 font-semibold">INDEX idx_chunks_tenant_clearance 
ON document_chunks (tenant_id, clearance_level);

Reciprocal Rank Fusion (RRF) in a Single SQL Query#

Rather than running vector search and text search in separate application threads and manually stitching the results, PostgreSQL executes Reciprocal Rank Fusion (RRF) directly in the database engine using Common Table Expressions (CTEs):

Reciprocal Rank Fusion formula: RRF(d) = SUM(1 / (k + rank(d))) where k = 60.

sql
WITH 
-- 1. Dense Semantic Vector Retrieval (Top 50)
dense_results AS (
    400 font-semibold">SELECT 
        id, 
        ROW_NUMBER() OVER (400 font-semibold">ORDER BY embedding <=> :query_vector) AS dense_rank
    400 font-semibold">FROM document_chunks
    400 font-semibold">WHERE tenant_id = :tenant_id
      AND clearance_level <= :user_clearance
    400 font-semibold">ORDER BY embedding <=> :query_vector
    LIMIT 50
),
-- 2. Sparse Lexical Full-Text Retrieval (Top 50)
sparse_results AS (
    400 font-semibold">SELECT 
        id, 
        ROW_NUMBER() OVER (400 font-semibold">ORDER BY ts_rank_cd(content_tsv, plainto_tsquery(400 font-semibold">class="text-emerald-300">'english', :query_text)) DESC) AS sparse_rank
    400 font-semibold">FROM document_chunks
    400 font-semibold">WHERE tenant_id = :tenant_id
      AND clearance_level <= :user_clearance
      AND content_tsv @@ plainto_tsquery(400 font-semibold">class="text-emerald-300">'english', :query_text)
    400 font-semibold">ORDER BY ts_rank_cd(content_tsv, plainto_tsquery(400 font-semibold">class="text-emerald-300">'english', :query_text)) DESC
    LIMIT 50
)
-- 3. Reciprocal Rank Fusion Merge
400 font-semibold">SELECT 
    c.id,
    c.section_name,
    c.content,
    c.is_table,
    c.page_number,
    c.bbox_metadata,
    COALESCE(1.0 / (60 + d.dense_rank), 0.0) + 
    COALESCE(1.0 / (60 + s.sparse_rank), 0.0) AS rrf_score
400 font-semibold">FROM dense_results d
FULL OUTER 400 font-semibold">JOIN sparse_results s ON d.id = s.id
400 font-semibold">JOIN document_chunks c ON c.id = COALESCE(d.id, s.id)
400 font-semibold">ORDER BY rrf_score DESC
LIMIT 30;

4. Air-Gapped Local LLM Inference Engine#

To guarantee zero data egress, the model generating responses must execute entirely inside the private VPC compute subnet. Modern quantized open-weight foundation models—such as Llama-3.1-70B-Instruct or Qwen-2.5-72B-Instruct—match or exceed proprietary frontier models on structured enterprise document comprehension and question-answering benchmarks.

Sizing and Quantization Calculations#

Running an unquantized (FP16/BF16) 70B parameter model requires 70 x 2 = 140 GB of VRAM just to store the model weights, demanding a costly 4 x 80 GB A100 configuration.

By employing Activation-aware Weight Quantization (AWQ) at 4-bit precision:

  • Model weights compress to approximately 36 GB to 38 GB of VRAM.
  • Activation tensors remain in FP16 precision, preventing quality degradation.
  • Two NVIDIA A100 (80GB) or four NVIDIA L40S (48GB) provide ample headroom for both weights and a high-concurrency PagedAttention Key-Value (KV) cache.
MetricFP16 Baseline (70B)AWQ 4-Bit (70B)Architectural Impact
Model VRAM Footprint~142 GB~38 GB73% reduction in baseline memory
Hardware Required4x A100 (80GB)2x A100 (80GB) or 4x L40S (48GB)Halves GPU infrastructure cost
Tokens / Second22 tok/sec58 tok/sec>2.6x inference throughput improvement
Time-To-First-Token (TTFT)~480 ms~195 msSub-200ms prompt ingestion phase
Perplexity DegradationBaseline< 0.12%Mathematically indistinguishable recall

Production vLLM Container Deployment#

The following docker-compose.yml runs an enterprise vLLM server inside the private compute subnet. Notice that network access is strictly confined to an internal Docker network with DNS loopback overrides:

yaml
version: 400 font-semibold">class="text-emerald-300">"3.8"

services:
  tei-embeddings:
    image: ghcr.io/huggingface/text-embeddings-inference:turing-1.5
    container_name: tei-bge-m3
    restart: always
    environment:
      - MODEL_ID=BAAI/bge-m3
      - MAX_CLIENT_BATCH_SIZE=64
      - MAX_BATCH_TOKENS=16384
    volumes:
      - /opt/models/bge-m3:/data
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    networks:
      - vpc_internal

  vllm-inference:
    image: vllm/vllm-openai:v0.5.4
    container_name: vllm-llama3-70b
    restart: always
    environment:
      - NCCL_DEBUG=INFO
      - HF_HUB_OFFLINE=1  400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Strictly enforces zero external model repository pulls
    command: &gt;
      --model /opt/models/Meta-Llama-3.1-70B-Instruct-AWQ
      --quantization awq
      --tensor-parallel-size 2
      --max-model-len 8192
      --gpu-memory-utilization 0.92
      --enforce-eager
      --port 8000
    volumes:
      - /opt/models/Meta-Llama-3.1-70B-Instruct-AWQ:/opt/models/Meta-Llama-3.1-70B-Instruct-AWQ:ro
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 2
              capabilities: [gpu]
    networks:
      - vpc_internal

networks:
  vpc_internal:
    driver: bridge
    internal: 400">true  400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Prevents external network routing

5. Enterprise Governance: Database Row-Level Security (RLS)#

A fatal flaw in naive RAG implementations is trusting the LLM to enforce access permissions. Instructing a model: "Only answer if the user has permission to see salary data" fails consistently under adversarial prompt injection attacks.

In a sovereign architecture, authorization is enforced at the data retrieval layer using PostgreSQL Row-Level Security (RLS):

sql
-- Enable Row-Level Security on document chunks
400 font-semibold">ALTER 400 font-semibold">TABLE document_chunks ENABLE ROW LEVEL SECURITY;

-- Create policy matching user400 font-semibold">class="text-emerald-300">'s tenant ID and clearance level 400 font-semibold">from session variables
400 font-semibold">CREATE POLICY tenant_isolation_policy ON document_chunks
    FOR 400 font-semibold">SELECT
    USING (
        tenant_id = current_setting('app.current_tenant_id400 font-semibold">class="text-emerald-300">', 400">true)
        AND (
            clearance_level = '400 font-semibold">public400 font-semibold">class="text-emerald-300">'
            OR (clearance_level = 'internal400 font-semibold">class="text-emerald-300">' AND current_setting('app.user_role400 font-semibold">class="text-emerald-300">', 400">true) IN ('employee400 font-semibold">class="text-emerald-300">', 'manager400 font-semibold">class="text-emerald-300">', 'admin400 font-semibold">class="text-emerald-300">'))
            OR (clearance_level = 'confidential400 font-semibold">class="text-emerald-300">' AND current_setting('app.user_role400 font-semibold">class="text-emerald-300">', 400">true) IN ('manager400 font-semibold">class="text-emerald-300">', 'admin400 font-semibold">class="text-emerald-300">'))
            OR (clearance_level = 'restricted400 font-semibold">class="text-emerald-300">' AND current_setting('app.user_role400 font-semibold">class="text-emerald-300">', 400">true) = 'admin')
        )
    );

When an enterprise query enters the RAG orchestrator, the application acquires a database connection from the pool and sets transaction-scoped session claims:

python
400 font-semibold">def query_with_rls(tenant_id: str, user_role: str, query_vector: list[float], query_text: str):
    with db_pool.get_connection() as conn:
        with conn.cursor() as cur:
            400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Inject session security variables 400 font-semibold">for current transaction only
            cur.execute(400 font-semibold">class="text-emerald-300">"SET LOCAL app.current_tenant_id = %s;", (tenant_id,))
            cur.execute(400 font-semibold">class="text-emerald-300">"SET LOCAL app.user_role = %s;", (user_role,))
            
            400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Execute hybrid search with 100% RLS policy enforcement
            cur.execute(HYBRID_SEARCH_SQL, {400 font-semibold">class="text-emerald-300">"query_vector": query_vector, 400 font-semibold">class="text-emerald-300">"query_text": query_text})
            results = cur.fetchall()
            400 font-semibold">return results

Chunks that the requesting user is not authorized to read are excluded by the PostgreSQL query planner prior to index scanning. They never enter the candidate pool, never consume LLM context tokens, and cannot be leaked via prompt jailbreaks.

6. End-to-End Latency Budget & Verification Checklist#

To prove that private air-gapped RAG does not compromise user experience, our production benchmarking measures the complete round-trip latency across all subsystem hops:

Execution StageComponent & SubsystemTarget Latency (p50)Target Latency (p99)
1. Ingress & AuthALB mTLS Termination + JWT Signature Validation4 ms9 ms
2. Query EmbeddingLocal BGE-M3 (TEI on NVIDIA L4)14 ms22 ms
3. Hybrid Retrievalpgvector HNSW + BM25 tsvector + RRF CTE18 ms31 ms
4. Cross-EncoderLocal Cross-Encoder Reranking (Top 30 -> Top 5)32 ms48 ms
5. Time-To-First-TokenvLLM Llama-3.1-70B-AWQ (Prefill Phase)185 ms240 ms
6. Token StreamingGeneration of 200 Tokens @ 55 tokens/sec363 ms420 ms
Total PipelineEnd-to-End User Experience616 ms770 ms

Operational Handover & Egress Verification Protocol#

Before declaring an enterprise RAG cluster production-ready, engineering leads must execute the following egress audit commands on the compute host:

bash
  400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># 1. Audit active network routing table (Ensure no 400 font-semibold">default 0.0.0.0/0 route exists)
ip route show

  400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># 2. Monitor 400 font-semibold">for 400">any unauthorized egress connection attempts across network interfaces
sudo tcpdump -i 400">any -n 400 font-semibold">class="text-emerald-300">"not (src net 10.100.0.0/16 and dst net 10.100.0.0/16)" -c 50

  400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># 3. Verify pgvector HNSW index health and memory footprint
psql 400 font-semibold">class="text-emerald-300">"$PG_DSN" -c 400 font-semibold">class="text-emerald-300">"
400 font-semibold">SELECT 
    schemaname, tablename, indexname, 
    pg_size_pretty(pg_relation_size(indexrelid)) AS index_size
400 font-semibold">FROM pg_stat_user_indexes 
400 font-semibold">WHERE indexname LIKE '%hnsw%';
"

  400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># 4. Probe local vLLM health and GPU memory allocation
curl -s http:400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic">//10.100.10.25:8000/health | jq .
nvidia-smi --query-gpu=memory.total,memory.used,memory.free,utilization.gpu --format=csv

Accelerate Your Sovereign Enterprise AI Architecture#

Deploying enterprise AI does not require sacrificing data residency or risking customer privacy. By engineering a mathematically isolated, air-gapped retrieval architecture inside your private cloud perimeter, your organization gains all the strategic advantages of generative intelligence while adhering strictly to global compliance standards.

KNetwork's AI & Data Systems Practice architects, deploys, and optimizes sovereign LLM infrastructure, deterministic document processing pipelines, and high-throughput vector search for Fortune 500 enterprises and regulated institutions worldwide.

Book a Technical Discovery Briefing with Our Systems Architects or explore our Artificial Intelligence & Data Engineering Practice to audit your current AI data pipeline.

Frequently Asked Strategic Questions

Technical and architectural governance answers for enterprise leadership.

D

Danisur Rahman

Practice Lead

Lead Systems Architect • KNetwork Advisory

Schedule Advisory Briefing

Advises enterprise technical leadership, CTOs, and heads of engineering on enterprise modernization, cloud migration governance, high-concurrency ledger design, and sovereign artificial intelligence compliance.