Artificial Intelligence & DataEvaluating Retrieval Precision in RAG: Setting Up Continuous Unit Tests with Synthetic Queries

Evaluating Retrieval Precision in RAG: Setting Up Continuous Unit Tests with Synthetic Queries

Eliminate silent retrieval degradation in enterprise RAG pipelines: Mean Reciprocal Rank (MRR), Hit Rate @ K, nDCG evaluation, automated synthetic query generation with LLM critique filters, and CI/CD quality gates.

D

Danisur Rahman

Verified
Principal AI Systems Architect•Sep 30, 2026•13 min read
Evaluating Retrieval Precision in RAG: Setting Up Continuous Unit Tests with Synthetic Queries

In modern enterprise software engineering, code is never deployed to production without passing rigorous, automated CI/CD test suites: unit tests, integration tests, fuzzers, and static analysis linters. If an engineer modifies an authentication handler or an SQL migration, automated test runners verify deterministic correctness in seconds.

Yet, Retrieval-Augmented Generation (RAG) systems—which power mission-critical internal search engines, automated customer support agents, and regulatory compliance portals—are routinely deployed into production as unmonitored black boxes.

Engineering teams tweak chunking parameters (e.g., changing token limits from 512 to 768), upgrade embedding models (e.g., from bge-base-en-v1.5 to text-embedding-3-large), alter distance metrics, or adjust vector database HNSW indexing configurations without measuring the downstream impact. The result is silent retrieval degradation: vector clusters shift, previously well-ranked passages slip down the retrieval hierarchy, and generation models begin hallucinating on queries that had functioned flawlessly the week before.

High-assurance AI engineering requires treating retrieval precision with the same architectural rigor as transactional database unit testing. By constructing Continuous Retrieval Regression Test Suites powered by automatically synthesized golden question-chunk datasets, teams can quantitatively benchmark Mean Reciprocal Rank (MRR), Hit Rate @ K, and Normalized Discounted Cumulative Gain (nDCG@K) inside CI/CD pipelines, automatically blocking regressions before they reach production.

At KNetwork's AI Development practice, we build enterprise AI systems that operate with mathematical predictability. In this engineering guide, we dissect the quantitative metrics of information retrieval, formulate synthetic query generation pipelines, construct an automated pytest regression test runner, and implement continuous CI gating for vector search pipelines.

1. The Anatomy of Retrieval Drift: Why RAG Degrades#

Retrieval quality in a vector database is not static; it is subject to continuous environmental and systemic entropy:

sh
RAG System Degradation Vectors in Production:

┌─────────────────────────────────────────────────────────────┐
│ 1. CORPUS DILUTION (Document Ingestion Drift)               │
│   Ingesting 10,000 400 font-semibold">new PDFs dilutes existing vector dense   │
│   neighborhoods, pushing relevant passages past Top-K cutoff│
│                                                             │
│ 2. CHUNKING BOUNDARY REGRESSIONS                            │
│   Altering character overlap or token splitting strategies  │
│   severs contextual entities across chunk boundaries        │
│                                                             │
│ 3. EMBEDDING MODEL RE-INDEXING MISMATCH                     │
│   Upgrading embedding model versions without full database  │
│   re-vectorization creates cross-dimensional distance errors│
│                                                             │
│ 4. HNSW 400 font-semibold">INDEX APPROXIMATION TRADEOFFS                       │
│   Adjusting efSearch or M parameters to improve query speed │
│   drops recall on long-tail domain queries                  │
└─────────────────────────────────────────────────────────────┘

When retrieval quality drops, the Large Language Model has zero opportunity to recover. If the correct ground-truth chunk is ranked at position 12 and the RAG pipeline injects only the Top 5 into the system prompt (K=5), the LLM is starved of context and must hallucinate an answer.

Treating retrieval precision as an auditable software metric requires transforming human document evaluation into automated regression assertions.

2. Quantitative Retrieval Metrics: Mathematical Foundations#

To establish automated test assertions in CI/CD, we must define the mathematical metrics that quantify retrieval performance:

sh
Retrieval Rank Evaluation Metrics:

Query: 400 font-semibold">class="text-emerald-300">"What is the maximum payload weight 400 font-semibold">for drone model Titan-X?"
Ground-Truth Chunk: chunk_titan_x_specs_04

Retrieved Ranking List:
Rank 1: chunk_battery_life_02   ❌ (Irrelevant)
Rank 2: chunk_motor_specs_09    ❌ (Irrelevant)
Rank 3: chunk_titan_x_specs_04  ✅ (TARGET HIT!)
Rank 4: chunk_warranty_01       ❌ (Irrelevant)
Rank 5: chunk_accessories_12    ❌ (Irrelevant)

METRIC CALCULATIONS (At K = 5):
• Hit Rate @ 5: 1.0 (Target chunk is present in Top 5)
• Reciprocal Rank (RR): 1 / 3 = 0.333 (Target is at position 3)
• Precision @ 5: 1 / 5 = 0.200 (1 relevant document out of 5 retrieved)

1. Context Recall (Hit Rate @ K)#

Measures the proportion of queries where the true ground-truth document chunk appears anywhere within the Top-K retrieved candidates:

Mathematical Formulation
Hit Rate @ K = (1 / |Q|) ∑[q=1..|Q|] 𝕀(TargetChunk_q ∈ TopK_q)

Where \mathbb{I} is the indicator function. Hit Rate @ K is the primary safety threshold: if a target chunk is not in Top-K, generation is guaranteed to fail.

2. Mean Reciprocal Rank (MRR)#

While Hit Rate confirms presence, MRR measures where the relevant document ranks. Higher ranks mean higher LLM attention allocation and lower prompt dilution:

Mathematical Formulation
MRR = (1 / |Q|) ∑[q=1..|Q|] (1 / rank)_q

If the ground truth is at Rank 1, reciprocal rank is 1.0; if at Rank 2, it is 0.5; if not in Top-K, it is 0.0. An enterprise RAG pipeline should target MRR ≥ 0.85.

3. Normalized Discounted Cumulative Gain (nDCG@K)#

When queries possess multiple relevant documents with graded relevance (e.g., a primary technical specification chunk vs an auxiliary overview chunk), nDCG@K penalizes placing highly relevant documents lower down the ranking:

Mathematical Formulation
DCG_K = ∑[i=1..K] \frac{2^{rel_i} - 1}{\log_2(i + 1)}
Mathematical Formulation
nDCG_K = \frac{DCG_K}{IDCG_K}

Where IDCG_K is the Ideal Discounted Cumulative Gain (the theoretical maximum score achieved by an optimal ranking order).

3. Synthetic Golden Dataset Generation: The Reverse-Sampling Engine#

Manual curation of thousands of evaluation queries by human domain experts is prohibitively slow and expensive. Furthermore, as the corporate knowledge base updates weekly, human test sets quickly fall out of sync with new documents.

The solution is an Automated Synthetic Query Generator: reverse-sampling realistic, multi-perspective questions directly from document chunks using frontier LLMs.

sh
Synthetic Golden Dataset Generation Pipeline:

┌─────────────────────────────────────────────────────────────┐
│ Corpus Chunk: chunk_id: 400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item1a_p44"      │
│ 400 font-semibold">class="text-emerald-300">"In Q3 2026, foreign exchange volatility reduced European   │
│ segment gross margin by 140 basis points, primarily driven  │
│ by Euro and British Pound depreciation against the USD."    │
└─────────────────────────────┬───────────────────────────────┘
                              │
                Reverse-Sampling Prompting Engine
                              │
              ┌───────────────┼───────────────┐
              ▼               ▼               ▼
      [Direct Factual] [Complex Reason] [Adversarial Query]
      400 font-semibold">class="text-emerald-300">"How much did FX "Which currencies 400 font-semibold">class="text-emerald-300">"What caused 140bps
      dilute European  drove European   margin compression
      margins in Q3?"  margin loss?400 font-semibold">class="text-emerald-300">"    in Q3 2026?"
              │               │               │
              └───────────────┼───────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│ EVALUATION GOLDEN RECORD:                                   │
│ {                                                           │
│   400 font-semibold">class="text-emerald-300">"query": 400 font-semibold">class="text-emerald-300">"Which currencies drove European margin loss?",  │
│   400 font-semibold">class="text-emerald-300">"target_chunk_id": 400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item1a_p44",       │
│   400 font-semibold">class="text-emerald-300">"expected_doc_id": 400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3.pdf",               │
│   400 font-semibold">class="text-emerald-300">"query_type": 400 font-semibold">class="text-emerald-300">"reasoning",                                │
│   400 font-semibold">class="text-emerald-300">"relevance_score": 1.0                                    │
│ }                                                           │
└─────────────────────────────────────────────────────────────┘

Reverse-Prompting Archetypes:#

  1. Direct Factual Questions: Queries targeting specific numeric metrics, dates, or named entities.
  2. Paraphrased / Semantic Analogies: Queries that express the same conceptual question without sharing exact keywords with the text, testing the dense embedding model's latent mapping.
  3. Multi-Hop Synthesis Queries: Questions whose answers require synthesizing two adjacent or cross-referenced chunks.
  4. Adversarial Distractors: Queries constructed to mimic irrelevant chunks in the corpus, testing the vector database's ability to resist false positives.

Autonomous Quality Filtering: The LLM-as-a-Judge Filter#

Generating synthetic queries naively can produce noisy or trivial questions (e.g., "What is discussed in this section?"). To ensure test suite integrity, every generated question must pass an automated Critique Gate:

python
400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
Synthetic Question Generation and Quality Critique Filter.
Generates evaluation pairs 400 font-semibold">from raw chunks and validates question answerability.
"400 font-semibold">class="text-emerald-300">""

400 font-semibold">from typing 400 font-semibold">import List, Dict, Optional
400 font-semibold">import json

400 font-semibold">class SyntheticQueryGenerator:
    400 font-semibold">def __init__(self, llm_client):
        self.client = llm_client

    400 font-semibold">def generate_evaluation_pairs(self, chunk_text: str, chunk_id: str, doc_id: str) -> List[Dict[str, Any]]:
        400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
        Synthesizes diverse questions 400 font-semibold">from chunk text and filters via self-critique.
        "400 font-semibold">class="text-emerald-300">""
        generation_prompt = f400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
You are an expert red-team AI evaluation engineer.
Given the document passage below, generate 3 distinct user questions that can be answered
ONLY using the facts directly stated in 400 font-semibold">this passage.

Passage:
\"\"\"{chunk_text}\"\"\"

Generate a JSON object matching 400 font-semibold">this schema:
{{
  "questions400 font-semibold">class="text-emerald-300">": [
    {{"400 font-semibold">type400 font-semibold">class="text-emerald-300">": "factual400 font-semibold">class="text-emerald-300">", "question400 font-semibold">class="text-emerald-300">": "...400 font-semibold">class="text-emerald-300">"}},
    {{"400 font-semibold">type400 font-semibold">class="text-emerald-300">": "paraphrase400 font-semibold">class="text-emerald-300">", "question400 font-semibold">class="text-emerald-300">": "...400 font-semibold">class="text-emerald-300">"}},
    {{"400 font-semibold">type400 font-semibold">class="text-emerald-300">": "reasoning400 font-semibold">class="text-emerald-300">", "question400 font-semibold">class="text-emerald-300">": "...400 font-semibold">class="text-emerald-300">"}}
  ]
}}
"400 font-semibold">class="text-emerald-300">""
        response = self.client.generate(prompt=generation_prompt, response_format=400 font-semibold">class="text-emerald-300">"json")
        candidate_questions = json.loads(response)[400 font-semibold">class="text-emerald-300">"questions"]

        verified_pairs = []
        400 font-semibold">for item in candidate_questions:
            q = item[400 font-semibold">class="text-emerald-300">"question"]
            400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Critique Pass: Can 400 font-semibold">this question be answered without the text? Is it ambiguous?
            400 font-semibold">if self._verify_question_quality(question=q, context=chunk_text):
                verified_pairs.append({
                    400 font-semibold">class="text-emerald-300">"query": q,
                    400 font-semibold">class="text-emerald-300">"target_chunk_id": chunk_id,
                    400 font-semibold">class="text-emerald-300">"expected_doc_id": doc_id,
                    400 font-semibold">class="text-emerald-300">"category": item[400 font-semibold">class="text-emerald-300">"400 font-semibold">type"]
                })

        400 font-semibold">return verified_pairs

    400 font-semibold">def _verify_question_quality(self, question: str, context: str) -> bool:
        400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"Verifies question is specific, self-contained, and provably answered by the chunk."400 font-semibold">class="text-emerald-300">""
        critique_prompt = f400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
Evaluate 400 font-semibold">if the following question is self-contained (does not use vague pronouns like '400 font-semibold">this document' or 'the table')
and is definitively answered by the provided text.
Question: "{question}400 font-semibold">class="text-emerald-300">"
Text: \"\"\"{context}\"\"\"

Respond with JSON: {{"is_high_quality400 font-semibold">class="text-emerald-300">": 400">true/400">false, "reason400 font-semibold">class="text-emerald-300">": "...400 font-semibold">class="text-emerald-300">"}}
"400 font-semibold">class="text-emerald-300">""
        eval_resp = self.client.generate(prompt=critique_prompt, response_format=400 font-semibold">class="text-emerald-300">"json")
        400 font-semibold">return json.loads(eval_resp).get(400 font-semibold">class="text-emerald-300">"is_high_quality", False)

4. Production Python Implementation: The RAG CI/CD Unit Test Suite#

Below is a complete, production-ready Python regression testing suite integrating synthetic dataset generation, vector database querying (Qdrant), metric calculation, and pytest regression assertions.

python
400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
Continuous Retrieval Evaluation Suite 400 font-semibold">for RAG Systems.
Calculates Hit Rate @ K, MRR, and nDCG, and executes automated pytest CI assertions.
"400 font-semibold">class="text-emerald-300">""

400 font-semibold">import math
400 font-semibold">from typing 400 font-semibold">import List, Dict, Any, Optional
400 font-semibold">from dataclasses 400 font-semibold">import dataclass
400 font-semibold">import pytest

@dataclass
400 font-semibold">class GoldenEvaluationPair:
    query: str
    target_chunk_id: str
    expected_doc_id: str
    category: str

@dataclass
400 font-semibold">class RetrievalResult:
    chunk_id: str
    score: float
    rank: int

400 font-semibold">class RetrievalEvaluator:
    @staticmethod
    400 font-semibold">def calculate_hit_rate(retrieved: List[RetrievalResult], target_id: str, k: int) -> float:
        400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"Returns 1.0 400 font-semibold">if target_id is present within Top K retrieved candidates, 400 font-semibold">else 0.0."400 font-semibold">class="text-emerald-300">""
        top_k = retrieved[:k]
        400 font-semibold">return 1.0 400 font-semibold">if 400">any(r.chunk_id == target_id 400 font-semibold">for r in top_k) 400 font-semibold">else 0.0

    @staticmethod
    400 font-semibold">def calculate_reciprocal_rank(retrieved: List[RetrievalResult], target_id: str, k: int) -> float:
        400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"Calculates 1 / rank 400 font-semibold">if target_id is within Top K, 400 font-semibold">else 0.0."400 font-semibold">class="text-emerald-300">""
        400 font-semibold">for rank_idx, result in enumerate(retrieved[:k], start=1):
            400 font-semibold">if result.chunk_id == target_id:
                400 font-semibold">return 1.0 / rank_idx
        400 font-semibold">return 0.0

    @staticmethod
    400 font-semibold">def calculate_ndcg(retrieved: List[RetrievalResult], target_id: str, k: int) -> float:
        400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"Calculates nDCG@K 400 font-semibold">for binary relevance (1 relevant target)."400 font-semibold">class="text-emerald-300">""
        dcg = 0.0
        400 font-semibold">for rank_idx, result in enumerate(retrieved[:k], start=1):
            400 font-semibold">if result.chunk_id == target_id:
                dcg = 1.0 / math.log2(rank_idx + 1)
                break
        
        400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># For single relevant item, ideal DCG is 1.0 / log2(1 + 1) = 1.0
        idcg = 1.0 / math.log2(2)
        400 font-semibold">return dcg / idcg

    400 font-semibold">def evaluate_test_suite(
        self,
        retrieval_engine,
        golden_dataset: List[GoldenEvaluationPair],
        k: int = 5
    ) -> Dict[str, float]:
        400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
        Executes all evaluation queries against the retrieval engine and returns aggregate metrics.
        "400 font-semibold">class="text-emerald-300">""
        total_hits = 0.0
        total_rr = 0.0
        total_ndcg = 0.0
        n_queries = len(golden_dataset)

        400 font-semibold">for pair in golden_dataset:
            400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Query the retrieval pipeline (Embedding + Vector Search + Reranker)
            retrieved_chunks = retrieval_engine.retrieve(query=pair.query, top_k=k)

            total_hits += self.calculate_hit_rate(retrieved_chunks, pair.target_chunk_id, k)
            total_rr += self.calculate_reciprocal_rank(retrieved_chunks, pair.target_chunk_id, k)
            total_ndcg += self.calculate_ndcg(retrieved_chunks, pair.target_chunk_id, k)

        400 font-semibold">return {
            f400 font-semibold">class="text-emerald-300">"hit_rate@{k}": total_hits / n_queries,
            f400 font-semibold">class="text-emerald-300">"mrr@{k}": total_rr / n_queries,
            f400 font-semibold">class="text-emerald-300">"ndcg@{k}": total_ndcg / n_queries,
            400 font-semibold">class="text-emerald-300">"sample_size": n_queries
        }

400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># --- Mock Vector Retrieval Engine 400 font-semibold">for Test Simulation ---

400 font-semibold">class MockVectorRetrievalPipeline:
    400 font-semibold">def __init__(self, accuracy_profile: str = 400 font-semibold">class="text-emerald-300">"high_precision"):
        self.accuracy_profile = accuracy_profile

    400 font-semibold">def retrieve(self, query: str, top_k: int = 5) -> List[RetrievalResult]:
        400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"Simulates vector search retrieval results."400 font-semibold">class="text-emerald-300">""
        400 font-semibold">if 400 font-semibold">class="text-emerald-300">"foreign exchange" in query.lower() or 400 font-semibold">class="text-emerald-300">"currencies" in query.lower():
            400 font-semibold">if self.accuracy_profile == 400 font-semibold">class="text-emerald-300">"high_precision":
                400 font-semibold">return [
                    RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item1a_p44", score=0.92, rank=1),
                    RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item7_p12", score=0.84, rank=2),
                    RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q2_item1a_p40", score=0.79, rank=3),
                ]
            400 font-semibold">else:
                400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># Degraded retrieval simulation (target ranked 3rd)
                400 font-semibold">return [
                    RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item7_p12", score=0.88, rank=1),
                    RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2025_q3_item1a_p42", score=0.81, rank=2),
                    RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item1a_p44", score=0.75, rank=3),
                ]

        400 font-semibold">return [
            RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"irrelevant_chunk_01", score=0.50, rank=1),
            RetrievalResult(chunk_id=400 font-semibold">class="text-emerald-300">"irrelevant_chunk_02", score=0.45, rank=2),
        ]

400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># --- Automated Pytest Regression Test Suite ---

GOLDEN_TEST_SUITE = [
    GoldenEvaluationPair(
        query=400 font-semibold">class="text-emerald-300">"Which currencies drove European margin loss in Q3 2026?",
        target_chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item1a_p44",
        expected_doc_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3.pdf",
        category=400 font-semibold">class="text-emerald-300">"financial_fx"
    ),
    GoldenEvaluationPair(
        query=400 font-semibold">class="text-emerald-300">"Foreign exchange impact on European gross margins",
        target_chunk_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3_item1a_p44",
        expected_doc_id=400 font-semibold">class="text-emerald-300">"sec_apple_2026_q3.pdf",
        category=400 font-semibold">class="text-emerald-300">"financial_fx"
    )
]

400 font-semibold">def test_rag_retrieval_meets_ci_performance_thresholds():
    400 font-semibold">class="text-emerald-300">""400 font-semibold">class="text-emerald-300">"
    CI/CD Gate: Verifies that vector retrieval meets strict production precision baselines.
    Blocks pull requests 400 font-semibold">if MRR drops below 0.85 or Hit Rate @ 5 drops below 0.95.
    "400 font-semibold">class="text-emerald-300">""
    evaluator = RetrievalEvaluator()
    production_pipeline = MockVectorRetrievalPipeline(accuracy_profile=400 font-semibold">class="text-emerald-300">"high_precision")

    metrics = evaluator.evaluate_test_suite(
        retrieval_engine=production_pipeline,
        golden_dataset=GOLDEN_TEST_SUITE,
        k=5
    )

    print(f400 font-semibold">class="text-emerald-300">"\n[CI EVALUATION RESULTS]: {metrics}")

    400 font-semibold">class=400 font-semibold">class="text-emerald-300">"text-slate-500 italic"># PRODUCTION QUALITY ASSERTIONS
    assert metrics[400 font-semibold">class="text-emerald-300">"hit_rate@5"] >= 0.95, (
        f400 font-semibold">class="text-emerald-300">"CRITICAL REGRESSION: Hit Rate @ 5 is {metrics['hit_rate@5']:.3f}, expected >= 0.950"
    )
    assert metrics[400 font-semibold">class="text-emerald-300">"mrr@5"] >= 0.85, (
        f400 font-semibold">class="text-emerald-300">"CRITICAL REGRESSION: MRR @ 5 is {metrics['mrr@5']:.3f}, expected >= 0.850"
    )
    assert metrics[400 font-semibold">class="text-emerald-300">"ndcg@5"] >= 0.85, (
        f400 font-semibold">class="text-emerald-300">"CRITICAL REGRESSION: nDCG @ 5 is {metrics['ndcg@5']:.3f}, expected >= 0.850"
    )

400 font-semibold">if __name__ == 400 font-semibold">class="text-emerald-300">"__main__":
    pytest.main([400 font-semibold">class="text-emerald-300">"-s", __file__])

5. Integrating Retrieval Tests into GitHub Actions CI/CD Pipeline#

To ensure that pull requests modifying chunking logic, vector database clients, or embedding configurations cannot be merged if retrieval performance regresses, we configure a dedicated GitHub Actions workflow:

yaml
name: Continuous RAG Retrieval Quality Gate

on:
  pull_request:
    paths:
      - 400 font-semibold">class="text-emerald-300">'ai/chunking/**'
      - 400 font-semibold">class="text-emerald-300">'ai/embeddings/**'
      - 400 font-semibold">class="text-emerald-300">'ai/retrieval/**'
      - 400 font-semibold">class="text-emerald-300">'tests/retrieval/**'

jobs:
  evaluate-retrieval-precision:
    runs-on: ubuntu-latest
    steps:
      - name: Checkout Source Code
        uses: actions/checkout@v4

      - name: Setup Python 3.11
        uses: actions/setup-python@v5
        with:
          python-version: 400 font-semibold">class="text-emerald-300">'3.11'
          cache: 400 font-semibold">class="text-emerald-300">'pip'

      - name: Install Dependencies
        run: |
          pip install --upgrade pip
          pip install pytest qdrant-client sentence-transformers

      - name: Start Ephemeral Vector DB Service
        run: |
          docker run -d -p 6333:6333 qdrant/qdrant:latest
          sleep 3

      - name: Seed Ephemeral Vector Index with Golden Chunks
        run: python -m tests.retrieval.seed_test_index

      - name: Run Retrieval Regression Benchmark
        run: |
          pytest -v tests/retrieval/test_retrieval_precision.py \
            --junitxml=reports/retrieval-results.xml

      - name: Publish Test Report
        uses: EnricoMi/publish-unit-test-result-action@v2
        400 font-semibold">if: always()
        with:
          files: reports/retrieval-results.xml

If an engineer introduces a change that drops MRR from 0.92 to 0.78, the GitHub Actions job fails immediately. The pull request cannot be merged until the retrieval parameters are calibrated or the indexing defect is resolved.

6. Continuous Golden Set Maintenance and Ground-Truth Refresh#

A static test suite inevitably becomes stale as new enterprise documentation is ingested. To maintain test validity:

sh
Automated Test Suite Maintenance Loop:

┌─────────────────────────────────────────────────────────────┐
│ 1. WEEKLY CORPUS AUDIT                                      │
│   Detect newly added documents and updated revisions        │
└─────────────────────────────┬───────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│ 2. AUTONOMOUS SYNTHETIC EXPANSION                           │
│   Worker invokes Claude 3.5 / GPT-4o to synthesize 5 Q/A    │
│   pairs per 100 400 font-semibold">new chunks                                  │
└─────────────────────────────┬───────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│ 3. PRODUCTION HARD-QUERY HARVESTING                         │
│   Query logs filtered: User queries with negative feedback  │
│   (400 font-semibold">class="text-emerald-300">"Answer was incorrect") are sanitized and labeled        │
└─────────────────────────────┬───────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│ 4. GOLDEN SUITE VERSIONING (Git LFS / S3)                   │
│   Updated evaluation dataset tagged: eval_golden_v2.4.json  │
└─────────────────────────────────────────────────────────────┘

  1. Production Hard-Query Ingestion: Real user queries that received negative feedback ("thumbs down" or regeneration requests) are automatically flagged, reviewed, ground-truth labeled, and added to the regression test suite.
  2. Deterministic Seed Locks: Ensure synthetic query generation uses deterministic system seeds (seed=42) and temperature 0.0 so test fixtures remain reproducible across CI runs.
  3. Sub-Domain Stratification: Maintain evaluation buckets across specific domain categories (e.g., Legal, Technical, Financial, HR). A regression test must report precision per category: an overall high score should not hide a 40% precision drop in technical engineering specs.

7. Architectural Checklist for RAG Evaluation Pipelines#

Before promoting a RAG system to production, ensure your testing infrastructure implements these quality controls:

  • [ ] Establish Automated CI Gate: Run retrieval unit tests on every pull request that modifies chunking strategies, embeddings, or vector search configurations.
  • [ ] Track MRR and Hit Rate @ K: Enforce minimum CI assertions (Hit Rate @ 5 ≥ 0.95, MRR @ 5 ≥ 0.85). Never rely solely on qualitative spot-checks.
  • [ ] Decouple Retrieval Testing from Generation: Test retrieval precision independently of LLM synthesis. If retrieval is broken, evaluating generation quality is meaningless.
  • [ ] Automate Synthetic Question Generation: Use multi-perspective prompting (factual, reasoning, and adversarial) to generate diverse golden datasets from source chunks.
  • [ ] Harvest Production Edge Cases: Capture failed queries from production audit logs and incorporate them into the permanent regression suite.
  • [ ] Stratify Evaluation by Category: Ensure precision metrics are tracked independently across distinct business domains to catch isolated topic regressions.

By treating retrieval precision as an automated, measurable, and continuous software engineering discipline, organizations eliminate silent RAG degradation and maintain provable accuracy across enterprise AI applications.

Frequently Asked Questions

Key questions answered regarding this architectural implementation.

D

Danisur Rahman

Lead Author

Principal AI Systems Architect • KNetwork Systems

Request Technical Review

Principal architect specializing in enterprise distributed systems, edge caching, and hardware integration pipelines. Leads engineering audits, high-concurrency database optimizations, and zero-trust VPC deployments across high-growth ventures.

Distributed BackendsEvent StreamingPrivate RAGIoT Telemetry
The Engineering Dispatch

Enjoyed this technical breakdown?

Subscribe to receive new architectural guides, system teardowns, and engineering benchmarks directly in your inbox.