Legal research is an information-heavy process.
For advocates and corporate legal teams, finding the right provision, case law, notification, circular, commentary, or procedure is often only the first step. The information also needs to be relevant to the question, grounded in the correct source, and presented with enough context to support the final answer.
This creates an interesting challenge for Generative AI.
A simple chatbot that sends a question to an LLM is not enough for legal applications.
The system needs reliable retrieval, structured legal data handling, source citations, protection against unsupported answers, and continuous evaluation.
We recently built a production-grade Agentic RAG application for the Indian legal and corporate domain to address this problem.
The application works with structured and semi-structured Indian legal data along with legal books and documents, and is designed to help advocates and corporate teams interact with this information using natural-language queries.
This article explains how we approached the problem, from data preparation and retrieval to RAG evaluation, regression testing, and production monitoring.

The objective was not simply to build a legal chatbot.
The goal was to build an AI system that could understand a user’s legal or corporate query, identify the relevant sources, retrieve the right information, and generate a grounded response with appropriate references.
A typical interaction looks like:
User Question → Query Understanding → Retrieval → Reranking → Context → LLM → Answer + Citation
For example, a user may ask a question related to:
The challenge is that the required information may not exist in one place.
A good answer may require information from legislation as well as case laws, notifications, circulars, or books.
This is where a carefully designed RAG architecture becomes important.
One of the biggest challenges was the nature of the source data.
A large part of the legal information was available in database tables rather than clean documents.
The application works with multiple independent legal data sources, including tables such as:
These tables contain text-heavy legal information and are not simply connected through a single relational structure that can be directly passed to an LLM.
Alongside this structured and semi-structured data, we also had legal books and PDF-based content.
This meant that we needed two different retrieval approaches:
Database-based retrieval for structured legal information.
Vector-based document retrieval for books and PDF content.
Before building the retrieval layer, the source data needed to be prepared carefully.
We migrated the relevant database data from SQL Server into PostgreSQL/Neon and designed the retrieval layer around PostgreSQL with PGVector.
For the book and PDF content, we used a separate vector retrieval approach with Pinecone.
The objective was to preserve important legal metadata while making the content searchable using semantic retrieval.
Data cleaning was also an important part of this stage.
We handled issues such as:
The goal was not simply to create embeddings.
The goal was to create searchable legal knowledge while preserving the information required later for citations and source verification.
Chunking is particularly important in legal RAG.
If chunks are too large, retrieval can bring unnecessary information and increase the context size.
If chunks are too small, important legal context can be separated.
Instead of applying the same chunking strategy to every table, we used a table-aware approach.
We treated database records as parent-level units and created child chunks when individual fields were too large.
For example, long legal fields such as case-law headnotes or detailed legislation headings required further splitting.
Other fields could often remain as a complete record because their structure already represented a meaningful unit of information.
This created a parent-child relationship:
Parent Record
↓
Relevant Field
↓
Child Chunks
Metadata was preserved across the chunks so that retrieved content could still be connected back to its original source.
This is important for legal applications because retrieval should not lose the identity or context of the original record.
For books and PDF documents, the chunking process was designed around the document structure so that retrieved sections could be traced back to the relevant book content.
The retrieval layer was designed as a combination of semantic and keyword-based search.
Pure semantic search can understand meaning well, but legal queries often contain exact terms, section names, case references, legislation names, and other terminology where keyword matching is extremely useful.
Therefore, we used a hybrid approach:
Semantic Retrieval + Keyword Retrieval
The retrieval strategy gave more weight to semantic understanding while still using keyword-based matching to improve exact-term retrieval.
BM25 was also incorporated for keyword-based retrieval.
This combination helps address two different types of legal queries:
“Explain the rules related to this topic.”
and
“Find information related to Section X of a specific Act.”
Both require different retrieval signals.
Another important part of the retrieval strategy was prioritizing legislation.
For many legal questions, the applicable legislation provides the primary foundation for the answer.
The retrieval process therefore gives priority to legislation-related results and then brings supporting information from other sources.
A simplified strategy looks like:
User Query
↓
Retrieve top legislation results
↓
Retrieve relevant results from other legal tables
↓
Combine results
↓
Rerank
↓
Send the best context to the LLM
The system was designed around a limited context window rather than simply passing a large number of retrieved chunks to the model.
We used a maximum context limit of 32 chunks, with legislation receiving priority and a smaller number of relevant chunks coming from other sources.
This helps control noise and token usage while keeping the most important legal information in context.
Retrieval gives us candidate information.
But candidate information is not necessarily the best information.
This is where reranking becomes useful.
After the initial retrieval stage, the retrieved results are reranked using Cohere reranking.
The idea is simple:
Initial Retrieval → Candidate Results → Reranking → Best Context
The reranker helps identify which retrieved chunks are most relevant to the actual user question.
This becomes especially important when several legal documents contain similar terminology but only a few directly address the question.
Legal books introduced another retrieval requirement.
Unlike database records, books contain longer documents with chapters, sections, pages, and different forms of legal explanation.
We therefore maintained a separate vector retrieval flow for book/PDF content using Pinecone.
The system can retrieve relevant book content and connect it back to the appropriate source.
The retrieved book chunks also retain their document context, allowing the application to provide a more useful reference instead of presenting the generated answer as if it came directly from the model.
The application is more than a simple:
Question → Vector Search → LLM
pipeline.
The system uses an agentic approach to determine how a query should be handled and which information sources are relevant.
Depending on the query, the system can work across different legal knowledge sources rather than treating the entire database as one large undifferentiated knowledge base.
This makes the retrieval process more controlled and allows the application to handle different types of legal questions more effectively.
The overall idea is:
User Query
↓
Query Understanding / Routing
↓
Relevant Legal Sources
↓
Hybrid Retrieval
↓
Reranking
↓
Context Selection
↓
LLM Generation
↓
Citation + Final Answer
This approach gives the application more control over the retrieval and generation process.
Legal questions are not always written using the exact terminology present in the source data.
A user may describe a legal concept using everyday language while the underlying documents use formal legal terminology.
To address this, query expansion can be used to generate additional search formulations before retrieval.
For example:
User Query → Expanded Search Queries → Retrieval → Reranking
This increases the chances of finding relevant information when the original wording does not directly match the source content.
We also tested query expansion and supervisor-style routing to improve the system’s ability to identify the appropriate information sources.
Retrieval is only one half of a RAG system.
The next challenge is making sure the LLM actually uses the retrieved legal information.
The generation layer was designed around grounded responses.
The model receives the selected context and is instructed to answer based on the available information rather than generating unsupported legal claims.
This is particularly important in legal applications.
An answer that sounds convincing but is not supported by the retrieved source can be more problematic than simply saying that the required information was not found.
The system therefore treats grounding as an important part of answer generation.
For a legal AI application, citations are not an optional feature.
Users need to understand where the information came from.
Where source URLs are available, the application can provide the relevant source reference.
For database-based information, the response can also be connected back to the relevant table and source record.
For book-based retrieval, the system maintains the connection between the generated response and the retrieved book content.
This gives users a way to move from:
AI Answer → Source → Original Legal Information
The objective is simple:
The AI should not just tell the user an answer.
It should also help the user understand where that answer came from.
Legal information changes over time.
A system that only knows the current version of a law may not be sufficient when users need to understand how a provision has changed.
One of the capabilities we incorporated was comparison between older and newer legal information.
This allows the application to distinguish between historical and current information where the underlying data supports it.
This is particularly useful for questions involving:
Version awareness is an important part of making legal RAG useful beyond simple document search.
Building the application was only one part of the work.
The next question was:
“How do we know whether the RAG system is actually working well?”
For this, we designed evaluation at multiple levels.
First, we evaluate the retrieval layer.
The key questions include:
Did we retrieve the relevant information?
Did we miss important information?
Are irrelevant chunks appearing too high in the results?
This is where metrics such as Context Precision and Context Recall become useful.
Context Precision helps us understand whether the retrieved information is relevant.
Context Recall helps us understand whether important information was missed.
Once the right context has been retrieved, we evaluate the generated answer.
Important dimensions include:
Faithfulness
Is the answer supported by the retrieved context?
Answer Relevancy
Does the response actually address the user’s question?
Answer Correctness
Is the final answer factually correct?
Completeness
Did the response cover the important parts of the question?
Style and Tone
Is the response appropriate for the intended users?
This gives us a broader view than simply asking whether the LLM generated a good-looking response.
To evaluate the system consistently, we created a Golden Dataset.
The dataset contains representative legal questions along with the information required to evaluate the expected behavior of the system.
This gives us a fixed benchmark.
Instead of testing the application with random questions every time, we can run the same evaluation set against different versions of the system.
For example:
Version A → Evaluation
Version B → Evaluation
Version C → Evaluation
The results can then be compared across versions.
This becomes the foundation for regression testing.
Manually reviewing every AI response does not scale.
For larger evaluation sets, we can use an LLM-based evaluator to judge responses against predefined criteria.
This approach can evaluate dimensions such as:
A rubric-based evaluation approach allows the evaluator to follow consistent criteria rather than simply giving an overall opinion.
We also use evaluation outputs and judge results to identify common failure patterns and areas where the RAG pipeline needs improvement.
Legal RAG systems continuously change.
We may change:
Every change can improve one part of the system while accidentally affecting another.
For example:
Before a change:
Faithfulness → 92%
Answer Correctness → 89%
Context Recall → 91%
After a change:
Faithfulness → 94%
Answer Correctness → 82%
Context Recall → 93%
The system improved in some areas but became weaker in answer correctness.
Without regression testing, this type of issue can easily go unnoticed.
The Golden Dataset therefore becomes a continuous benchmark for the application.
Build → Evaluate → Compare → Improve → Test Again
We are also moving evaluation closer to the engineering lifecycle.
Instead of treating evaluation as a separate manual activity, RAG evaluation can be integrated into CI/CD.
A simplified workflow looks like:
Code Change
↓
Build & Unit Tests
↓
RAG Evaluation
↓
Compare with Benchmarks
↓
Quality Gates
↓
Deploy or Block
For example, an application can define minimum acceptable thresholds for important metrics.
If the new version meets the required thresholds, deployment can continue.
If an important metric drops below the defined threshold, the deployment can be stopped for investigation.
This makes RAG quality part of the deployment process.
Evaluation should not stop after deployment.
Once the system is being used by real users, we also need to monitor operational and quality signals.
Important monitoring areas include:
Latency
P50 and P95 latency help us understand both typical and slower user experiences.
Token Usage and Cost
Monitoring input tokens, output tokens, total token usage, and cost helps us understand whether the application is operating efficiently.
Error Rate
We monitor failed requests and service-level issues.
Reliability
The application needs to continue working when users depend on it.
Rate Limits
External APIs and services can have request or token limits, so traffic and throttling need to be monitored.
Quality
RAG evaluation metrics can also be tracked over time to identify degradation.
This gives us a complete view of the system:
Quality + Safety + Performance + Cost + Reliability
When an AI response is incorrect, simply knowing that the answer was wrong is not enough.
We need to understand why.
For this reason, the application also considers observability across the RAG pipeline.
A typical trace can help us understand:
User Query
→ Query Processing
→ Retrieved Sources
→ Retrieved Chunks
→ Reranking
→ Final Context
→ LLM Response
→ Citation
→ Evaluation Result
This makes debugging much easier.
For example, if an answer is incorrect, we can investigate whether:
The wrong documents were retrieved.
Important information was missed.
The reranker selected the wrong chunks.
The LLM failed to follow the retrieved context.
The answer was incomplete.
The citation was incorrect.
This changes RAG debugging from guessing to tracing the actual pipeline.

Building a legal RAG application is not simply about connecting a vector database to an LLM.
The difficult part is building a system that can work reliably with complex information.
For a legal and corporate use case, the system needs:
Reliable data preparation
Structured and semantic retrieval
Hybrid search
Reranking
Source-aware responses
Legal version awareness
Grounded generation
Continuous evaluation
Regression testing
Observability
Production monitoring
Each component contributes to the overall quality of the application.
A strong RAG system is therefore not just:
LLM + Vector Database
It is an engineering system built around:
Data → Retrieval → Context → Generation → Evaluation → Monitoring → Continuous Improvement
At a high level, the complete workflow looks like this:
User
↓
Agentic Query Understanding
↓
Query Expansion / Routing
↓
Hybrid Retrieval
Semantic Search + Keyword/BM25
↓
Legislation + Relevant Legal Sources
↓
Reranking
↓
Context Selection
↓
LLM
↓
Grounded Answer
↓
Citations / Source References
↓
Evaluation
Retrieval + Application + Safety + Operational
↓
Regression Testing
↓
Monitoring
↓
Continuous Improvement
This is the difference between building a RAG demo and building a RAG application that is designed for real-world use.
One of the biggest lessons from this project is that RAG quality cannot be measured using a single metric.
A system may retrieve relevant information but generate an incorrect answer.
It may generate a correct answer but miss important information.
It may provide a good answer but take too long to respond.
It may perform well during testing but degrade after deployment.
That is why evaluation needs to happen at multiple levels.
Retrieval Evaluation asks:
“Did we find the right information?”
Application Evaluation asks:
“Did we generate a correct, relevant, complete answer?”
Safety Evaluation asks:
“Is the response safe and does it protect sensitive information?”
Operational Evaluation asks:
“Is the system fast, reliable, and cost-efficient?”
Regression Testing asks:
“Did our latest change improve the system without breaking existing behavior?”
Production Monitoring asks:
“Is the system continuing to perform well over time?”
Together, these create a much stronger evaluation strategy.
Our approach to this legal Agentic RAG application was not to treat the LLM as the complete solution.
The LLM is one component of a larger system.
The real engineering work happens across the entire pipeline:
Preparing the legal data.
Designing the right chunking strategy.
Building hybrid retrieval.
Prioritizing relevant legislation.
Using reranking to improve context selection.
Connecting answers back to their sources.
Handling legal versions.
Evaluating retrieval and generation.
Building a Golden Dataset.
Running regression tests.
Integrating evaluation into CI/CD.
And continuously monitoring the application after deployment.
For legal AI, the goal is not simply to generate an impressive answer.
The goal is to build a system where the answer can be retrieved, grounded, evaluated, traced, and continuously improved.
That is what makes RAG suitable for serious real-world applications.

Recent Comments