Building a Production-Grade Agentic RAG Application for Legal & Corporate Research

  • Home
  • Blog
  • Building a Production-Grade Agentic RAG Application for Legal & Corporate Research

Legal research is an information-heavy process.

For advocates and corporate legal teams, finding the right provision, case law, notification, circular, commentary, or procedure is often only the first step. The information also needs to be relevant to the question, grounded in the correct source, and presented with enough context to support the final answer.

This creates an interesting challenge for Generative AI.

A simple chatbot that sends a question to an LLM is not enough for legal applications.

The system needs reliable retrieval, structured legal data handling, source citations, protection against unsupported answers, and continuous evaluation.

We recently built a production-grade Agentic RAG application for the Indian legal and corporate domain to address this problem.

The application works with structured and semi-structured Indian legal data along with legal books and documents, and is designed to help advocates and corporate teams interact with this information using natural-language queries.

This article explains how we approached the problem, from data preparation and retrieval to RAG evaluation, regression testing, and production monitoring.

1. The Problem We Wanted to Solve

The objective was not simply to build a legal chatbot.

The goal was to build an AI system that could understand a user’s legal or corporate query, identify the relevant sources, retrieve the right information, and generate a grounded response with appropriate references.

A typical interaction looks like:

User Question → Query Understanding → Retrieval → Reranking → Context → LLM → Answer + Citation

For example, a user may ask a question related to:

  • A particular legislation or provision
  • A case law
  • A notification or circular
  • A legal procedure
  • Corporate or regulatory requirements
  • Differences between an older and newer law
  • Information contained inside legal books

The challenge is that the required information may not exist in one place.

A good answer may require information from legislation as well as case laws, notifications, circulars, or books.

This is where a carefully designed RAG architecture becomes important.

2. Understanding the Data

One of the biggest challenges was the nature of the source data.

A large part of the legal information was available in database tables rather than clean documents.

The application works with multiple independent legal data sources, including tables such as:

  • Article
  • CaseLaws
  • Circular
  • Commentary
  • Procedure Details
  • Legislation
  • Notifications

These tables contain text-heavy legal information and are not simply connected through a single relational structure that can be directly passed to an LLM.

Alongside this structured and semi-structured data, we also had legal books and PDF-based content.

This meant that we needed two different retrieval approaches:

Database-based retrieval for structured legal information.

Vector-based document retrieval for books and PDF content.

3. Preparing the Legal Data

Before building the retrieval layer, the source data needed to be prepared carefully.

We migrated the relevant database data from SQL Server into PostgreSQL/Neon and designed the retrieval layer around PostgreSQL with PGVector.

For the book and PDF content, we used a separate vector retrieval approach with Pinecone.

The objective was to preserve important legal metadata while making the content searchable using semantic retrieval.

Data cleaning was also an important part of this stage.

We handled issues such as:

  • HTML cleanup
  • NULL and empty values
  • Long text fields
  • Source URLs
  • Metadata preservation
  • Invalid records
  • Table-specific fields

The goal was not simply to create embeddings.

The goal was to create searchable legal knowledge while preserving the information required later for citations and source verification.

4. Chunking Strategy

Chunking is particularly important in legal RAG.

If chunks are too large, retrieval can bring unnecessary information and increase the context size.

If chunks are too small, important legal context can be separated.

Instead of applying the same chunking strategy to every table, we used a table-aware approach.

We treated database records as parent-level units and created child chunks when individual fields were too large.

For example, long legal fields such as case-law headnotes or detailed legislation headings required further splitting.

Other fields could often remain as a complete record because their structure already represented a meaningful unit of information.

This created a parent-child relationship:

Parent Record
↓
Relevant Field
↓
Child Chunks

Metadata was preserved across the chunks so that retrieved content could still be connected back to its original source.

This is important for legal applications because retrieval should not lose the identity or context of the original record.

For books and PDF documents, the chunking process was designed around the document structure so that retrieved sections could be traced back to the relevant book content.

5. Building the Retrieval Layer

The retrieval layer was designed as a combination of semantic and keyword-based search.

Pure semantic search can understand meaning well, but legal queries often contain exact terms, section names, case references, legislation names, and other terminology where keyword matching is extremely useful.

Therefore, we used a hybrid approach:

Semantic Retrieval + Keyword Retrieval

The retrieval strategy gave more weight to semantic understanding while still using keyword-based matching to improve exact-term retrieval.

BM25 was also incorporated for keyword-based retrieval.

This combination helps address two different types of legal queries:

“Explain the rules related to this topic.”

and

“Find information related to Section X of a specific Act.”

Both require different retrieval signals.

6. Legislation-First Retrieval

Another important part of the retrieval strategy was prioritizing legislation.

For many legal questions, the applicable legislation provides the primary foundation for the answer.

The retrieval process therefore gives priority to legislation-related results and then brings supporting information from other sources.

A simplified strategy looks like:

User Query

↓

Retrieve top legislation results

↓

Retrieve relevant results from other legal tables

↓

Combine results

↓

Rerank

↓

Send the best context to the LLM

The system was designed around a limited context window rather than simply passing a large number of retrieved chunks to the model.

We used a maximum context limit of 32 chunks, with legislation receiving priority and a smaller number of relevant chunks coming from other sources.

This helps control noise and token usage while keeping the most important legal information in context.

7. Reranking the Retrieved Information

Retrieval gives us candidate information.

But candidate information is not necessarily the best information.

This is where reranking becomes useful.

After the initial retrieval stage, the retrieved results are reranked using Cohere reranking.

The idea is simple:

Initial Retrieval → Candidate Results → Reranking → Best Context

The reranker helps identify which retrieved chunks are most relevant to the actual user question.

This becomes especially important when several legal documents contain similar terminology but only a few directly address the question.

8. Working with Legal Books and PDFs

Legal books introduced another retrieval requirement.

Unlike database records, books contain longer documents with chapters, sections, pages, and different forms of legal explanation.

We therefore maintained a separate vector retrieval flow for book/PDF content using Pinecone.

The system can retrieve relevant book content and connect it back to the appropriate source.

The retrieved book chunks also retain their document context, allowing the application to provide a more useful reference instead of presenting the generated answer as if it came directly from the model.

9. Agentic Query Handling

The application is more than a simple:

Question → Vector Search → LLM

pipeline.

The system uses an agentic approach to determine how a query should be handled and which information sources are relevant.

Depending on the query, the system can work across different legal knowledge sources rather than treating the entire database as one large undifferentiated knowledge base.

This makes the retrieval process more controlled and allows the application to handle different types of legal questions more effectively.

The overall idea is:

User Query

↓

Query Understanding / Routing

↓

Relevant Legal Sources

↓

Hybrid Retrieval

↓

Reranking

↓

Context Selection

↓

LLM Generation

↓

Citation + Final Answer

This approach gives the application more control over the retrieval and generation process.

10. Query Expansion

Legal questions are not always written using the exact terminology present in the source data.

A user may describe a legal concept using everyday language while the underlying documents use formal legal terminology.

To address this, query expansion can be used to generate additional search formulations before retrieval.

For example:

User Query → Expanded Search Queries → Retrieval → Reranking

This increases the chances of finding relevant information when the original wording does not directly match the source content.

We also tested query expansion and supervisor-style routing to improve the system’s ability to identify the appropriate information sources.

11. Grounded Answer Generation

Retrieval is only one half of a RAG system.

The next challenge is making sure the LLM actually uses the retrieved legal information.

The generation layer was designed around grounded responses.

The model receives the selected context and is instructed to answer based on the available information rather than generating unsupported legal claims.

This is particularly important in legal applications.

An answer that sounds convincing but is not supported by the retrieved source can be more problematic than simply saying that the required information was not found.

The system therefore treats grounding as an important part of answer generation.

12. Citations and Source Traceability

For a legal AI application, citations are not an optional feature.

Users need to understand where the information came from.

Where source URLs are available, the application can provide the relevant source reference.

For database-based information, the response can also be connected back to the relevant table and source record.

For book-based retrieval, the system maintains the connection between the generated response and the retrieved book content.

This gives users a way to move from:

AI Answer → Source → Original Legal Information

The objective is simple:

The AI should not just tell the user an answer.

It should also help the user understand where that answer came from.

13. Old vs. New Law Comparison

Legal information changes over time.

A system that only knows the current version of a law may not be sufficient when users need to understand how a provision has changed.

One of the capabilities we incorporated was comparison between older and newer legal information.

This allows the application to distinguish between historical and current information where the underlying data supports it.

This is particularly useful for questions involving:

  • Changes in legislation
  • Replaced provisions
  • Earlier versions of laws
  • Current applicability
  • Historical legal references

Version awareness is an important part of making legal RAG useful beyond simple document search.

14. RAG Evaluation

Building the application was only one part of the work.

The next question was:

“How do we know whether the RAG system is actually working well?”

For this, we designed evaluation at multiple levels.

Retrieval Evaluation

First, we evaluate the retrieval layer.

The key questions include:

Did we retrieve the relevant information?

Did we miss important information?

Are irrelevant chunks appearing too high in the results?

This is where metrics such as Context Precision and Context Recall become useful.

Context Precision helps us understand whether the retrieved information is relevant.

Context Recall helps us understand whether important information was missed.

Generation Evaluation

Once the right context has been retrieved, we evaluate the generated answer.

Important dimensions include:

Faithfulness
Is the answer supported by the retrieved context?

Answer Relevancy
Does the response actually address the user’s question?

Answer Correctness
Is the final answer factually correct?

Completeness
Did the response cover the important parts of the question?

Style and Tone
Is the response appropriate for the intended users?

This gives us a broader view than simply asking whether the LLM generated a good-looking response.

15. Building a Golden Dataset

To evaluate the system consistently, we created a Golden Dataset.

The dataset contains representative legal questions along with the information required to evaluate the expected behavior of the system.

This gives us a fixed benchmark.

Instead of testing the application with random questions every time, we can run the same evaluation set against different versions of the system.

For example:

Version A → Evaluation

Version B → Evaluation

Version C → Evaluation

The results can then be compared across versions.

This becomes the foundation for regression testing.

16. LLM-as-a-Judge and G-Eval

Manually reviewing every AI response does not scale.

For larger evaluation sets, we can use an LLM-based evaluator to judge responses against predefined criteria.

This approach can evaluate dimensions such as:

  • Correctness
  • Relevance
  • Faithfulness
  • Completeness
  • Style
  • Other application-specific criteria

A rubric-based evaluation approach allows the evaluator to follow consistent criteria rather than simply giving an overall opinion.

We also use evaluation outputs and judge results to identify common failure patterns and areas where the RAG pipeline needs improvement.

17. Regression Testing

Legal RAG systems continuously change.

We may change:

  • Embedding models
  • Retrieval strategies
  • Chunking
  • Prompts
  • LLMs
  • Rerankers
  • Query expansion
  • Source data

Every change can improve one part of the system while accidentally affecting another.

For example:

Before a change:

Faithfulness → 92%
Answer Correctness → 89%
Context Recall → 91%

After a change:

Faithfulness → 94%
Answer Correctness → 82%
Context Recall → 93%

The system improved in some areas but became weaker in answer correctness.

Without regression testing, this type of issue can easily go unnoticed.

The Golden Dataset therefore becomes a continuous benchmark for the application.

Build → Evaluate → Compare → Improve → Test Again

18. Evaluation in CI/CD

We are also moving evaluation closer to the engineering lifecycle.

Instead of treating evaluation as a separate manual activity, RAG evaluation can be integrated into CI/CD.

A simplified workflow looks like:

Code Change

↓

Build & Unit Tests

↓

RAG Evaluation

↓

Compare with Benchmarks

↓

Quality Gates

↓

Deploy or Block

For example, an application can define minimum acceptable thresholds for important metrics.

If the new version meets the required thresholds, deployment can continue.

If an important metric drops below the defined threshold, the deployment can be stopped for investigation.

This makes RAG quality part of the deployment process.

19. Production Monitoring

Evaluation should not stop after deployment.

Once the system is being used by real users, we also need to monitor operational and quality signals.

Important monitoring areas include:

Latency

P50 and P95 latency help us understand both typical and slower user experiences.

Token Usage and Cost

Monitoring input tokens, output tokens, total token usage, and cost helps us understand whether the application is operating efficiently.

Error Rate

We monitor failed requests and service-level issues.

Reliability

The application needs to continue working when users depend on it.

Rate Limits

External APIs and services can have request or token limits, so traffic and throttling need to be monitored.

Quality

RAG evaluation metrics can also be tracked over time to identify degradation.

This gives us a complete view of the system:

Quality + Safety + Performance + Cost + Reliability

20. Observability and Debugging

When an AI response is incorrect, simply knowing that the answer was wrong is not enough.

We need to understand why.

For this reason, the application also considers observability across the RAG pipeline.

A typical trace can help us understand:

User Query

→ Query Processing

→ Retrieved Sources

→ Retrieved Chunks

→ Reranking

→ Final Context

→ LLM Response

→ Citation

→ Evaluation Result

This makes debugging much easier.

For example, if an answer is incorrect, we can investigate whether:

The wrong documents were retrieved.

Important information was missed.

The reranker selected the wrong chunks.

The LLM failed to follow the retrieved context.

The answer was incomplete.

The citation was incorrect.

This changes RAG debugging from guessing to tracing the actual pipeline.

21. Why This Approach Matters

Building a legal RAG application is not simply about connecting a vector database to an LLM.

The difficult part is building a system that can work reliably with complex information.

For a legal and corporate use case, the system needs:

Reliable data preparation

Structured and semantic retrieval

Hybrid search

Reranking

Source-aware responses

Legal version awareness

Grounded generation

Continuous evaluation

Regression testing

Observability

Production monitoring

Each component contributes to the overall quality of the application.

A strong RAG system is therefore not just:

LLM + Vector Database

It is an engineering system built around:

Data → Retrieval → Context → Generation → Evaluation → Monitoring → Continuous Improvement

22. The Overall Architecture

At a high level, the complete workflow looks like this:

User

↓

Agentic Query Understanding

↓

Query Expansion / Routing

↓

Hybrid Retrieval

Semantic Search + Keyword/BM25

↓

Legislation + Relevant Legal Sources

↓

Reranking

↓

Context Selection

↓

LLM

↓

Grounded Answer

↓

Citations / Source References

↓

Evaluation

Retrieval + Application + Safety + Operational

↓

Regression Testing

↓

Monitoring

↓

Continuous Improvement

This is the difference between building a RAG demo and building a RAG application that is designed for real-world use.

23. What We Learned

One of the biggest lessons from this project is that RAG quality cannot be measured using a single metric.

A system may retrieve relevant information but generate an incorrect answer.

It may generate a correct answer but miss important information.

It may provide a good answer but take too long to respond.

It may perform well during testing but degrade after deployment.

That is why evaluation needs to happen at multiple levels.

Retrieval Evaluation asks:

“Did we find the right information?”

Application Evaluation asks:

“Did we generate a correct, relevant, complete answer?”

Safety Evaluation asks:

“Is the response safe and does it protect sensitive information?”

Operational Evaluation asks:

“Is the system fast, reliable, and cost-efficient?”

Regression Testing asks:

“Did our latest change improve the system without breaking existing behavior?”

Production Monitoring asks:

“Is the system continuing to perform well over time?”

Together, these create a much stronger evaluation strategy.

Conclusion

Our approach to this legal Agentic RAG application was not to treat the LLM as the complete solution.

The LLM is one component of a larger system.

The real engineering work happens across the entire pipeline:

Preparing the legal data.

Designing the right chunking strategy.

Building hybrid retrieval.

Prioritizing relevant legislation.

Using reranking to improve context selection.

Connecting answers back to their sources.

Handling legal versions.

Evaluating retrieval and generation.

Building a Golden Dataset.

Running regression tests.

Integrating evaluation into CI/CD.

And continuously monitoring the application after deployment.

For legal AI, the goal is not simply to generate an impressive answer.

The goal is to build a system where the answer can be retrieved, grounded, evaluated, traced, and continuously improved.

That is what makes RAG suitable for serious real-world applications.