RAG Knowledge Base assessment and evaluation with RAGAS framework
Standard RAG that is based on AWS BeedRock consists of the following parts: pdf corpus of data strored in S3, connected Beedrock Embedding Model producer with default chunk strategy setup with fixed size limitation and overlap. After RAG is connected to live system (chat bot, API, content engine, etc.) and operates we are comming to next challage - to measure and tune the accuracy of RAG.
Evaluation Infrastructure
When it comes to evaluation and tunning, it can be made using AWS RAG evaluation and 3rd-party toolboxes, In this post I’ll give an overview of ragas framework that allows to define a set of metrics that will be evaluated across different configurations of RAG, output results into xls file, have historical data of metrics.
It is very useful both for RAG evaluation and tuning. Changing knowledge base processing, like chunk size, overlap, embedding model, etc. we can reevaluate new assembled version using ragas framework and compare with previous versions.
Evaluation Metrics
Here is a list of RAGAS evaluations metrics:
- Faithfulness: This measures the factual consistency of the generated answer against the given context. It is calculated from answer and retrieved context. The answer is scaled to (0,1) range. Higher the better.
- Answer Relevance: This metric focuses on assessing how pertinent the generated answer is to the given prompt. A lower score is assigned to answers that are incomplete or contain redundant information and higher scores indicate better relevancy. This metric is computed using the question, the context and the answer. Please note, that even though in practice the score will range between 0 and 1 most of the time, this is not mathematically guaranteed, due to the nature of the cosine similarity ranging from -1 to 1.
- Context Precision: This is a metric that evaluates whether all of the ground-truth relevant items present in the contexts are ranked higher or not. Ideally all the relevant chunks must appear at the top ranks. This metric is computed using the question, ground_truth and the contexts, with values ranging between 0 and 1, where higher scores indicate better precision.
- Context Recall: This metric measures the extent to which the retrieved context aligns with the annotated answer, treated as the ground truth. It is computed based on the ground truth and the retrieved context, and the values range between 0 and 1, with higher values indicating better performance.
- Context entities recall: This metric gives the measure of recall of the retrieved context, based on the number of entities present in both ground_truths and contexts relative to the number of entities present in the ground_truths alone. Simply put, it is a measure of what fraction of entities are recalled from ground_truths. This metric is useful in fact-based use cases like tourism help desk, historical QA, etc. This metric can help evaluate the retrieval mechanism for entities, based on comparison with entities present in ground_truths, because in cases where entities matter, we need the contexts which cover them.
- Answer Semantic Similarity: The concept of Answer Semantic Similarity pertains to the assessment of the semantic resemblance between the generated answer and the ground truth. This evaluation is based on the ground truth and the answer, with values falling within the range of 0 to 1. A higher score signifies a better alignment between the generated answer and the ground truth.
- Answer Correctness: The assessment of Answer Correctness involves gauging the accuracy of the generated answer when compared to the ground truth. This evaluation relies on the ground truth and the answer, with scores ranging from 0 to 1. A higher score indicates a closer alignment between the generated answer and the ground truth, signifying better correctness. Answer correctness encompasses two critical aspects: semantic similarity between the generated answer and the ground truth, as well as factual similarity. These aspects are combined using a weighted scheme to formulate the answer correctness score. Users also have the option to employ a ‘threshold’ value to round the resulting score to binary, if desired.
- Aspect Critique: This is designed to assess submissions based on predefined aspects such as harmlessness and correctness. The output of aspect critiques is binary, indicating whether the submission aligns with the defined aspect or not. This evaluation is performed using the ‘answer’ as input.
Also we can define our own custom metrics it a numeric or PROMT form for ragas.
Evaluation setup
Specify or detect Knowledge Base ID
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
import botocore
import boto3
session = boto3.Session()
bedrock_client = session.client('bedrock-agent')
try:
response = bedrock_client.list_knowledge_bases(
maxResults=1 # We only need to retrieve the first Knowledge Base
)
knowledge_base_summaries = response.get('knowledgeBaseSummaries', [])
if knowledge_base_summaries:
kb_id = knowledge_base_summaries[0]['knowledgeBaseId']
print(f"Knowledge Base ID: {kb_id}")
else:
print("No Knowledge Base summaries found.")
except botocore.exceptions.ClientError as e:
print(f"Error: {e}")
Initialize BedRock Models:
We’ll be using several different models, each with its own purpose:
- Initialize bedrock model amazon.nova-pro-v1:0 as your large language model to perform query completions using the RAG pattern.
- Initialize bedrock model us.amazon.nova-2-lite-v1:0 as your large language model to perform RAG evaluation.
- Initialize bedrock model amazon.titan-embed-text-v1 as your large language embedding model to create embeddings for RAG evaluation. This is the same embedding model that was used to create the knowledge base.
- Initialize LangChain retriever integrated with knowledge bases.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
import boto3
import pprint
from langchain_aws import ChatBedrock
from langchain_aws import BedrockEmbeddings
from langchain_aws.retrievers.bedrock import AmazonKnowledgeBasesRetriever
pp = pprint.PrettyPrinter(indent=2)
bedrock_client = boto3.client('bedrock-runtime')
llm_for_text_generation = ChatBedrock(model_id="amazon.nova-pro-v1:0", client=bedrock_client)
llm_for_evaluation = ChatBedrock(model_id="us.amazon.nova-2-lite-v1:0", client=bedrock_client)
bedrock_embeddings = BedrockEmbeddings(model_id="amazon.titan-embed-text-v1", client=bedrock_client)
AmazonKnowledgeBasesRetriever allows get advanced information about documents metadata:
1
2
3
4
5
6
7
retriever = AmazonKnowledgeBasesRetriever(
knowledge_base_id=kb_id,
retrieval_config={"vectorSearchConfiguration": {"numberOfResults": 5}},
# endpoint_url=endpoint_url,
# region_name="us-east-1",
# credentials_profile_name="<profile_name>",
)
Creating PROMT with context and question as variables:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
from langchain.prompts import PromptTemplate
PROMPT_TEMPLATE = """
Human: You are a financial advisor AI system, and provides answers to questions by using fact based and statistical information when possible.
Use the following pieces of information to provide a concise answer to the question enclosed in <question> tags.
If you don't know the answer, just say that you don't know, don't try to make up an answer.
<context>
{context}
</context>
<question>
{question}
</question>
The response should be specific and use statistics or numbers when possible.
Assistant:"""
prompt = PromptTemplate(template=PROMPT_TEMPLATE,
input_variables=["context", "question"])
used amazon.nova-pro-v1:0 model
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
def format_docs(docs): # concatenate the text from the page_content field in the output from retriever.invoke
return "\n\n".join(doc.page_content for doc in docs)
chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt
| llm_for_text_generation
| StrOutputParser()
)
query = "Provide a list of ten risks for AnyCompany financial as a numbered list. Do not include descriptions."
response = chain.invoke(query)
print(response)
PROMT execution results v1:
- Supplier Risk
- Market Risks (Commodity Prices)
- Market Risks (Foreign Exchange Rates)
- Market Risks (Equity Prices)
- Operational Risk
- Regulatory Risk
- Strategic Risk
- Reputation Risk
- Legal Risk
- Environmental, Social, and Governance (ESG) Risk
PROMT execution results v2 (once AWS released new model version - same day):
- Market risk due to changes in market prices.
- Credit risk from issuer defaults.
- Liquidity risk from inability to sell securities timely or at fair value.
- Concentration risk with top five borrowers accounting for 15% of total loan portfolio.
- Estimation risk in allowance for doubtful accounts.
- Trading securities portfolio unrealized loss of $50 million as of December 31, 2021.
- Significant investment positions in Company A ($400 million), Company B ($350 million), and Company C ($250 million), representing 22% of the total investment portfolio.
- Risk from leasing office space from a significant shareholder (PersonA).
- Risk associated with the use of estimates in financial reporting.
- Risk from material related party transactions conducted at arm’s length.
Preparing GroundTruth Data:
To perform evaluation, we need to have a ground truth data. The best way is to prepare such corpus of data based on your documents domain knowledge with domain experts and business analysts. Another technique can be usage of highly tunned model (that is perfectly trained under your domain), or high-cost comercial model (just for corpus generation task) not for reglar usage.
So once question and ground_truths pairs are prepared, they are ready for inference:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
from datasets import Dataset
import time
import random
# Define questions and ground truths for RAGAS evaluation
questions = [
"What was the primary reason for the increase in net cash by operating activities for AnyCompany Financial in 2021?",
"Which year did AnyCompany Financial have the highest net cash used in investing activities, what was the primary reason?",
"What was the primary source of cash inflows from financing activities for AnyCompany Financial in 2021?",
"Calculate the year-over-year percentage change in cash and cash equivalents for AnyCompany Financial from 2020 to 2021.",
"With the information provided, what can you infer about AnyCompany Financial's overall financial health and growth prospects?"
]
ground_truth = [
"An increase in net cash provided by operating activities was primarily due to increases in net income and favorable changes in operating assets and liabilities.",
"AnyCompany Financial had the highest net cash used in investing activities in 2021, at $360 million, compared to $290 million in 2020 and $240 million in 2019. The primary reason was an increase in purchases of property, plant, and equipment and marketable securities.",
"The primary source of cash inflows from financing activities for AnyCompany Financial in 2021 was an increase in proceeds from the issuance of common stock and long-term debt.",
"To calculate the year-over-year percentage change in cash and cash equivalents from 2020 to 2021: \
2020 cash and cash equivalents: $350 million \
2021 cash and cash equivalents: $480 million \
Percentage change = (2021 value - 2020 value) / 2020 value * 100 \
= ($480 million - $350 million) / $350 million * 100 \
= 37.14% increase",
"Based on information provided, AnyCompany Financial appears to be in a healthy financial position with good growth prospects. The company increased its net cash provided by operating activities, indicating strong profitability and efficient management of working capital. AnyCompany Financial has been investing in long-term assets, such as property, plant, and equipment, and marketable securities, which suggests plans for future growth and expansion. The company was able to finance its growth through the issuance of common stock and long-term debt, indicating confidence from investors and lenders. Overall, AnyCompany Financial's steady increase in cash and cash equivalents over the past three years provides a strong foundation for future growth and investment opportunities."
]
def get_model_response(query, chain, retriever, max_retries=5, wait_time=15):
"""Get response from the model with fixed wait time between retries"""
for attempt in range(max_retries):
try:
# Configure Nova Lite with increased tokens
nova_config = {
"schemaVersion": "messages-v1",
"messages": [{
"role": "user",
"content": [{"text": query}]
}],
"inferenceConfig": {
"maxTokens": 2048,
"temperature": 0.5,
"topP": 0.9,
"topK": 20
}
}
# Try to invoke with config override
try:
answer = chain.invoke(
query,
config_override={"model_kwargs": nova_config}
)
except AttributeError:
# If config_override doesn't work, try direct invocation
answer = chain.invoke(query)
context = [docs.page_content for docs in retriever.invoke(query)]
print(f"Successfully processed query on attempt {attempt + 1}")
return answer, context
except Exception as e:
if attempt == max_retries - 1:
print(f"Failed after {max_retries} attempts for query: {query[:50]}...")
print(f"Error: {str(e)}")
return None, None
print(f"Attempt {attempt + 1} failed, waiting {wait_time} seconds before retry...")
time.sleep(wait_time)
# Process questions one at a time with fixed delay
answers = []
contexts = []
print("Starting to process questions...")
for i, query in enumerate(questions, 1):
print(f"\nProcessing question {i}/{len(questions)}")
print(f"Query: {query[:100]}...")
answer, context = get_model_response(query, chain, retriever)
if answer is not None:
answers.append(answer)
contexts.append(context)
print(f"Successfully processed question {i}")
else:
print(f"Failed to process question {i}")
if i < len(questions):
print(f"Waiting 60 seconds before next question...")
time.sleep(60)
# Create dataset for RAGAS evaluation
data = {
"question": questions[:len(answers)],
"ground_truth": ground_truth[:len(answers)],
"answer": answers,
"contexts": contexts
}
# Convert to dataset
dataset = Dataset.from_dict(data)
# Print dataset information
print("\nDataset Creation Summary:")
print(f"Total questions processed: {len(dataset)} out of {len(questions)}")
print(f"Columns available: {dataset.column_names}")
# Print sample entry
if len(dataset) > 0:
print("\nSample Entry (First Question):")
print(f"Question: {dataset[0]['question']}")
print(f"Ground Truth: {dataset[0]['ground_truth']}")
print(f"Model Answer: {dataset[0]['answer']}")
else:
print("\nNo entries were successfully processed into the dataset.")
Retriever actions:
The Retriever is connected to Knowledge Base - it populates context with 5 most closed document chunks. These chunks will be included into final results evaluation table for expert review of measurement, metrics, Q&A and sampling.
Compare answer results from LLM & Ground truth:
1
2
3
4
5
6
7
8
i = 0
for answer in answers:
i = i + 1
print(str(i) + ').' + questions[i - 1] + '\n')
print("LLM:" + answer + '\n')
print("Ground truth: " + ground_truth[i - 1] + '\n')
dataset.to_csv('DATASET.csv')
Answers by Model and GroundTruth to Questions:
1
2
3
1).What was the primary reason for the increase in net cash by operating activities for AnyCompany Financial in 2021?
LLM:The primary reason for the increase in net cash provided by operating activities for AnyCompany Financial in 2021 was an increase in net income and favorable changes in operating assets and liabilities. Specifically, net cash provided by operating activities rose from $880 million in 2020 to $1,100 million in 2021.
Ground truth: An increase in net cash provided by operating activities was primarily due to increases in net income and favorable changes in operating assets and liabilities.
1
2
3
2).Which year did AnyCompany Financial have the highest net cash used in investing activities, what was the primary reason?
LLM:AnyCompany Financial had the highest net cash used in investing activities in 2021, with a net cash used of $360 million. This was primarily due to an increase in purchases of property, plant, and equipment and marketable securities.
Ground truth: AnyCompany Financial had the highest net cash used in investing activities in 2021, at $360 million, compared to $290 million in 2020 and $240 million in 2019. The primary reason was an increase in purchases of property, plant, and equipment and marketable securities.
Running the evaluation with RAGAs
In the next task, we import all the metrics we want to use from ragas.metrics. Then, using the evaluate () function and simply passing in the relevant metrics and the prepared dataset.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
import warnings
import logging
warnings.filterwarnings('ignore') # ignore warnings related to pydantic v1 to v2 migration
logging.getLogger('root').setLevel(logging.CRITICAL)
from ragas import evaluate
from ragas.metrics import (
faithfulness,
answer_relevancy,
context_recall,
context_precision,
context_entity_recall,
answer_similarity,
answer_correctness
)
from ragas.metrics.critique import (
harmfulness,
maliciousness,
coherence,
correctness,
conciseness
)
# specify the metrics here
metrics = [
faithfulness,
answer_relevancy,
context_precision,
context_recall,
context_entity_recall,
answer_similarity,
answer_correctness,
harmfulness,
maliciousness,
coherence,
correctness,
conciseness
]
try:
result = evaluate(
dataset=dataset,
metrics=metrics,
llm=llm_for_evaluation,
embeddings=bedrock_embeddings,
)
df = result.to_pandas()
except Exception as e:
# Handle any exceptions that occur during the evaluation
print(f"An error occurred: {e}")
Saving RAGAS scores to xls
1
2
3
4
5
import pandas as pd
pd.options.display.max_colwidth = 10
df.style.set_sticky(axis="columns")
df.style.to_excel('styled.xlsx', engine='openpyxl')
Scores analysis and further RAG tuning
Here is one single entry question example from evaluation results (full xls you can dowload and check at Links section).
Question: Which year did AnyCompany Financial have the highest net cash used in investing activities, what was the primary reason?
Ground Truth: AnyCompany Financial had the highest net cash used in investing activities in 2021, at $360 million, compared to $290 million in 2020 and $240 million in 2019. The primary reason was an increase in purchases of property, plant, and equipment and marketable securities.
Answer: AnyCompany Financial had the highest net cash used in investing activities in 2021, with a net cash used of $360 million. The primary reason for this increase was due to an increase in purchases of property, plant, and equipment and marketable securities, compared to $ 290 million in 2020 and $240 million in 2019.
| faithfulness | answer_relevancy | context_precision | context_recall | context_entity_recall | answer_similarity | answer_correctness | harmfulness | maliciousness | coherence | correctness | conciseness |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.666667 | 0.881489 | 0.500000 | 1.000000 | 1.000000 | 0.988129 | 0.847032 | 0 | 0 | 1 | 1 | 1 |
Conclusions
Please note that the scores above give a relative idea on the performance of your RAG application and should be used with caution and not as standalone scores.
Also note that we have used only 5 question/answer pairs for evaluation. As a best practice, you should use enough data to cover different aspects of your document for evaluating model.
Based on the scores, we can review other components of RAG workflow to further optimize the scores.
Few recommended options are to review your chunking strategy, prompt instructions, adding more numberOfResults for additional context and so on.


