Project Overview
This case study focuses on building a multilingual Retrieval-Augmented Generation (RAG) system using two advanced open-source language models: meta-llama/Meta-Llama-3.1-8B-Instruct and mistralai/Mistral-7B-Instruct-v0.3. The system aims to answer questions in both Hindi and English by retrieving relevant information and generating accurate responses in the appropriate language. A key challenge is ensuring the system’s performance in both languages, particularly when the original data is available only in English. To evaluate the system, we focus on two main aspects: the accuracy of the generated responses and the effectiveness of language detection.
Objective
The primary objective is to create a robust RAG system that can handle bilingual queries, retrieve pertinent information from a corpus embedded in English, and generate answers in the language of the query. This system aims to demonstrate the versatility of the тАЬmeta-llama/Meta-Llama-3.1-8B-InstructтАЭ and тАЬmistralai/Mistral-7B-Instruct-v0.3тАЭ models in a multilingual context.
Language Models:
1. Meta-LLaMA (Meta-Llama-3.1-8B-Instruct):
Meta-Llama-3-8B-Instruct is an 8 billion parameter language model developed by Meta AI. It’s based on the Transformer architecture and is likely an evolution of the Llama 2 series. This model is specifically designed for instruction-following tasks and is trained on a vast corpus of internet text, potentially including code and multilingual data. It’s known for its efficient performance and strong capabilities across various NLP tasks, particularly in areas like code generation and multilingual processing. However, its relatively smaller size compared to larger language models may limit its performance on extremely complex tasks.
2. Mistral (Mistral-7B-Instruct-v0.3):
Mistral-7B-Instruct-v0.3 is a 7 billion parameter language model created by Mistral AI. Built on the Transformer architecture with potential optimizations, this model is trained on a diverse dataset and is optimized for instruction-following tasks. Despite its compact size, Mistral-7B is known for its highly efficient performance and strong capabilities across a wide range of NLP tasks. It offers a good balance between model size and performance, making it suitable for various applications. Like the Llama model, its smaller size may pose limitations for very complex tasks, but its efficiency and strong instruction-following abilities make it a powerful tool for many NLP applications.
Document Embedding:
Document Embedding consists of two key components: the vector database and the embedding model.
We can embed documents using the embedding model and store them in a vector database for future reference.
The section below explains this in more detail.
1. Embedding Model - LaBSE (Language-Agnostic BERT Sentence Embedding):
LaBSE, or Language-agnostic BERT Sentence Embedding, is a multilingual sentence embedding.
It’s designed to create semantically meaningful vector representations of sentences across 109 languages. It’s designed to create semantically meaningful vector representations of sentences across 109 languages.
2. Vector Database - Chroma:
Chroma is an open-source embedding database designed for building AI applications with embeddings and natural language processing. It provides a simple, fast, and scalable solution for storing and querying high-dimensional vector data. Chroma is particularly well-suited for applications like semantic search, recommendation systems, and similarity matching.
Language Detection and Instruction Prompting:
The system includes an instruction in the prompt that guides the model to detect the query’s language and respond accordingly. This instruction enables the automatic conversion of retrieved information into the language of the query, whether Hindi or English.
Question Answering:
Questions were posed in both Hindi and English to assess the models’ capabilities. The system was designed to automatically detect the language of the query and use the corresponding model to generate the response.
Documents Used:
The financial reports used in this study were sourced from publicly available open-source files. These files are accessible to anyone interested in exploring or replicating the study. The documents can be downloaded from the following links:
https://s2.q4cdn.com/470004039/files/doc_financials/2022/q4/_10-K-2022-(As-Filed).pdf
https://s2.q4cdn.com/470004039/files/doc_earnings/2023/q4/filing/_10-K-Q4-2023-As-Filed.pdf
Output:
Here are the results for the questions asked in English Language:
Questions | Actual Answer | Meta-Llama-3-8B-Instruct Response | Mistral-7B-Instruct-v0.3 Response |
|---|---|---|---|
What is the total operating income for 2022? | The total operating income for 2022 is $119,437 million. | The total operating income for 2022 is $119,437 million. | The total operating income for 2022 is $119,437 (in millions). |
How does the company measure share-based compensation? | The Company generally measures share-based compensation based on the closing price of the CompanyтАЩs common stock on the date of grant, and recognizes expense on a straight-line basis for its estimate of equity awards that will ultimately vest. | The company recognizes share-based compensation expense on a straight-line basis for its estimate of equity awards that will ultimately vest. | The company measures share-based compensation on a straight-line basis for its estimate of equity awards that will ultimately vest. |
Can you provide information about Restricted Marketable Securities? | The Company considers marketable securities to be restricted when withdrawal or general use is legally restricted. | The Company considers marketable securities to be restricted when withdrawal or general use is legally restricted. | The Company considers marketable securities to be restricted when withdrawal or general use is legally restricted. |
Can you provide information about the 2022 Employee Stock Plan? | The 2022 Employee Stock Plan provides for broad-based equity grants to employees, including executive officers, and permits the granting of restricted stock units (тАЬRSUsтАЭ), stock grants, performance-based awards, stock options and stock appreciation rights. | The 2022 Employee Stock Plan provides for broad-based equity grants to employees, including executive officers, and permits the granting of restricted stock units, stock grants, performance-based awards, stock options, and stock appreciation rights. | The 2022 Employee Stock Plan is a shareholder-approved plan that provides for broad-based equity grants to employees, including executive officers. It permits the granting of Restricted Stock Units (RSUs), stock grants, performance-based awards, stock options, and stock appreciation rights. RSUs generally vest over four years. |
What is the total amount of income taxes paid in 2023, 2022, and 2021? | The total amount of income taxes paid in 2023, 2022, and 2021 is $16,741 million, $19,300 million, and $14,527 million, respectively. | The total amount of income taxes paid in 2023, 2022, and 2021 is $16,741 million, $19,300 million, and $14,527 million, respectively. | The total amount of income taxes paid in 2023 was $16,741 million, in 2022 was $19,300 million, and in 2021 was $14,527 million. |
Here are the results for the questions asked in Hindi Language:
Questions | Actual Answer | Meta-Llama-3-8B-Instruct Response | Mistral-7B-Instruct-v0.3 |
|---|---|---|---|
2022 рдХреЗ рд▓рд┐рдП рдХреБрд▓ рдкрд░рд┐рдЪрд╛рд▓рди рдЖрдп рдХреНрдпрд╛ рдереА? | 2022 рдХреЗ рд▓рд┐рдП рдХреБрд▓ рдкрд░рд┐рдЪрд╛рд▓рди рдЖрдп $119,437 рдорд┐рд▓рд┐рдпрди рд╣реИред | рдХреБрд▓ рдкрд░рд┐рдЪрд╛рд▓рди рдЖрдп 2022 рдореЗрдВ $119,437 рдорд┐рд▓рд┐рдпрди рдереАред | 2022 рдХреЗ рд▓рд┐рдП рдХреБрд▓ рдкрд░рд┐рдЪрд╛рд▓рди рдЖрдп $119,103 рдерд╛ред |
рдХрдВрдкрдиреА рд╢реЗрдпрд░-рдЖрдзрд╛рд░рд┐рдд рдореБрдЖрд╡рдЬреЗ рдХреА рдорд╛рдк рдХреИрд╕реЗ рдХрд░рддреА рд╣реИ? | рдХрдВрдкрдиреА рдЖрдо рддреМрд░ рдкрд░ рдХрдВрдкрдиреА рдХреЗ рд╕рд╛рдорд╛рдиреНрдп рд╕реНрдЯреЙрдХ рдХреЗ рд╕рдорд╛рдкрди рдореВрд▓реНрдп рдХреЗ рдЖрдзрд╛рд░ рдкрд░ рд╢реЗрдпрд░-рдЖрдзрд╛рд░рд┐рдд рдореБрдЖрд╡рдЬреЗ рдХреЛ рдорд╛рдкрддреА рд╣реИ рдЕрдиреБрджрд╛рди рдХреА рддрд╛рд░реАрдЦ, рдФрд░ рдЗрдХреНрд╡рд┐рдЯреА рдкреБрд░рд╕реНрдХрд╛рд░реЛрдВ рдХреЗ рдЕрдиреБрдорд╛рди рдХреЗ рд▓рд┐рдП рд╕реАрдзреА рд░реЗрдЦрд╛ рдХреЗ рдЖрдзрд╛рд░ рдкрд░ рд╡реНрдпрдп рдХреЛ рдкрд╣рдЪрд╛рдирддреА рд╣реИ рдЬреЛ рдЕрдВрддрддрдГ рдирд┐рд╣рд┐рдд рд╣реЛрдВрдЧреЗред | рдХрдВрдкрдиреА рд╢реЗрдпрд░-рдЖрдзрд╛рд░рд┐рдд рдореБрдЖрд╡рдЬреЗ рдХреА рдорд╛рдк рдХреЛ рд╕traight-line рдо╨╡╤В╨╛╨┤ рд╕реЗ рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИ, рдЬрд┐рд╕рдореЗрдВ рдХрдВрдкрдиреА рдЕрдкрдиреЗ рдкреВрд░реЗ рд╕рдВрднрд╛рд╡рд┐рдд рд╢реЗрдпрд░реЛрдВ рдХреЗ рд▓рд┐рдП рдореБрдЖрд╡рдЬреЗ рдХреА рдЕрдиреБрдорд╛рдирд┐рдд рд░рд╛рд╢рд┐ рдХреЛ рд╕рдорд╛рди рд░реВрдк рд╕реЗ рд╡рд┐рддрд░рд┐рдд рдХрд░рддреА рд╣реИред | рдХрдВрдкрдиреА рд╢реЗрдпрд░-рдЖрдзрд╛рд░рд┐рдд рдореБрдЖрд╡рдЬреЗ рдХреЛ рд╕рд╛рдорд╛рдиреНрдпрддрд╛ рдХреЗ рдЕрдиреБрд╕рд╛рд░ рд╢реБрджреНрдз рд╕реНрдЯреНрд░реИрдЯ рд▓рд╛рдЗрди рдореЗрдВ рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИ рдФрд░ рдЙрд╕рдХреЗ рд╕рд╛рде рдЕрдЧрд▓реЗ рд╕реЗ рд╡рд┐рдХрд▓реНрдк рд╕реЗ рд╡рд┐рд╢реЗрд╖ рдкреНрд░рдХрд╛рд░ рдХреЗ рд╢реЗрдпрд░ рдХреЛ рд╕реНрдерд╛рдирд╛рдВрддрд░рд┐рдд рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИ рдЬрд┐рд╕рдХреЗ рдЕрдВрддрд░реНрдЧрдд рдЕрдм рд╡рд┐рдХрд▓реНрдк рд╣реЛ рд╕рдХрддрд╛ рд╣реИ рдХрд┐ рдЙрди рд╢реЗрдпрд░реЛрдВ рдХреЛ рдЕрдкрдиреА рдШреЛрд╖рд┐рдд рдХрд░ рджреА рдЧрдИ рд╕рдордп рд╕реЗ рд╡реЗ рдлрд┐рд░ рд╕реЗ рд╡рд┐рдХрд▓реНрдк рд╣реЛ рд╕рдХрддреЗ рд╣реИрдВ рдпрд╛ рдЙрдиреНрд╣реЗрдВ рдЕрдкрдиреА рдШреЛрд╖рд┐рдд рдХрд░ рджреА рдЧрдИ рд╕рдордп рд╕реЗ рд╡реЗ рдлрд┐рд░ рд╕реЗ рд╡рд┐рдХрд▓реНрдк рдирд╣реАрдВ рд╣реЛ рд╕рдХрддреЗред |
рд╕реАрдорд┐рдд рд╡рд┐рдкрдгрдиреАрдп рдкреНрд░рддрд┐рднреВрддрд┐рдпреЛрдВ рдХреЗ рдмрд╛рд░реЗ рдореЗрдВ рдЬрд╛рдирдХрд╛рд░реА рджреЗрдВред | рдЬрдм рдирд┐рдХрд╛рд╕реА рдпрд╛ рд╕рд╛рдорд╛рдиреНрдп рдЙрдкрдпреЛрдЧ рдХрд╛рдиреВрдиреА рд░реВрдк рд╕реЗ рдкреНрд░рддрд┐рдмрдВрдзрд┐рдд рд╣реЛ рддреЛ рдХрдВрдкрдиреА рд╡рд┐рдкрдгрди рдпреЛрдЧреНрдп рдкреНрд░рддрд┐рднреВрддрд┐рдпреЛрдВ рдХреЛ рдкреНрд░рддрд┐рдмрдВрдзрд┐рдд рдорд╛рдирддреА рд╣реИред | рдХрдВрдкрдиреА рдиреЗ рд╕реАрдорд┐рдд рд╡рд┐рдкрдгрдиреАрдп рдкреНрд░рддрд┐рднреВрддрд┐рдпреЛрдВ рдХреЛ рдРрд╕рд╛ рдорд╛рдирд╛ рдЬрд╛рддрд╛ рд╣реИ рдЬрд┐рд╕рдХреЗ рд▓рд┐рдП рдирд┐рдХрд╛рд╕ рдпрд╛ рд╕рд╛рдорд╛рдиреНрдп рдЙрдкрдпреЛрдЧ рдХрд╛рдиреВрдиреА рд░реВрдк рд╕реЗ рд╕реАрдорд┐рдд рд╣реИред | рд╕реАрдорд┐рдд рд╡рд┐рдкрдгрдиреАрдп рдкреНрд░рддрд┐рднреВрддрд┐рдпреЛрдВ рдХреЗ рдмрд╛рд░реЗ рдореЗрдВ, рдХрдВрдкрдиреА рдЕрдкрдиреЗ рд╡рд┐рдкрдгрдиреАрдп рд╕реЗрдХреБрд░рд┐рдЯреАрдЬ рдХреЛ рд╕реАрдорд┐рдд рд╣реЛрдирд╛ рд╕рдордЭрддреА рд╣реИ рдЬрдмрдХрд┐ рдЙрдирдХрд╛ рдЕрдиреБрдкрд▓рдмреНрдз рдЙрдкрдпреЛрдЧ рдпрд╛ рд╕реАрдорд┐рдд рдХрд░рдирд╛ рд╣реЛ рд╕рдХрддрд╛ рд╣реИред рдпрд╣ рдХрдВрдкрдиреА рдХреЛрдиреНрд╕реЛрд▓рд┐рдбреЗрдЯреЗрдб рдмреИрд▓реЗрдВрд╕ рд╢реЗрдЯреНрд╕ рдореЗрдВ рдЕрдкрдиреЗ рдХреБрд░реЗрдВрдЯ рдпрд╛ рдЕрдХреНрд░реВрдЕрд░реНрдЯ рдорд╛рд░реНрдХреЗрдЯреЗрдмрд▓ рд╕реЗрдХреБрд░рд┐рдЯреАрдЬ рдХреЗ рд░реВрдк рдореЗрдВ рд╢рд╛рдорд┐рд▓ рдХрд░рддреА рд╣реИ, рдЬрд┐рд╕рдХреЗ рдЕрдзрд┐рдХрд╛рд░рд┐рдпреЛрдВ рдХреЗ рдЕрдзрд┐рдХрд╛рд░ рдХреЗ рдЕрдиреБрд╕рд╛рд░ рдкреНрд░рд╛рдкреНрдд рд╣реБрдП рд╕реЗрдХреБрд░рд┐рдЯреАрдЬ рдХрд╛ рд╡рд░реНрдЧ рд╣реЛрддрд╛ рд╣реИред |
2022 рдХреЗ рдХрд░реНрдордЪрд╛рд░реА рд╕реНрдЯреЙрдХ рдпреЛрдЬрдирд╛ рдХреЗ рдмрд╛рд░реЗ рдореЗрдВ рдЬрд╛рдирдХрд╛рд░реА рджреЗрдВред | 2022 рдХрд░реНрдордЪрд╛рд░реА рд╕реНрдЯреЙрдХ рдпреЛрдЬрдирд╛ рдХрд╛рд░реНрдпрдХрд╛рд░реА рдЕрдзрд┐рдХрд╛рд░рд┐рдпреЛрдВ рд╕рд╣рд┐рдд рдХрд░реНрдордЪрд╛рд░рд┐рдпреЛрдВ рдХреЛ рд╡реНрдпрд╛рдкрдХ-рдЖрдзрд╛рд░рд┐рдд рдЗрдХреНрд╡рд┐рдЯреА рдЕрдиреБрджрд╛рди рдкреНрд░рджрд╛рди рдХрд░рддреА рд╣реИ, рдФрд░ рдкреНрд░рддрд┐рдмрдВрдзрд┐рдд рд╕реНрдЯреЙрдХ рджреЗрдиреЗ рдХреА рдЕрдиреБрдорддрд┐ рджреЗрддреА рд╣реИ рдЗрдХрд╛рдЗрдпрд╛рдБ ("рдЖрд░рдПрд╕рдпреВ"), рд╕реНрдЯреЙрдХ рдЕрдиреБрджрд╛рди, рдкреНрд░рджрд░реНрд╢рди-рдЖрдзрд╛рд░рд┐рдд рдкреБрд░рд╕реНрдХрд╛рд░, рд╕реНрдЯреЙрдХ рд╡рд┐рдХрд▓реНрдк рдФрд░ рд╕реНрдЯреЙрдХ рдкреНрд░рд╢рдВрд╕рд╛ рдЕрдзрд┐рдХрд╛рд░ред | 2022 Employee Stock Plan рдХреЗ рддрд╣рдд рдХрд░реНрдордЪрд╛рд░рд┐рдпреЛрдВ рдХреЛ рдмреНрд░реЙрдб-рдмреЗрд╕реНрдб рдЗрдХреНрд╡рд┐рдЯреА рдЧреНрд░реИрдВрдЯреНрд╕ рдкреНрд░рджрд╛рди рдХрд┐рдП рдЬрд╛рддреЗ рд╣реИрдВ, рдЬрд┐рд╕рдореЗрдВ рдХрд╛рд░реНрдпрдХрд╛рд░реА рдЕрдзрд┐рдХрд╛рд░рд┐рдпреЛрдВ рдХреЛ рднреА рд╢рд╛рдорд┐рд▓ рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИред рдЗрд╕ рдпреЛрдЬрдирд╛ рдХреЗ рддрд╣рдд RSUs, рд╕реНрдЯреЙрдХ рдЧреНрд░реИрдВрдЯреНрд╕, рдкреНрд░рджрд░реНрд╢рди-рдЖрдзрд╛рд░рд┐рдд рдЕрд╡рд╛рд░реНрдбреНрд╕, рд╕реНрдЯреЙрдХ рдСрдкреНрд╢рдиреНрд╕ рдФрд░ рд╕реНрдЯреЙрдХ рдРрдкреНрд░реЗрд╢рд┐рдПрд╢рди рд░рд╛рдЗрдЯреНрд╕ рдЬреИрд╕реЗ рдкреНрд░рдХрд╛рд░ рдХреЗ рдЗрдХреНрд╡рд┐рдЯреА рдЕрд╡рд╛рд░реНрдбреНрд╕ рдкреНрд░рджрд╛рди рдХрд┐рдП рдЬрд╛рддреЗ рд╣реИрдВред | The 2022 Employee Stock Plan is a shareholder-approved plan that provides for broad-based equity grants to employees, including executive officers. It permits the granting of restricted stock units (RSUs), stock grants, performance-based awards, stock options, and stock appreciation rights. RSUs granted under the plan generally vest over four years. |
2023, 2022 рдФрд░ 2021 рдореЗрдВ рднреБрдЧрддрд╛рди рдХрд┐рдП рдЧрдП рдЖрдпрдХрд░ рдХреА рдХреБрд▓ рд░рд╛рд╢рд┐ рдХреНрдпрд╛ рд╣реИ? | 2023, 2022 рдФрд░ 2021 рдореЗрдВ рднреБрдЧрддрд╛рди рдХрд┐рдП рдЧрдП рдЖрдпрдХрд░ рдХреА рдХреБрд▓ рд░рд╛рд╢рд┐ рдХреНрд░рдорд╢рдГ $16,741 рдорд┐рд▓рд┐рдпрди, $19,300 рдорд┐рд▓рд┐рдпрди рдФрд░ $14,527 рдорд┐рд▓рд┐рдпрди рд╣реИред | 2023 рдореЗрдВ рднреБрдЧрддрд╛рди рдХрд┐рдП рдЧрдП рдЖрдпрдХрд░ рдХреА рдХреБрд▓ рд░рд╛рд╢рд┐ 16,741 рдХрд░реЛрдбрд╝, 2022 рдореЗрдВ 19,300 рдХрд░реЛрдбрд╝ рдФрд░ 2021 рдореЗрдВ 14,527 рдХрд░реЛрдбрд╝ рдереАред | 2023 рдореЗрдВ рднреБрдЧрддрд╛рди рдХрд┐рдП рдЧрдП рдЖрдпрдХрд░ рдХреА рдХреБрд▓ рд░рд╛рд╢рд┐ $16,741 рдорд┐рд▓рд┐рдпрди рд░реВрдкрдпреЗ рдереАред |
Performance Evaluation
1. Language Detection:
The models effectively detected the language of the queries and provided accurate responses in the corresponding language.
2. Response Accuracy in English:
Meta-Llama-3-8B-Instruct performs well in answering queries based on the English PDFs in comparison of Mistral-7B-Instruct-v0.3 in English.
3. Response Accuracy in Hindi:
Meta-Llama-3-8B-Instruct performs better in Hindi compared to Mistral-7B-Instruct-v0.3.
Mistral-7B-Instruct-v0.3 struggles considerably with Hindi.
4. Overall Performance:
Meta-Llama-3-8B-Instruct consistently outperforms Mistral-7B-Instruct-v0.3 in both languages.
Both models show better performance in English compared to Hindi, which is expected given that they are likely trained primarily on English data.
Challenges
Prompt Engineering:
The effectiveness of the responses was also influenced by the quality of the prompts. In some cases, the prompts used were not optimized, leading to less accurate or incomplete answers.
Partial Hindi Translation:
The Mistral model occasionally failed to fully translate words in Hindi, resulting in responses that included both Hindi and English words.
Conclusion
The integration of the meta-llama/Meta-Llama-3.1-8B-Instruct and mistralai/Mistral-7B-Instruct-v0.3 models in a Multilingual RAG application has proven effective in handling complex, cross-lingual question-answering tasks. By embedding English documents and enabling Multilingual queries, the project demonstrated the potential of these models in a real-world application.
However, the challenges with partial translations in Hindi responses indicate that improvement is needed, especially in handling specific languages. The effectiveness of prompts also played a crucial role in the quality of the generated responses, highlighting the need for careful prompt engineering. Overall, the project succeeded in providing accurate and contextually appropriate answers in both languages.
This case study underscores the effectiveness of combining state-of-the-art language models in creating a bilingual RAG system and highlights the importance of continuous refinement to address specific linguistic challenges.