- This blog post is a “how I built this” follow-up to A Chat-First Approach to Digital Government Services: The Irembo Chatbot
- The code for this blog post can be found in this Github repository: https://github.com/danksky/irembo-chatbot

“ChatGPT, generate an image of a digital Rwanda”
The goal was simple: to build a system that make IremboGov portal services more accessible to Rwandan citizen by combining conversational AI with a Retrieval-Augmented Generation (RAG) framework.
This post outlines the architecture and key components — OpenAI’s models for language processing, Pinecone for vector storage, and LangChain for managing conversational flows. It also digs into the trade-offs and decisions that shaped the project. If you’re an Irembo developer or manager, you might be curious about how this chatbot project came together and what practical takeaways it offers. This write-up is for anyone who is exploring ways to enhance their own systems or just looking to understand the tech better.
Overview of Components#
GPT and Embedding Model#
I wanted to optimize my two model types, a GPT model and an embedding model, for three key points: most cost-effectiveness, speed, and accuracy.
To experiment for speed and accuracy, I used a script to:
- download the page content from the Irembo support article: Frequently Asked Questions About Driving Licenses
- load the page content into an
InMemoryVectorStore - invoke a query on my LangChain StateGraph flow with various combinations of embedding and GPT models:
- The question I asked was:
Human: How much do I pay for an application for a provisional driving license, and which country is this document about?
- The answer I was looking for was:
AI: The application for a provisional driving license costs RWF 10,000. This document pertains to Rwanda.
GPT Model: Why I chose Open AI GPT-4o#
I chose GPT-4o (specifically gpt-4o-mini-2024–07–18) for this project because it performs well and is cost-effective for the use case. It was straightforward to integrate the GPT-4 LLM model into conversational flows and retrieval logic with LangChain. This setup ensures that the LLM can respond naturally while leveraging the context provided by the Retrieval-Augmented Generation framework.
You can also see in this gist that I tried meta-llama/Llama-3.3–70B-Instruct but it wasn’t free and I didn’t want a paid HuggingFace subscription.
I also tried CohereForAI/c4ai-command-r7b-12–2024 but it wasn’t accurate enough for me.
Embeddings Model: Why I chose Open AI#
To enable effective context retrieval from the Pinecone vector database, I experimented with a variety of embedding models, including:
- The good: Open AI’s
text-embedding-3-largegot the job done, accurately. - The bad: These models simply didn’t work well enough to accurately store the vectors such that the necessary information was retrievable during my testing. The context that was retrieved was wrong, no matter how much I adjusted the
kvalue for my vector database’ssimilarity_search.
– HuggingFace’sall-MiniLM-L6-v2
– Open AI’stext-embedding-3-small
– Google’stext-embedding-004 - The ugly: Mistral AI’s
mistral-embedgave me grief, because the Langchain package,langchain-mistralai, was requesting something from HuggingFace and the request was breaking.
After testing these options, I ultimately chose OpenAI’s text-embedding-3-large. It provided the most reliable results for retrieving relevant content, so I trusted it to do the same when I changed my vector database from InMemoryVectorStore to PineconeVectorStore.
Pinecone#
I chose Pinecone as a scalable and persistent vector database, eliminating the need to rebuild the vector database on each deployment. This persistence allows us to reuse embeddings, reducing the API consumption for OpenAI’s embedding model. Pinecone’s generous free tier will hopefully offer enough capacity to host our demo as users query Irembo’s support articles via the live chatbot.
Setting up Pinecone involved creating an index optimized for storing embeddings from OpenAI’s text-embedding-3-large model. LangChain’s PineconeVectorStore integration made it straightforward to connect Pinecone to the chatbot, enabling seamless similarity-based retrieval of stored documents.
State Graph#

Graph Setup
The state graph forms the backbone of the chatbot, organizing how it processes user input and generates responses. Built using LangChain’s StateGraph, the design ensures that each stage of the interaction is modular and manageable, from understanding the user’s query to retrieving relevant context and generating a final response.
The graph includes three primary nodes:
query_or_respond: Determines if a query requires retrieval or can be answered directly using the conversational context. This node acts as the decision point. It evaluates the user’s query to decide whether to proceed with retrieval or skip directly to response generation. By ensuring retrieval is only used when necessary, this step optimizes both performance and cost.retrieve: Fetches the most relevant information from the Pinecone vector database when needed. If the decision is made to fetch additional context, this node interfaces with the Pinecone database. It performs a similarity search on the stored embeddings to find the most relevant documents. The results are passed along to the next step for integration into the response.generate: Uses the LLM to create a response, incorporating any retrieved context. At this stage, the LLM generates a complete response. It combines the query, retrieved context (if any), and the conversation history to produce a concise, accurate answer tailored to the user’s needs.
Node Ordering
The nodes are connected through conditional edges to maintain a logical flow. For example:
query_or_respond→retrieve→generate: For queries requiring external context.query_or_respond→generate: For straightforward conversational queries.
By leveraging LangChain’s add_conditional_edges and checkpointer (using MemorySaver), the system handles transitions seamlessly while preserving conversation history across sessions.
This modular design makes the chatbot robust, efficient, and easy to extend. Each node focuses on a specific task, ensuring clarity and maintainability in the overall workflow.
LangChainTracer
LangChainTracer was used to debug and monitor the flow of the chatbot during development. It tracked how queries moved through the state graph, providing insights into transitions between nodes like query_or_respond, retrieve, and generate. This made it easier to spot bottlenecks or errors, especially in complex interactions.
Adding LangChainTracer was simple — just a few lines to set it up as a callback in LangChain’s configuration. This low-effort integration proved invaluable for ensuring that each node performed as expected and that retrieval and generation steps were correctly synchronized. It also allowed for easy performance tuning and iterative improvements throughout the project.
Gradio Interface
The chatbot’s interface was built using Gradio’s ChatInterface, which simplifies the process of creating conversational applications. This component allowed me to focus on integrating the backend functionality without needing to build a custom front-end from scratch.
One major advantage of using Gradio is its seamless deployment capability with gradio deploy, enabling the app to be published directly to Hugging Face Spaces. This made it easy to share a live, interactive version of the chatbot for demos and testing.
The integration of ChatInterface provided a clean, user-friendly way to manage conversations while maintaining support for features like session tracking through Gradio’s request parameter. This ensured that the chatbot could handle user interactions efficiently and maintain context across sessions.
How the Pinecone Database Was Populated#
To populate the Pinecone vector database with Irembo’s support articles, I developed a Python script that automates the process of scraping, processing, and storing the content. Here’s an overview of the key components and their roles:
Scraping Irembo’s Support Articles
The script begins by fetching the sitemap from Irembo’s support portal, which lists all available articles. By parsing this sitemap, the script identifies the URLs of individual support articles for subsequent processing.
Processing and Cleaning the Content
For each article URL, the script performs the following steps:
HTML Retrieval: Downloads the HTML content of the article page.
Content Extraction: Utilizes Beautiful Soup to parse the HTML and extract the main content, specifically targeting elements that contain the article’s text.
Text Cleaning: Processes the extracted text to remove any extraneous whitespace, HTML tags, or irrelevant sections, ensuring that only the meaningful content remains.
Splitting Documents into Chunks
To enhance the efficiency of the vector database and improve the relevance of search results, the cleaned text is divided into smaller chunks. This segmentation allows the system to retrieve specific sections of an article that are most pertinent to a user’s query.
Generating Embeddings
Each text chunk is then converted into a numerical vector representation using OpenAI’s text-embedding-3-large model. These embeddings capture the semantic meaning of the text, enabling effective similarity searches within the vector database.
Upserting into Pinecone
The final step involves inserting (upserting) these embeddings into the Pinecone vector database. Each entry in the database includes:
ID: A unique identifier for the text chunk.
Embedding: The vector representation of the text.
Metadata: Additional information such as the source URL and the original text content, which can be useful for reference and debugging.
By automating this pipeline, the system ensures that the Pinecone database remains up-to-date with the latest support articles from Irembo, facilitating accurate and efficient retrieval of information in response to user queries.
For a detailed implementation, you can refer to the scrape_irembo.py script in the GitHub repository.
Why Not Just Use BotoPress?#
BotoPress is an appealing option for quickly building chatbots, especially with its straightforward drag-and-drop interface and built-in integrations with platforms like WhatsApp, Facebook Messenger, and others. It would have streamlined the process of deploying a chatbot to multiple channels without much additional effort.
However, the focus of this project was to deeply explore and customize a Retrieval-Augmented Generation (RAG) workflow. By building the system from scratch, I was able to directly control how data is stored in Pinecone, how retrieval integrates with the chatbot’s conversational logic, and how embeddings are managed. This level of granularity wouldn’t have been possible with BotoPress, which abstracts much of the implementation.
Moreover, developing this solution allowed me to break down the components of a RAG workflow and understand how systems like BotoPress might function under the hood. While BotoPress could still be considered for future iterations, this hands-on approach provided valuable insights and flexibility for tailoring the chatbot to Irembo’s specific needs.
Lessons Learned#
poetryfor Dependency Management: Poetry streamlined managing dependencies and environments, ensuring a consistent and reproducible setup.Experimentation is key: Testing multiple models (HuggingFace, OpenAI, Mistral, etc.) and tools like Pinecone highlighted the importance of flexibility during development.
LangChain simplifies complexity: LangChain’s abstractions, like StateGraph and VectorStores, made integrating retrieval and LLM workflows straightforward while keeping the system modular.
Pinecone saves resources: Pinecone’s persistent storage eliminated the need to regenerate embeddings, reducing API usage and speeding up deployments.
Custom builds offer insight: Building the chatbot manually provided a deeper understanding of RAG workflows, making it easier to customize and optimize.
Conclusion#
Building this chatbot was an exercise in understanding and customizing a Retrieval-Augmented Generation (RAG) workflow to meet specific requirements. By combining tools like OpenAI’s models, Pinecone for vector storage, and LangChain for conversational logic, I was able to create a modular, scalable, and cost-effective solution for improving access to IremboGov portal services.
The project highlighted the value of experimentation, from testing different embedding models to refining retrieval logic and integrating state graph workflows. It also reinforced the importance of tools like Poetry for dependency management, LangChainTracer for debugging, and Gradio for a seamless user interface.
This approach not only delivered a functional chatbot but also provided a deeper understanding of the underlying technology, which can be applied to future projects. If you’re exploring similar solutions or have questions about the tools and techniques discussed here, feel free to reach out — I’d love to hear your thoughts or collaborate further.


