A RAG chatbot that knows when to look something up exactly instead of searching for it.
Complete Streamlit app. Runs locally with a free Groq API key.
Overview
A conversational assistant over the Netflix titles catalogue. Questions about directors are answered by an exact lookup over the dataset; everything else goes through retrieval-augmented generation with conversation memory.
Problem
Pure vector search is unreliable for exact attributes: a question like 'what did this director make' wants every matching row, not the four most similar chunks.
Approach
Route the question. Director questions hit a deterministic pandas lookup that returns complete answers. Open questions go to a ConversationalRetrievalChain over a Chroma index of MiniLM embeddings, answered by Llama 3.1 8B on Groq at temperature 0.
Architecture
Select a component to see what it does. Blue packets show the direction data moves.
- Question to Router
- Router to Exact lookup (director)
- Router to Chroma + MiniLM (open)
- Chroma + MiniLM to Llama 3.1 8B (context)
- Exact lookup to Answer + sources
- Llama 3.1 8B to Answer + sources
What it does
- Hybrid routing between exact lookup and semantic retrieval.
- Local sentence-transformer embeddings (all-MiniLM-L6-v2), so no embedding API is needed.
- Conversation memory for follow-up questions.
- Retrieved source titles shown under each answer, with a badge for which route answered.
Technical challenges
- Exact-match routing must not be triggered by short titles that happen to appear inside an unrelated question.
Engineering decisions
- Temperature 0 and k=4 retrieval to keep answers grounded in retrieved rows.
Limitations
- Needs a Groq API key (free tier) for generation.
- The Chroma index is rebuilt in memory on cold start.