Screenshot of RAG Document Assistant

2026 · Solo

RAG Document Assistant

Ask questions over your own documents — and a hand-labeled, automatically-scored eval harness keeps the answers honest.

Phases 1–5 substantially complete

The RAG pipeline (src/ingest.py → src/retrieve.py → src/generate.py) chunks documents with a RecursiveCharacterTextSplitter (500 chars/50 overlap), embeds with sentence-transformers into a persistent Chroma store, and grounds llama3.2's answers in the top-k retrieved chunks — the same retrieve()/generate_answer() functions power both the CLI (src/ask.py) and a two-page Streamlit app (chat + an Evaluation Dashboard). The eval half runs a 32-question hand-labeled dataset through that identical code path, scores it with Ragas (faithfulness, context_recall) using a local Ollama model as judge, and a pytest regression gate fails the suite if either metric drops more than 0.05 against the committed baseline — run for real (a ~33-minute full pass), not just structurally checked. An experiment runner tries a chunk-size/top-k/embedding-model change against the same dataset in its own isolated Chroma collection, logged to a CSV, without ever touching the real vector store.

Features

  • Local-first RAG pipeline — sentence-transformers embeddings, Chroma vector store, llama3.2 via Ollama for generation, no API keys or cloud cost
  • Two-page Streamlit app — chat UI with document upload/rebuild-in-place, plus an Evaluation Dashboard showing baseline scores and every experiment run
  • Document scoping — narrows retrieval to one file with a higher top-k, added after vague questions in a multi-document corpus retrieved the wrong document's boilerplate
  • 32-question hand-labeled eval dataset, including deliberately unanswerable questions, scored with Ragas (faithfulness, context_recall) via a local Ollama judge
  • Regression gate — a pytest suite that regenerates and rescores the full dataset from scratch and fails on a >0.05 metric drop vs. baseline
  • Isolated experimentation — chunk size / top-k / embedding-model sweeps run in their own Chroma collections, logged to a CSV, never touching the real data
  • Table-of-contents filtering at ingest — drops navigational pages so a section title repeated in a TOC can't out-rank that section's own real content

Tech stack

RAG Pipeline

  • Python
  • LangChain
  • sentence-transformers
  • Chroma
  • Ollama (llama3.2)

Evaluation

  • Ragas (faithfulness, context_recall)
  • pytest regression gate
  • local Ollama judge

App

  • Streamlit
  • CLI (src/ask.py)