Skip to content
Zain Mahmood
Selected work

EldenLore RAG

A lore chatbot is easy to build and almost always subtly wrong, because a language model will happily invent a plausible answer. The engineering here is the retrieval pipeline and the gates that make it refuse instead of guessing.

Type
Retrieval engineering
State
Complete · v1.0.0
Role
Solo, end to end
  • Python
  • ChromaDB
  • all-MiniLM-L6-v2
  • Groq
  • Streamlit

At a glance

  • Refuses when retrieval is weak instead of guessing
  • Multi-pass retrieval with a HyDE fallback
  • Citations back to the source passage on every answer

What I owned

Solo project. Corpus cleaning and chunking, the retrieval pipeline, the grounding and refusal logic, and the Streamlit interface.

Decisions, and what they cost

  1. A relevance gate that refuses instead of answering weakly

    A naive RAG pipeline always answers: retrieve the top-k chunks, hand them to the model, get prose back. If retrieval was weak, the model fills the gap from its own priors and the answer looks exactly as confident as a correct one. A gate between retrieval and generation refuses in character when passages are too weak, and a hallucination guard checks entity grounding on the way out, re-asking in strict mode if it fails.

    What I gave upThe bot answers fewer questions. A confident wrong answer about lore is worse than an admission of ignorance, because the user cannot tell the two apart.

  2. Multi-pass retrieval rather than one embedding lookup

    One vector search over the raw question misses badly on lore queries full of pronouns and indirect references. Retrieval fuses several passes: the raw query, LLM-expanded variants, entity-relation queries against a curated knowledge graph, and a category-filtered search, with HyDE as a fallback. Pronouns are rewritten first, so 'Who was her twin?' becomes 'Who was Malenia's twin?' before it hits the index.

    What I gave upSeveral model calls per question instead of one, so it is slower and burns far more rate limit. Expansion and reranking run on a small fast model to keep that affordable.

  3. Sentence-boundary chunking with de-duplication

    The corpus is roughly 6,400 lore passages from wiki sources, full of navigation artifacts and near-duplicate text. Chunks are cut on sentence boundaries so a chunk is never half a sentence, then de-duplicated and scrubbed before embedding.

    What I gave upChunks are uneven in length, which makes the embedding space less uniform than fixed windows. Retrieval quality was better in practice, so I kept it.

Known limitations

  • Runs on the free Groq tier, so requests can fail under rate limiting. Those surface as an explicit retry rather than a degraded answer.
  • Coverage is bounded by the corpus. Anything outside it gets an in-character refusal.
  • No hosted demo. The repo runs locally.

Where it stands

v1.0.0 released. Answers use llama-3.3-70b-versatile; expansion and reranking use llama-3.1-8b-instant. Answers carry citations back to the retrieved passages, and a lore quiz mode runs on the same retrieval path.