EldenLore RAG
A lore chatbot is easy to build and almost always subtly wrong, because a language model will happily invent a plausible answer. The engineering here is the retrieval pipeline and the gates that make it refuse instead of guessing.
- Type
- Retrieval engineering
- State
- Complete · v1.0.0
- Role
- Solo, end to end
- Python
- ChromaDB
- all-MiniLM-L6-v2
- Groq
- Streamlit
At a glance
- Refuses when retrieval is weak instead of guessing
- Multi-pass retrieval with a HyDE fallback
- Citations back to the source passage on every answer
What I owned
Solo project. Corpus cleaning and chunking, the retrieval pipeline, the grounding and refusal logic, and the Streamlit interface.
Decisions, and what they cost
A relevance gate that refuses instead of answering weakly
A naive RAG pipeline always answers: retrieve the top-k chunks, hand them to the model, get prose back. If retrieval was weak, the model fills the gap from its own priors and the answer looks exactly as confident as a correct one. A gate between retrieval and generation refuses in character when passages are too weak, and a hallucination guard checks entity grounding on the way out, re-asking in strict mode if it fails.
What I gave upThe bot answers fewer questions. A confident wrong answer about lore is worse than an admission of ignorance, because the user cannot tell the two apart.
Multi-pass retrieval rather than one embedding lookup
One vector search over the raw question misses badly on lore queries full of pronouns and indirect references. Retrieval fuses several passes: the raw query, LLM-expanded variants, entity-relation queries against a curated knowledge graph, and a category-filtered search, with HyDE as a fallback. Pronouns are rewritten first, so 'Who was her twin?' becomes 'Who was Malenia's twin?' before it hits the index.
What I gave upSeveral model calls per question instead of one, so it is slower and burns far more rate limit. Expansion and reranking run on a small fast model to keep that affordable.
Sentence-boundary chunking with de-duplication
The corpus is roughly 6,400 lore passages from wiki sources, full of navigation artifacts and near-duplicate text. Chunks are cut on sentence boundaries so a chunk is never half a sentence, then de-duplicated and scrubbed before embedding.
What I gave upChunks are uneven in length, which makes the embedding space less uniform than fixed windows. Retrieval quality was better in practice, so I kept it.
Known limitations
- Runs on the free Groq tier, so requests can fail under rate limiting. Those surface as an explicit retry rather than a degraded answer.
- Coverage is bounded by the corpus. Anything outside it gets an in-character refusal.
- No hosted demo. The repo runs locally.
Where it stands
v1.0.0 released. Answers use llama-3.3-70b-versatile; expansion and reranking use llama-3.1-8b-instant. Answers carry citations back to the retrieved passages, and a lore quiz mode runs on the same retrieval path.