02 · Retrieval-augmented AI
RAG Knowledge Assistant.
Answers questions over internal docs: documents are chunked and embedded into pgvector, the top-k nearest passages are retrieved by semantic search, and Azure OpenAI writes an answer that cites its sources.
01 · The problem
Engineers lose time digging through long documentation, runbooks, and past incidents to answer the same questions. Keyword search misses answers that use different wording.
02 · What I built
- An ingestion pipeline that splits documents into chunks, embeds them, and stores the vectors in PostgreSQL with pgvector.
- Semantic search that embeds each question and retrieves the top-k most similar chunks.
- Answer generation on Azure OpenAI, grounded only in the retrieved chunks and citing each source.
- A Node.js API that serves questions from internal tools.
03 · Key decisions
- 01
pgvector instead of a separate vector database
Embeddings live next to the existing relational data, so there is one system to run, and similarity search can be combined with ordinary SQL filters.
- 02
Retrieve first, then generate
The model only sees the top-k retrieved passages, which keeps answers grounded in real documents instead of the model's general knowledge.
- 03
Cite every answer
Citations let people check the source and build trust in the tool, and make wrong answers easy to trace back to the passage that caused them.
04 · Results
- Answers grounded in internal documents, with citations
- Meaning-based search that finds answers keyword search misses
- One Postgres instance for both relational data and vectors