← All projects

02 · Retrieval-augmented AI

RAG Knowledge Assistant.

Answers questions over internal docs: documents are chunked and embedded into pgvector, the top-k nearest passages are retrieved by semantic search, and Azure OpenAI writes an answer that cites its sources.

Top-k retrieval · cited answers
Azure OpenAIpgvectorPostgreSQLNode.js

01 · The problem

Engineers lose time digging through long documentation, runbooks, and past incidents to answer the same questions. Keyword search misses answers that use different wording.

02 · What I built

  • An ingestion pipeline that splits documents into chunks, embeds them, and stores the vectors in PostgreSQL with pgvector.
  • Semantic search that embeds each question and retrieves the top-k most similar chunks.
  • Answer generation on Azure OpenAI, grounded only in the retrieved chunks and citing each source.
  • A Node.js API that serves questions from internal tools.

03 · Key decisions

  1. 01

    pgvector instead of a separate vector database

    Embeddings live next to the existing relational data, so there is one system to run, and similarity search can be combined with ordinary SQL filters.

  2. 02

    Retrieve first, then generate

    The model only sees the top-k retrieved passages, which keeps answers grounded in real documents instead of the model's general knowledge.

  3. 03

    Cite every answer

    Citations let people check the source and build trust in the tool, and make wrong answers easy to trace back to the passage that caused them.

04 · Results

  • Answers grounded in internal documents, with citations
  • Meaning-based search that finds answers keyword search misses
  • One Postgres instance for both relational data and vectors