← All projects

01 · Async processing

DocReady.

Analyzes DOCX and PDF files for knowledge-base migration readiness. AWS SQS queues analysis jobs for concurrent processing, and an LLM verdict engine scores each document.

Queue-backed concurrency
PythonAWS SQSGemini AIReact
View source ↗

01 · The problem

Before moving documents into a knowledge base, teams need to know which files are actually ready. Checking DOCX and PDF files by hand is slow and inconsistent, and running an LLM inline on every upload would make the API slow and fragile.

02 · What I built

  • An upload endpoint that accepts DOCX and PDF files and returns straight away with a job reference.
  • An AWS SQS queue that holds analysis jobs, decoupling upload handling from the LLM processing pipeline.
  • Workers that pull jobs concurrently, parse each document, and send it to Gemini for analysis.
  • An LLM verdict engine that returns a structured readiness score (READY / NOT READY) with strengths, issues, and improvement suggestions.
  • A React interface to upload files and review each verdict.

03 · Key decisions

  1. 01

    Queue the work instead of doing it inline

    LLM calls are slow and can fail. Putting them behind SQS keeps uploads fast, and a failed job can be retried without the user uploading again.

  2. 02

    Structured output, not free text

    Asking the model for a fixed shape (score, verdict, strengths, issues) makes results comparable across documents and simple to render.

  3. 03

    Scale by adding workers

    Several workers read from the same queue, so a batch of documents is processed in parallel and throughput grows with the number of workers.

04 · Results

  • Multi-document batches processed concurrently
  • A consistent, structured readiness report for every file
  • Upload latency independent of LLM latency