← All notes

Jul 2026 · 5 min read

Keep slow LLM work off the request path with a queue

How DocReady uses AWS SQS so uploads return instantly, workers scale independently, and a failed LLM call is retried instead of lost.

AWS SQSAsyncLLM

LLM calls take seconds and sometimes fail. If an upload endpoint waits for the model, every slow response ties up a request, and every failure turns into an error for the user.

Accept fast, process later

In DocReady, the upload endpoint stores the file, puts a small job message on an SQS queue, and returns immediately. Workers do the slow part.

python
import json
import boto3

sqs = boto3.client("sqs")

def enqueue_analysis(queue_url: str, document_id: str) -> None:
    sqs.send_message(
        QueueUrl=queue_url,
        MessageBody=json.dumps({"document_id": document_id}),
    )

Workers that can fail safely

A worker only deletes a message after it has finished. If it crashes or the model call fails, the message becomes visible again after the visibility timeout and another worker retries it.

python
def work(queue_url: str) -> None:
    while True:
        resp = sqs.receive_message(QueueUrl=queue_url, MaxNumberOfMessages=5, WaitTimeSeconds=20)
        for msg in resp.get("Messages", []):
            job = json.loads(msg["Body"])
            analyze_document(job["document_id"])   # parse + LLM verdict
            sqs.delete_message(QueueUrl=queue_url, ReceiptHandle=msg["ReceiptHandle"])

What this buys

  • Uploads stay fast no matter how slow the model is
  • Throughput scales by running more workers
  • Failures retry automatically, and a dead-letter queue catches documents that keep failing