PT

Nº 05of 22Featured

Project Wraeclast

RAG assistant: daily ingestion on GitHub Actions, vector search in Postgres with pgvector, and retrieval-grounded answers

Year
2026
Status
IN PROGRESS
Role
Personal project · solo
Stack
Python · FastAPI · PostgreSQL · pgvector +4
wraeclast
PW

Python

No public screen — generated cover

01The project

Personal assistant for a game whose meta shifts with every patch. A daily job collects the game economy, the owner's character and community content, summarizes each document into structured JSON with an LLM, and stores the embedding in Postgres with pgvector. On a question, the FastAPI API embeds the text, retrieves the nearest passages by cosine distance and assembles a context block with prices, a farm ranking and the character profile — the LLM answers from that block only. There is no trained model and no fine-tuning: the intelligence is the curated corpus that grows every day. The Next.js site has a today screen, a farms screen and a crafting bench screen, plus a graph of the knowledge collected.

02What was built

  1. Vector search in Postgres with pgvector: 1024-dimension embeddings (Matryoshka truncation so they fit the HNSW index) retrieved by cosine distance and filterable by topic

  2. No fine-tuning and no model of my own: the chat context is assembled from the retrieved passages plus prices, farms and the character profile, and the prompt orders the model to answer from that context only and to flag when the data is missing

  3. Heavy ingestion kept out of the API: the daily cron runs on GitHub Actions, with no execution time limit; only the reads and /chat run as serverless functions on Vercel

  4. 350 tests passing locally under pytest and CI running ruff + pytest on every push; the 80% per-module coverage target is written down and dated in the ROADMAP

  5. /chat is gated by a token compared in constant time with hmac.compare_digest, and it fails closed when the token is not configured

03Decisions

Decision 01

Question: Why RAG and not fine-tuning?

Answer: The content changes with every game patch. Retraining a model on every change is expensive and goes stale fast; a curated corpus that gains new documents every day is current by construction, and the LLM only comes in to turn text into JSON and to answer grounded in what was retrieved.

Decision 02

Question: Why does ingestion run on GitHub Actions instead of inside the API?

Answer: A serverless function has an execution time ceiling, and the daily ingestion (scraping, embeddings and LLM curation) does not fit under it. The cron job runs on Actions and only writes to the database; the API keeps reads and chat, which are fast enough.

Decision 03

Question: Why truncate the embedding to 1024 dimensions?

Answer: The provider returns 3072 dimensions by default and pgvector's HNSW index has a practical limit below that. Matryoshka truncation (the dimensions parameter on the call itself) makes the vector fit the index without switching providers or giving up similarity search.

04Stack

  • Python
  • FastAPI
  • PostgreSQL
  • pgvector
  • Next.js
  • TypeScript
  • GitHub Actions
  • Vercel
Next project →Nº 06

Transactional e-commerce in Spring Boot: outbox with exponential retry, webhook idempotency, and rate limiting in Redis behind a circuit breaker