SÃO PAULO, BR · AVAILABLE FOR WORK · BACKEND, WEB & AI

LeonardoSouza

Backend, Web & AI

COMPUTER SCIENCE · INSPER · FIFTH SEMESTER · SÃO PAULO

I build backend, web, and AI that run in production: an LLM assistant over the user's own data, vector search with pgvector for answers grounded in the retrieved context, an outbox with exponential backoff and idempotency, and a Next.js front-end in a system with real users. Three.js and GLSL come in when the problem is visualization — not before.

curl -sO https://portfolio-souzxxxs-projects.vercel.app/leonardo-souza-cv-en.pdf
49k
matches in the ML pipeline
218
test files
12
domain modules
4
distributed services
souzxxx

#02 · SELECTED WORK

Selected projects

A selection of what holds up in a technical conversation: an AI product in production, vector search with pgvector, transactional e-commerce with an outbox and idempotency, and web with real users. The external links in this section are checked automatically every week.

FinanceHub — Dashboard — balance, projection and the week's summary

Dashboard — balance, projection and the week's summary · EXPAND

1 of 4

FinanceHub — Dashboard — balance, projection and the week's summary

Dashboard — balance, projection and the week's summary

BUILD 001 · YEAR 2026 · REPO: PRIVATE

LIVE

FinanceHub

End-to-end personal finance platform with AI, in production

Personal finance app with Supabase auth, a balance and projection dashboard, transaction management, per-category budgets, recurring entries, import/export and a conversational AI assistant (Luna). Turbo monorepo with a Python backend and a Next.js frontend, continuously deployed on Vercel. The screenshots alongside are from the authenticated platform in production.

  • 01Luna: context assembled from the authenticated user's own transactions, categories and budgets, through parameterized queries filtered by user_id
  • 02LLM through the OpenAI SDK pointed at Groq (Llama 3.3 70B), a 500-token output cap and history truncated to the last 10 messages
  • 03Per-route rate limiting with slowapi: 30 req/h on chat, 10/h on insights, 5/h on the monthly report
  • 04Row level security in Postgres on Luna's digest tables, on top of Supabase auth
  • 05Turbo monorepo: FastAPI API + Next.js frontend, continuously deployed on Vercel
DECISION 01

Why Groq and not OpenAI directly?

The SDK is the same; only the base_url changes. Llama 3.3 70B on Groq delivers far lower latency at a cost a personal project can carry, and switching providers is one line of configuration if the trade-off changes.

DECISION 02

Why truncate history at 10 messages?

The month's financial context already fills the system prompt. With no cap, a long conversation drives cost per request up and answer quality down — 10 messages covers the continuity of a finance chat.

STACK: Next.js · TypeScript · Python · Turbo · Supabase · PostgreSQL · Groq · Llama 3.3 · Vercel

VIEW LIVE ↗

PRIVATE REPOSITORY

QuintoAndar Pricing — Public calculator — 8 fields and the price the API estimates

Public calculator — 8 fields and the price the API estimates · EXPAND

1 of 4

QuintoAndar Pricing — Public calculator — 8 fields and the price the API estimates

Public calculator — 8 fields and the price the API estimates

BUILD 002 · YEAR 2026 · REPO: PRIVATE

SHIPPED

QuintoAndar Pricing

Home prices served by a FastAPI API on AWS: 9.32% MAPE against actual 2010 sales, every request logged, continuous deploy with a healthcheck

Fourth-semester Insper sprint with QuintoAndar as the partner: a team of four had to answer what a home sells for and keep that answer running in production. The dataset is Ames Housing — 1,285 residential properties from 2006 to 2009, 80 features — and the final test ran against 175 actual 2010 sales, a temporal holdout whose ids are disjoint from the training set. Two models came out of it: a 4-feature linear regression to explain price to agents and owners (11.9% MAPE, R² 0.832) and a Gradient Boosting model for the public calculator. Of the 52 commits in the application repository, 20 are mine — the largest share on the team: the FastAPI API, prediction logging, the 36 tests, the validation script and report against the 2010 data, and the minimal 8-field model.

  • 01Gradient Boosting picked among 5 models in chronological validation: $16,117 MAE and 9.83% MAPE on the 2009 holdout, against $57,077 and 34.69% for the median baseline
  • 02Final validation against the 175 actual 2010 sales: $15,810 MAE, 9.32% MAPE, 26,071 RMSE and R² 0.894, with 92% of the properties within 20% error
  • 03Public calculator cut from 12 fields to 8 by forward selection, at the knee of the error curve
  • 04FastAPI 0.115 API with /predict, /predict/simples, /metrics and /health, Pydantic contracts, and the preprocessing and features ported over from the model repo so they match training exactly
  • 05Every request written to SQLite — 422s included — with latency, model version, input and output; the real-usage report closed at 300 calls with a p95 of 11.7 ms and 94.3% success
  • 0636 pytest tests across 8 files, and CI/CD that runs the suite on every push and, on merge to main, does SSH + rsync and docker compose up --build --wait on EC2, with a healthcheck after the deploy
DECISION 01

Why a chronological split and not a random one?

Home prices move over time. A random split puts a 2009 sale in training and a 2007 sale in test, and the model ends up evaluated while knowing the future. The cut is by date, validation was the year 2009, and the final test ran against the 175 sales from 2010 — a year training never saw, with disjoint ids.

DECISION 02

Why log the requests that fail with 422 as well?

Logging only successful predictions measures the model, not the service. Recording the 422s alongside them is what shows which contract the client is getting wrong and how often — that is where the 94.3% success rate over the report's 300 calls comes from, instead of an impression that everything was fine.

DECISION 03

Why not retrain after the 2010 validation?

The model trained through 2009 came in at 9.32% MAPE on a year it had never seen, against 9.83% on the 2009 holdout — so there was no degradation to fix. Retraining with no sign of degradation swaps a measured model for a new, unmeasured one; the decision to leave it alone is written down in the validation report.

STACK: Python · FastAPI · scikit-learn · Pydantic · SQLite · Docker · GitHub Actions · AWS EC2

PRIVATE REPOSITORY

TE

Java 21

BUILD 003 · YEAR 2026 · REPO: PRIVATE

IN PROGRESS

Transactional e-commerce (under NDA)

Transactional e-commerce in Spring Boot: outbox with exponential retry, webhook idempotency, and rate limiting in Redis behind a circuit breaker

A Brazilian brand's own store — name under contract — as a Java 21 / Spring Boot 4.1 modular monolith with a Next.js storefront. The system moves money, so most of the engineering is about the unhappy path: payment webhooks validated by signature and deduplicated against a processed-webhooks table; transactional emails in an outbox, claimed in a short transaction and sent outside any lock; retries with exponential backoff and an idempotency key at the provider; per-key rate limiting in Redis through an atomic Lua script that fails open when Redis goes down. 12 domain modules, Flyway migrations, and integration tests against real Postgres and Redis with Testcontainers.

  • 01Transactional outbox: claim in a short transaction, send without holding a lock, exponential backoff from 30s doubling to a 1h ceiling with jitter
  • 02End-to-end idempotency: each webhook processed exactly once, an Idempotency-Key at the email provider
  • 03Per-key rate limiting in Redis with an atomic Lua script that fails open, behind a circuit breaker that probes in half-open
  • 04Payment webhooks validated by HMAC signature; the NF-e webhook by its own secret token under constant-time comparison — both fail closed
  • 0512 domain modules, Flyway, and integration tests with Testcontainers (real Postgres + Redis)
DECISION 01

Why an outbox instead of calling the email provider inside the transaction?

An HTTP call inside a transaction holds a pool connection while it waits on the network. The outbox breaks that into three steps: a short transaction that claims the row, a send with no lock at all, and a short transaction that records the result. The Idempotency-Key is the row id, so a resend after a crash is a no-op.

DECISION 02

Why does the rate limiter fail OPEN?

If Redis goes down, blocking all traffic takes checkout down with it — letting traffic through for a few minutes costs less than not selling. It logs ERROR so the failure raises an alert, and the circuit breaker keeps every request from paying the connection timeout while Redis is out.

DECISION 03

Why a modular monolith and not microservices?

Solo project, ~100 orders/month. Microservices here would buy network latency and deployment complexity without solving a single problem I actually have. The modules have clear boundaries for the day splitting them pays off.

STACK: Java 21 · Spring Boot · PostgreSQL · Redis · Next.js · TypeScript · Testcontainers · Flyway · Docker

PRIVATE REPOSITORY

CC

Next.js

BUILD 004 · YEAR 2026 · REPO: PRIVATE

LIVE

Construction company internal portal

Homeowner's manual module in a Next.js portal in production: auditable PDF import, PDF/Excel export, and three performance optimizations

A Next.js 16 and React 19 portal that puts a construction company's internal systems behind a single login — post-handover support, the client portal, site evaluation, personnel administration, and the homeowner's manual. It is a team project in a private repository: of the 583 commits, around 150 are mine, and 108 of those sit in the homeowner's manual module — which is what I describe here. The document that module produces is read by the buyer as part of the contract, so the requirement is not screen count: it is that the warranty data be correct and that the path to a correction be auditable.

  • 01Homeowner's manual module: a real 197-page PDF manual turns into a structured catalog of 33 building-system records, 146 warranties, 39 preventive maintenance items, 78 care instructions and 131 warranty-voiding cases
  • 02The import script never writes to the database: it emits a .sql that a human reads as a diff before it becomes a migration, and it checks the extraction by structural counts instead of by eyeballing text
  • 03Three performance optimizations of mine in that same module: the N+1 gone (one query per development instead of one per unit), the PDF logo embedded once instead of per page, and the record list loaded without pulling every record's full text
  • 04A team system in production: all 204 API routes go through a single wrapper that centralizes authentication, role, app permission and capability, re-read from the database on every request
  • 05The suite runs against a real, disposable Postgres, with the actual migrations and an explicit guard so it can never point at the production database — 181 Vitest test files and 6 end-to-end Playwright suites
DECISION 01

How do you import warranty data from a PDF without risking a bad extraction landing in a contractual document?

The import script never writes to the database. It emits a .sql you review as a diff, and it only becomes a migration after someone reads it; the check is structural (how many records, how many warranties, how many preventive items), because eyeballing hundreds of lines of text is not verification, it is hope.

DECISION 02

Why does this go in the portfolio as a module and not as a whole product?

It is a team project: most of the repository's commits belong to another developer. What I claim is what the history confirms as mine — the homeowner's manual module and its optimizations. The rest of the system is here as context, not as authorship.

STACK: Next.js · React · TypeScript · Drizzle ORM · PostgreSQL · Tailwind · Vitest · Playwright · Vercel

PRIVATE REPOSITORY

Sentinel — screenshot

EXPAND

Sentinel — screenshot

screenshot

BUILD 005 · YEAR 2026

LIVE

Sentinel

Real-time monitoring: FastAPI backend metrics streamed over WebSocket and rendered in 3D

Full-stack 3D dashboard where my GitHub repositories orbit a reactive core like satellites. Repository size maps to stars + forks, color maps to language, and orbital speed maps to the most recent push. Hand-written GLSL shaders draw the plasma, the holographic grids and the glitch layers. The FastAPI backend streams CPU, RAM and disk metrics over WebSocket, and those are what drive the core's pulse and color.

  • 01Metrics streamed over WebSocket with automatic reconnection and a retry cap
  • 02Hand-written GLSL shaders (plasma, holographic grid, glitch)
  • 03Fly-to camera with GSAP, bloom and chromatic aberration
  • 04A FastAPI backend pushes CPU/RAM/disk metrics that drive the scene

STACK: Next.js 16 · Three.js · React Three Fiber · GLSL · FastAPI · WebSocket · TypeScript · Docker

PW

Python

BUILD 006 · YEAR 2026

IN PROGRESS

Project Wraeclast

RAG assistant: daily ingestion on GitHub Actions, vector search in Postgres with pgvector, and retrieval-grounded answers

Personal assistant for a game whose meta shifts with every patch. A daily job collects the game economy, the owner's character and community content, summarizes each document into structured JSON with an LLM, and stores the embedding in Postgres with pgvector. On a question, the FastAPI API embeds the text, retrieves the nearest passages by cosine distance and assembles a context block with prices, a farm ranking and the character profile — the LLM answers from that block only. There is no trained model and no fine-tuning: the intelligence is the curated corpus that grows every day. The Next.js site has a today screen, a farms screen and a crafting bench screen, plus a graph of the knowledge collected.

  • 01Vector search in Postgres with pgvector: 1024-dimension embeddings (Matryoshka truncation so they fit the HNSW index) retrieved by cosine distance and filterable by topic
  • 02No fine-tuning and no model of my own: the chat context is assembled from the retrieved passages plus prices, farms and the character profile, and the prompt orders the model to answer from that context only and to flag when the data is missing
  • 03Heavy ingestion kept out of the API: the daily cron runs on GitHub Actions, with no execution time limit; only the reads and /chat run as serverless functions on Vercel
  • 04350 tests passing locally under pytest and CI running ruff + pytest on every push; the 80% per-module coverage target is written down and dated in the ROADMAP
  • 05/chat is gated by a token compared in constant time with hmac.compare_digest, and it fails closed when the token is not configured
DECISION 01

Why RAG and not fine-tuning?

The content changes with every game patch. Retraining a model on every change is expensive and goes stale fast; a curated corpus that gains new documents every day is current by construction, and the LLM only comes in to turn text into JSON and to answer grounded in what was retrieved.

DECISION 02

Why does ingestion run on GitHub Actions instead of inside the API?

A serverless function has an execution time ceiling, and the daily ingestion (scraping, embeddings and LLM curation) does not fit under it. The cron job runs on Actions and only writes to the database; the API keeps reads and chat, which are fast enough.

DECISION 03

Why truncate the embedding to 1024 dimensions?

The provider returns 3072 dimensions by default and pgvector's HNSW index has a practical limit below that. Matryoshka truncation (the dimensions parameter on the call itself) makes the vector fit the index without switching providers or giving up similarity search.

STACK: Python · FastAPI · PostgreSQL · pgvector · Next.js · TypeScript · GitHub Actions · Vercel

#03 · INDEX

More projects

What I built exploring languages, paradigms, and domains — from Prolog to Python, from a compiler to embedded firmware, from games to process automation.

  1. #07STATUS: SHIPPED

    Software Design — Microservices

    Distributed system: Gateway (Java) + User Service (Python) + Connections (Java) + Frontend

    2026 · Java · Spring

  2. #08STATUS: SHIPPED

    CCA center management

    Spring Boot 4 / Java 21 + React 19 for social-education centers: public sign-up, enrollment, daily attendance and three access roles

    2025 · Java 21 · Spring Boot

  3. #09STATUS: SHIPPED

    ML-Copa

    World Cup prediction with XGBoost, adaptive Elo and Dixon-Coles

    2026 · Python · XGBoost

  4. #10STATUS: IN PROGRESS

    CacaOS

    C++ firmware for an ESP32 with a touch screen: 8 LVGL mini-apps and an SDL2 simulator for developing without the board

    2026 · C++ · PlatformIO

  5. #11STATUS: SHIPPED

    Custom language compiler

    Compiler in Java: lexer, recursive-descent parser, tree-walking interpreter and a NASM x86 32-bit assembly generator

    2026 · Java · Assembly x86

  6. #12STATUS: LIVE

    USP-Fono

    Partnership with USP — web application for speech-language pathology

    2026 · JavaScript · React

  7. #13STATUS: SHIPPED

    PredictFlow

    Next.js 15 frontend for a sales pipeline dashboard: CSV import, Chart.js dashboards and overdue-deal alerts

    2025 · Next.js 15 · React 19

  8. #14STATUS: SHIPPED

    Universe

    Full-stack application in Next.js and TypeScript

    2026 · Next.js · TypeScript

  9. #15STATUS: SHIPPED

    Soli

    Full-stack social application (JS + Python)

    2025 · JavaScript · Python

  10. #16STATUS: SHIPPED

    Delivery Tracker

    Real-time delivery tracking

    2025 · Python · JavaScript

  11. #17STATUS: SHIPPED

    Pokédex

    Pokédex in React + TypeScript + Vite

    2026 · React · TypeScript

  12. #18STATUS: SHIPPED

    RFQ Automation

    Request-For-Quote automation

    2024 · Python · Automation

  13. #19STATUS: SHIPPED

    MD-Project

    Discrete Mathematics in Prolog

    2026 · Prolog · Logic Programming

  14. #20STATUS: SHIPPED

    Hardware/Software Systems

    Low-level projects and computer architecture

    2026 · C · Assembly

  15. #21STATUS: SHIPPED

    Pokémon Showdown PS

    System inspired by Pokémon Showdown

    2026 · HTML · JavaScript

  16. #22STATUS: SHIPPED

    Calculus Project

    Calculus solved and visualized in Python

    2025 · Python · Math

#04 · EDUCATION

Insper · Computer Science

B.S. in Computer Science. Five semesters, from discrete math to distributed systems, applied AI, large-scale data, and low-level architecture.

S1

1st Semester

Insper IntroProgramming and logic fundamentals
Core CoursesMath, algorithms and basic data structures
SprintSocial network in Django 5 + PostgreSQL with Google sign-in; comments, likes, content reporting and profile
S2

2nd Semester

Bits and ProcessorsComputer architecture: ALU, datapath and assembly language
PhishingAppSecurity application for phishing detection
Effective ProgrammingOOP, design patterns and code quality
SprintPredictFlow: sales pipeline front end in Next.js over a FastAPI/MongoDB API; JWT authentication and dashboards
S3

3rd Semester

ARQOBJAdvanced object-oriented architecture
Linear AlgebraMathematical foundations for computer graphics and ML
Discrete MathDiscrete mathematics — implemented in Prolog (MD-Project)
Artificial IntelligenceQ-Learning, SARSA and agents for NQueens, Frozen Lake and SPFC
SPRINTCCA student group management in Spring Boot/Java 21 + React; enrollment and waitlist modules
S4

4th Semester

Software DesignMicroservice architecture with Docker (Java Gateway + Python User Service + JS Front)
Machine LearningSupervised modeling, validation and models in production
Languages & ParadigmsFormal study of paradigms beyond imperative and object-oriented
Hardware/Software SystemsOperating system interface, syscalls and low-level programming
SprintReal estate pricing with QuintoAndar: API in FastAPI + Gradient Boosting in production on AWS, CI/CD and 36 tests
S5

5th Semester (current)

Platforms, Microservices and APIsAPI design, service-to-service communication, scalable platforms
AI StartupAI product from scratch: validation, applied LLMs and go-to-market
Big DataLarge-scale data: modeling, pipelines and distributed queries
Algorithm Analysis and Technical InterviewsComplexity, data structures and problem solving under pressure
Games and InteractionGame loops, real-time interaction and user experience

#05 · TOOLING

Technologies I use

Languages, front-end, back-end, ML, and infra — from what I write to where it runs.

LANGUAGES

TypeScript

Python

Java

JavaScript

Prolog

C

C++

FRONT-END

Next.js

React

Three.js

R3F

Tailwind

Framer Motion

GLSL Shaders

Vite

LVGL

BACK-END

Node.js

FastAPI

Spring

WebSocket

REST APIs

PostgreSQL

pgvector

Drizzle ORM

ML

XGBoost

Pandas

NumPy

Scikit-Learn

RAG

Embeddings

INFRA

Docker

Vercel

Turbo

GitHub Actions

Render

#06 · CONTACT · SÃO PAULO, BR · SINCE 2026

Let's build something that holds up in production.

I'm Leonardo Souza, a Computer Science student at Insper (fifth semester), in São Paulo. I work on all three fronts: backend, web, and applied AI. On the backend, an outbox queue with exponential backoff and an idempotency key, rate limiting in Redis with a circuit breaker that fails open, and payment webhooks validated by signature. On the web, a Next.js and TypeScript front-end in a system with real users — including a module that generates a document the end client reads as part of the contract. In AI, an LLM assistant over the authenticated user's data and vector search in Postgres with pgvector, where the answer is grounded in what was retrieved, no fine-tuning. I'd rather measure than assume — and I'd rather have a system that defends itself than one that only works on the happy path.