PT

Nº 03of 22Featured

QuintoAndar Pricing

Home prices served by a FastAPI API on AWS: 9.32% MAPE against actual 2010 sales, every request logged, continuous deploy with a healthcheck

Year
2026
Status
SHIPPED
Role
API, tests and validation · team of 4
Stack
Python · FastAPI · scikit-learn · Pydantic +4
quintoandar-precificacao
QuintoAndar Pricing — Public calculator — 8 fields and the price the API estimates
4 screens

1 of 4

QuintoAndar Pricing — Public calculator — 8 fields and the price the API estimates

Public calculator — 8 fields and the price the API estimates

01The project

Fourth-semester Insper sprint with QuintoAndar as the partner: a team of four had to answer what a home sells for and keep that answer running in production. The dataset is Ames Housing — 1,285 residential properties from 2006 to 2009, 80 features — and the final test ran against 175 actual 2010 sales, a temporal holdout whose ids are disjoint from the training set. Two models came out of it: a 4-feature linear regression to explain price to agents and owners (11.9% MAPE, R² 0.832) and a Gradient Boosting model for the public calculator. Of the 52 commits in the application repository, 20 are mine — the largest share on the team: the FastAPI API, prediction logging, the 36 tests, the validation script and report against the 2010 data, and the minimal 8-field model.

02What was built

  1. Gradient Boosting picked among 5 models in chronological validation: $16,117 MAE and 9.83% MAPE on the 2009 holdout, against $57,077 and 34.69% for the median baseline

  2. Final validation against the 175 actual 2010 sales: $15,810 MAE, 9.32% MAPE, 26,071 RMSE and R² 0.894, with 92% of the properties within 20% error

  3. Public calculator cut from 12 fields to 8 by forward selection, at the knee of the error curve

  4. FastAPI 0.115 API with /predict, /predict/simples, /metrics and /health, Pydantic contracts, and the preprocessing and features ported over from the model repo so they match training exactly

  5. Every request written to SQLite — 422s included — with latency, model version, input and output; the real-usage report closed at 300 calls with a p95 of 11.7 ms and 94.3% success

  6. 36 pytest tests across 8 files, and CI/CD that runs the suite on every push and, on merge to main, does SSH + rsync and docker compose up --build --wait on EC2, with a healthcheck after the deploy

03Decisions

Decision 01

Question: Why a chronological split and not a random one?

Answer: Home prices move over time. A random split puts a 2009 sale in training and a 2007 sale in test, and the model ends up evaluated while knowing the future. The cut is by date, validation was the year 2009, and the final test ran against the 175 sales from 2010 — a year training never saw, with disjoint ids.

Decision 02

Question: Why log the requests that fail with 422 as well?

Answer: Logging only successful predictions measures the model, not the service. Recording the 422s alongside them is what shows which contract the client is getting wrong and how often — that is where the 94.3% success rate over the report's 300 calls comes from, instead of an impression that everything was fine.

Decision 03

Question: Why not retrain after the 2010 validation?

Answer: The model trained through 2009 came in at 9.32% MAPE on a year it had never seen, against 9.83% on the 2009 holdout — so there was no degradation to fix. Retraining with no sign of degradation swaps a measured model for a new, unmeasured one; the decision to leave it alone is written down in the validation report.

04Screens

  • comparativo-mae.png

    MAE of the 5 models on the 2009 chronological holdout

  • previsto-vs-real.png

    Predicted vs. actual for the Gradient Boosting model

  • latencia.png

    Latency per call in production (p95 11.7 ms)

1 of 4

QuintoAndar Pricing — Public calculator — 8 fields and the price the API estimates

Public calculator — 8 fields and the price the API estimates

05Stack

  • Python
  • FastAPI
  • scikit-learn
  • Pydantic
  • SQLite
  • Docker
  • GitHub Actions
  • AWS EC2
Next project →Nº 04

TECNA's internal super app on Next.js 16: four new systems of mine — Quality with an offline field screen, Planning with critical path, Safety and Project Management —, the homeowner's manual and the website's CMS