Data engineering and automation · São Paulo, Brazil
I work with public data — budgets, economic indicators, municipal records. The kind of data that is technically open and practically unusable until somebody sits down and makes it trustworthy.
That's the whole job, as far as I'm concerned: raw data in, reliable interfaces out. Everything below is some version of that.
At the São Paulo State Government, in Budget Planning Management, I built and applied the methodology that monitored public spending and flagged financial anomalies. Oracle, SQL, dbt, and a quantity of XML that I did not choose and have made peace with.
Now at FGV IBRE, I build internal platforms that take economic data analysis from a manual process to one that runs on its own and tells you when it didn't.
Automacao-de-Relatorios — A reporting process that lived in manual steps, rebuilt as a pipeline: multiple Excel sources reconciled, business rules applied, charts generated, everything rendered into a Word document through a Streamlit interface. There's an optional AI layer that writes the prose and is not allowed anywhere near the arithmetic.
sp-gov-budget-analytics — The dbt framework behind the public spending work, with a runnable demo that puts a language model on top of a real government pipeline using nothing but context engineering. No fine-tuning, no magic.
ETL-PROJECT-PYTHON-GCP — An ETL pipeline for Python and BigQuery, built so that Brazilian government exports from different agencies stop disagreeing with each other. Third version. The one I'd hand to someone else.
The rest of my repositories are a working library rather than a gallery. Some of them are old, and it shows.
In 2026 I spent time in Belgium and the Netherlands and talked to developers there. The gap I came home with wasn't about tooling — it was about discipline. Tests that exist. CI that runs. Documentation that describes the code that's actually in the repository.
So that's what I'm building now, and I'm doing it in public, including the parts where my older work doesn't measure up yet. I'd rather fix that where people can see it than quietly delete the evidence.
I'm also moving toward backend, with PHP and Drupal. Partly because structured content and reliable interfaces are the same problem I already solve from the data side, just approached from the other end. Partly because Drupal was born in Antwerp and now carries a good share of public-sector web in Europe and Canada — and public-sector data is where I already know my way around.
J'apprends aussi le français :) ... C'est sympa de regarder des vlogs en français, surtout depuis mon voyage.
Working with: Python · SQL · Oracle · dbt · BigQuery · pandas · R · Streamlit · Docker
Learning properly, one at a time: PHP · Drupal · Java
I've left things off this list on purpose. I'd rather be asked about four things I can defend than twelve I can name.
🎤 AI Tinkerers SP + Banco BMG, March 2026 — From Oracle to Insight: How I Built an AI-Augmented Budget Intelligence Pipeline for the São Paulo State Government
✍️ I write about data engineering, public data and the unglamorous parts on Medium
💬 LinkedIn — open to talking with people working on public-sector data or Drupal, in Europe or Canada especially
Made with coffee and data ☕


