About Me
I’m Rafael Kovashikawa, an economist-turned-data-professional with a soft spot for macro signals, efficient pipelines and making cool things with LLMs.
Always curious about tech, data and dogs.
I will use this space to post about things that I’ve found interesting along the journey and, of course, my projects.
What I do best
- Data Science & Modeling · Develop and productionize time‑series (ARIMA, VaR), tree‑based and econometric models for CPI forecasts, credit risk and trading signals.
- Data Engineering · Build and orchestrate ETL on Airflow + Kubernetes.
- ML in Production · Deploy credit-risk and anomaly-detection models (FastAPI, Docker).
- LLMs · RAG pipeline, LangChain + vector DBs.
Tech Stack
Programming: Python, R, MongoDB, SQL, JavaScript, TypeScript, CSS, HTML, MCP (Model Context Protocol), Pydantic AI, Django
Data Viz & Analysis: Plotly, Matplotlib, Tableau, Excel, Bloomberg Terminal
Cloud & Automation: AWS, Redis (Pub/Sub), Apache Airflow, GitHub Actions (CI/CD), Google Cloud Platform
ML & AI: scikit‑learn, XGBoost, LightGBM, Econometrics (Statsmodels), LLMs, NLP, MCP Servers, Context Engineering, AI Agent Development
Languages: Native Portuguese; fluent English and Spanish
Professional Experience
Fuse — AI Data Strategy for Marketers — Data Scientist & AI Engineer
Jan 2026 – Present
- Architected the enterprise Salesforce integration as an outbound hosted-MCP client: OAuth 2.0 + PKCE authorization flow, cross-tenant refresh-token assertions, and an untrusted-content envelope guarding against prompt injection from CRM record text. Shipped as a distributable 2GP managed package.
- Built 6 production MCP servers (Klaviyo, SEO, Salesforce, TikTok Ads, LinkedIn Ads, Exa) exposing 30+ analytics tools to LLM agents.
- Designed a multi-agent ReAct orchestration layer for three research agents: an 8-action loop gated by a complexity classifier, adaptive retry across 9 error classes, and per-datasource circuit breakers.
- Built a sandboxed executor for LLM-generated analysis code that intercepts hallucinated imports and redirects agents to the real API, so they self-repair instead of failing.
- Engineered context system across 16 agents (per-API field references, deterministic tool-routing evals) to resolve vocabulary collisions across 12+ MCP servers and cut hallucinated queries.
- Instrumented OpenTelemetry tracing and per-tool credit metering, turning LLM spend into a billable, auditable unit.
GYRA+ — Head of Data
Aug 2024 – Jan 2025 · Brazil
- Led the strategic re‑architecture of data models and ETL/ELT pipelines (Python, Airflow, AWS), cutting AWS costs for the data team by 50 %.
- Built a high‑performance simulation API (FastAPI, Polars, MongoDB) deployed on Kubernetes, enabling rapid credit‑portfolio stress testing.
- Delivered a production‑ready RAG pipeline (LangChain, MongoDB) that extracts and structures key information from legal documents.
JGP Asset Management — Macro Data Scientist
Jan 2021 – Jul 2024 · Brazil
- Designed LATAM CPI forecasting models and global macro dashboards that out‑sped Bloomberg.
- Engineered resilient pipelines (Python, SQL, Airflow) to ingest and curate macro datasets for trading research.
- Developed NLP models to score Brazilian Central Bank communications and integrate sentiment into trading strategies.
- Automated data workflows, eliminating manual steps and improving data timeliness.
Órama (acquired by BTG Pactual)— Data Analytics
Sep 2019 – Aug 2020
- Mined client data with Python and SQL, applying NLP and k‑means clustering to improve investment recommendations.
IBM — Data Operations
Apr 2019 – Aug 2019
- Automated ticket‑management workflows with Python, Selenium, and SQL, eliminating +90% of manual effort
Education
MIT — Data Science & Statistics MicroMasters (edX)
May 2024 – Feb 2026
Fundação Getúlio Vargas (EPGE) — B.Sc. Economics
2017 – 2021
Awards
| Award | What I Built |
|---|---|
| 🏆 Google Cloud AI Hackathon (Dec 2025) | Built MedAnnotator, a multimodal medical annotation system. [Post] [Code] |
| 🏆 Big Data Hackathon (Sep 2021) | XP Inc. & Microsoft Azure — built a 14 GB big‑data application on Azure. |
| 🏆 Cryptocurrency Datathon (Sep 2020) | FGV & Ripple — ML strategies for Bitcoin trading with NLP on news and Reddit. |
Outside the terminal
I was born and raised in Rio de Janeiro, Brazil 🇧🇷
Now I live in Washington, DC 🇺🇸
My favorite things are:
🦮 dogs · 🏃 running · ⚽ Fluminense fan
Need a hand with AI and data problems?
Hit me on LinkedIn or open an issue and let’s chat.