Advertising Optimization Engine

View Repository Source

A machine-learning engine that decides which ad to show and when — written entirely in NumPy from scratch, with 23 verified mathematical spec checks.

Most machine-learning projects lean on a stack of existing libraries and call it a day. This project is a complete advertising-decision engine (which ad to show, to whom, and when) written entirely in Python using only numpy, working through the mathematics directly rather than inheriting from a library. This serves as the mature realization of my 2018 graduate-fellowship project.

Off-Policy RL from Scratch

I implemented modern off-policy reinforcement learning from scratch. This includes V-trace targets and a Retrace policy update over a safeguarded replay buffer, functioning as the decision layer of a complete optimization engine. Special functions, numerical integration, and optimizers (like Adam and Nelder-Mead) were all implemented in-house.

23/23 Spec Checks Verified

All 23 mathematical claims about the system were individually verified against a formal spec:

  • Analytic gradients matched to finite differences.
  • A maximum-likelihood fit successfully recovered true parameters where a naive fit was measurably biased.
  • Every deliberate deviation from the spec was confirmed by direct enumeration.

Modeling Real User Responses

User responses are modeled using a constrained self-exciting point process fitted by maximum likelihood. This is paired with a two-headed neural network whose loss correctly handles not-yet-observed outcomes and automatically balances the two distinct training objectives.

Weak Identification and Honest Limits

The system predicts individual outcomes well. The dual-head network learned effectively, boosting conversion AUC from 0.55 to 0.95 and achieving a predicted-vs-true timing correlation of 0.93.

However, the branching ratio — the parameter carrying the entire economic claim — came back 46% high on data generated by the exact model being fitted, at a higher likelihood than the truth. I identified this as a weak identification issue, not an optimizer failure, and explicitly published it as a limitation rather than quietly engineering past it.