Owned the statistics on the company's experiment platform, and the models behind personalized offers.
I owned the statistics behind the A/B tests at a 27-market car-parts retailer, and the models that decided which offers a shopper saw.
Release decisions rested on numbers I was accountable for, and the suggestions shoppers saw came from models I designed.
AUTODOC · full-time · autodoc.de
Context
AUTODOC is Europe's largest online car-parts retailer, selling across 27 markets. I owned the statistics end to end on the company's experiment platform, and the models behind the offers shoppers saw. Two years, 27 markets, one platform.
The problem
Two jobs at once: make the company's tests trustworthy, and make the suggestions fit the car in front of the shopper.
What I did
- Experiments — agreed the effect worth detecting before anyone built the feature, on a fixed error budget, and capped how long a test could run, because checking early has a cost
- Experiments — replaced averages-based maths with resampling, because money data has long tails and a simple test on raw averages misleads
- Experiments — put every metric under a versioned definition, so two teams could not quietly measure the same thing differently
- Experiments — kept a permanently frozen group of customers, to measure the effect of everything shipped rather than one test at a time
- Personalization — built the models behind personalized offers
- Personalization — made cohorts reproduce exactly, so a result can be re-checked months later
Results
- Test results that held up when re-checked
- Test results teams could defend to leadership without a statistician in the room
- One shared definition for each metric
- Suggestions chosen by model
- Cohorts that reproduce exactly on re-run
Skills used
A/B testing, CUPED, Sequential testing, Python, SQL
Next
Next case: Scaleo. If you have a similar problem, see how I work with clients, read the about page, or get in touch.