Most of what I learn comes from the gap between what a system was supposed to do and what it measurably did. These are the ones worth writing down: an eval that said the opposite of the received wisdom, a bug that only appeared once the thing was deployed, a number that turned out to be measuring the wrong clock.