trolls.dev

The Night Shift · Department

code-results

Notes from building code-results — tying every deploy to the business metrics that moved after it.

7 essays filed here

  1. 01

    You ship a change, and conversion climbs nine percent. Then someone splits the number by cohort and finds the truth: new users converted worse, returning users converted worse, every segment converted worse. The total rose anyway — because the release quietly changed who showed up, not how well the product worked. It's Simpson's paradox, and it's the trap a single headline metric is built to walk you into. code-results reads the number the only way that can't lie to you: split, and weighted by the mix that actually moved.

    6 min read
  2. 02

    Ship one thing on a quiet afternoon and the number that moves after it is yours to read. Ship two things twenty minutes apart and the lift has no owner — it belongs to both changes and neither, and any clean split you draw between them is a number you made up. The honest move isn't to guess which deploy earned the win. It's to notice that, for now, the question can't be answered — and to say so out loud instead of picking a winner.

    6 min read
  3. 03

    Ship a change, watch the metric, and sometimes the line just… doesn't move. Most tools treat that as a non-event — no green arrow, nothing to celebrate, nothing to file. But a confirmed 'no effect' is a real result, and burying it is how a team pays twice: once to build the thing, and again to relearn that it did nothing. code-results stamps the flat line as a finding, not a gap.

    6 min read
  4. 04

    A change is a trade, not a gift. You ship to move one metric, you watch that metric, it moves, and you call it a win — while the number you never instrumented quietly moves the other way. The bill goes unread because it was never on the screen. code-results reads every deploy against the whole handful a business would notice, so the win and the cost arrive together, on the same board.

    6 min read
  5. 05

    The answer to 'did it work?' isn't worth the same on Monday as it is three weeks later. Not because the number changes — because you do. The reason you shipped, the alternatives you weighed, the file you were nervous about: all of it decays out of memory on a curve. A verdict that arrives after the why is gone is a fact you can't act on. code-results is built to catch the answer while the reason is still lit.

    6 min read
  6. 06

    Sign-ups are up since you shipped. That feels like a win — but 'up' is a level, and a level is the one thing that can't tell you whether your change did anything. The real question is whether the line broke where your deploy landed, and that's a question of alignment, not analytics.

    5 min read
  7. 07

    Most teams can tell you exactly what they shipped last quarter and almost nothing about whether it mattered. The feedback loop dies at the merge button — and that's the loop code-results is built to close.

    5 min read