Discussion about this post

User's avatar
Ariel Sama's avatar

The gate vs. gauge distinction is the clearest frame I've seen for what makes AI quality work fundamentally different from traditional software QA. The point about gating on the tail rather than the average is particularly sharp. That's exactly where the Air Canada failure lives: the average response was probably fine.

The pharmaceutical release lab analogy landed for me because it reframes the question. It's not 'do we need testers again', it's 'we already have the apparatus, we just haven't pointed it at language models yet.'

This connects directly to what I'm navigating with my own trading system. The deterministic Python layer has clean gates. The path toward agentic execution is where the gauge problem starts, and your framing of calibrating the judge before trusting the scores is the piece I hadn't fully worked out.

2 more comments...

No posts

Ready for more?