The Normalization of Inexplicable Failures
218 points - today at 3:26 PM
SourceComments
I am also big on testing (the correct things). And nine-nines (big on Elixir).
And... I'm also big on agent-assisted dev. Which requires pretty much every check in the book to stay productive in. And that's fine to me. I've seen bugs that I wouldn't have made myself. And I've also seen my own bugs fixed. They've all gotten fixed in short order. I don't see why this is a problem.
Raise your personal standards.
Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.
That may be tolerable for some user-facing app. But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows EVERYTHING and EVERYONE down.
It’s also tightly connected to a normalization of lack of accountability.
> This isn't "getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" -- you still have to do the hard part.
This is probably losing the younger portion of the audience by now. ;)
> For many users, however, the actual experience is roughly just "stupid thing sucks." Software already feels capricious; more failures just change the rate of frustration.
I am betting author does not use cloud services much. It is not just "users", it's developers as well. Github is returning 5xx? AWS service does not work? Your email did not get delivered? Nothing we (developers) can do, "stupid thing sucks".
I bought a new electric car recently. For the most part I've been quite happy with it. Shortly after I bought it, it started popping up a warning message saying "check EV system" every time I started it. By the time I brought it into the dealership, the warning had gone away, and the technician just told me something to the effect of "eh, I guess it just does that sometimes, let us know if it happens again." Hardware fault? Software bug? Who can say?
Like most modern cars, it has connectivity and Google Maps built into the infotainment system. The vast majority of the time, it works fine. Sometimes it says it has no connectivity (meaning no traffic data and suboptimal routes) for the duration of a drive, even in areas with a strong cell signal where it normally works fine. Sometimes the car says it has connectivity, but Google Maps still thinks it's offline. Sometimes Maps will actually load and display a route, but the "start navigation" button just spins forever as though it's still waiting for something. Are these related issues? Is there a common cause that might be fixable? Who can say?
(Conveniently enough, the warranty specifically does not cover any failures of software or firmware to operate correctly.)
Not sure why anyone would feel the need to dunk on this thing without showing a real failure example.
I agree with that Evals are a scarce commodity rn. A business needs to define what good looks like. This is a laborious, and sometimes politically controversial, process.
Given good Evals, frontier llms are a magical tool that can automate tasks and do them more accurately than human ops teams.
Exactly which llm to use depends on the mixture of speed, cost and quality of the output.
Jev makes a claim to expand some regions of the Pareto frontier. I look forward to testing if this is true.
There are many areas of work we can’t automate rn. We cannot create good Evals either because time horizons are too long, or it’s too difficult to create good Evals.
That doesn’t mean there’s anything wrong with building good ai engineering systems in areas where it works magically.
Wasting your time and resources is a power signifier, but getting you to waste your own time and your own resources is hegemony.
"Normal accidents, or system accidents, are... inevitable in extremely complex systems. Given the characteristic of the system involved, multiple failures that interact with each other will occur, despite efforts to avoid them... while operator error is a very common problem, many failures relate to organizations rather than technology, and major accidents almost always have very small beginnings. Such events appear trivial to begin with before unpredictably cascading through the system to create a large event with severe consequences." [1]
If you are pitching that your service can do potentially "whatever the client wants" you have such a thin basis on which to provide contracts and guarantees as a provider. The narrower the function, the clearer you can be about what's supposed to happen and why things might have gone wrong.
When you're using probabilities as the fundamental approach to computation, all of that goes out the window. Nondeterminism is powerful because it's insanely flexible, but the cost of that flexibility is predictability and expectation. Determinism was humanity's primary choice for formalisms and technology precisely because it reduces complex problems and situations to repeatable mechanics that are easy to understand. Deterministic tools can't do a lot in the grand scheme of things, but it is precisely these limitations that make them work well in concert and keep them comprehensible.
Lovely post, and lovely phrase (normalizing of inexplicable failures).
Excellent post!
Literally go no further than the age old advice of "have you tried turning it off and then back on again?", and then that actually working.
Did this person never experience the effects of rocking the boat just a little too much? Daring to do a little too good of a job? How?
(Yes, storing gold on a what the room door claims is a toilet is a best practice, as thieves won't look for it there. But now the security through obscurity got leaked /s )