articles · software & hmi
Design the fault state first: a device's worst day, legible to people and agents.
Products are judged in their worst five minutes, not their best. Most interfaces are still designed for the demo. And as more devices are monitored, questioned and operated through software, the fault state increasingly has to explain itself to an assistant as clearly as to the person standing at the device.
skeelx · updated oct 2026 · 4 min read
Open any HMI design file and you'll find the happy path polished to a shine: the dashboard with good data, the workflow that completes, the animation that lands. Somewhere in a corner, if at all, live the screens nobody wanted to design: the sensor dropout, the failed update, the value out of range. Yet those screens are where trust is decided. Nobody remembers the interface from the day everything worked.
The demo-state bias
The bias is structural: demos sell projects, so demo states get design attention. But a device in the field spends real time degraded: connectivity flapping, a sensor aging, an operator doing something the flow never imagined. If those states are afterthoughts, the product's worst day is also its most confusing, which is precisely backwards. The worst day is when the operator most needs the interface to be calm, legible and honest.
Normal, degraded, fault: one model
We design the full state model as one artifact: every state the system can be in, what the operator sees, what they can and cannot do, and how they get back. Degraded states are designed to keep useful work possible. A device that goes dark because one input is stale has chosen drama over service. Fault states are designed to answer three questions instantly: what happened, is anything at risk, what do I do now. If a state can't answer those, it isn't designed yet. The questions don't change when the one asking is an assistant on the operator's phone; only the form of the answer does.
Recovery is the feature
The most under-designed flow in hardware is recovery: resume after interruption, safe restart after power loss, re-pairing without a manual. Recovery designed well converts failures into non-events; designed badly, it converts minor faults into support calls and returns. We treat "back to normal in one obvious path" as a requirement with the same standing as any spec line, and test it on hardware, with the real latency and the real gloves, not in a browser mock. That's the core of our HMI practice, and nowhere does it matter more than in robotics, where the operator meets the machine through its worst five minutes.
The agent-native lens: faults a machine can read
When the buyer sends an assistant. An assistant comparing devices may read the troubleshooting guide as closely as the feature list. A fault table that says what each code means and how to clear it reads as a maker that expects to support what it sells. After the sale, the person asking "why is it flashing amber?" may be asking an assistant rather than the manual, so give it something exact: each state named, what it means, whether anything is at risk and what clears it. On a connected device, expose that status as a callable tool (through MCP's tools/list and tools/call, or a plain API) so the answer comes back as "degraded: sensor stale, logging continues" rather than a bare hex code.
When the business runs on agents. Once the state model is written as data rather than drawn as screens, agents can work it from both ends. In design, they can walk every pairing of state and event and list the ones with no designed response, which is where confusing screens come from. In service, they can watch fleet telemetry for units stuck in a degraded state, group similar fault reports and draft a triage with the likely cause and recovery step. Any action that changes a device in the field (a restart, a rollback, a re-provision) is proposed by the agent, approved by a person and logged with that person's name. Anything touching safety, or needing hands on the hardware, stays a human call from the start.
The worst day, read twice
Increasingly, a device's worst day is read by software before a person sees it: a status call from an assistant, a fleet alert, a support agent drafting the reply. A state model that answers the three questions answers them for every one of those readers, and for the operator at the device. Design the fault state first. The demo will still be fine. The Tuesday afternoon when everything goes sideways is the day your product earns its next order.