← Lab / 003 · Automation · · 2 min read

Reliable Automation Around Human Workflows

Automation should support human decision-making instead of blindly replacing it. Inputs, confidence, failure states, idempotency, logging and approval.

  • automation
  • reliability
  • observability
  • security-operations

The most useful automation I have seen does not try to remove people from a workflow. It removes the repetitive parts and puts the decision, with its evidence, in front of the person who should make it.

This matters most where the cost of a wrong action is real, which includes security operations. These are the properties I look for when I design automation around human work.

  1. 01PROPOSEgather inputs, form a recommendation
  2. 02VERIFYvalidate, attach confidence and evidence
  3. 03APPROVEa person decides when it matters
  4. 04ACTdo the work, safely and repeatably
  5. 05RECORDlog what happened and why

Inputs come first

Automation is only as sound as what it is fed. Inputs can be missing, malformed, stale or simply wrong. Validating them at the boundary, and refusing to continue on bad input, is far cheaper than debugging a confident but mistaken action afterwards.

Confidence is information

A result is rarely simply true or false. Passing along how sure the system is, and why, lets a person weigh it properly. An automation that always sounds certain trains people to either trust it blindly or ignore it.

Automation that hides its uncertainty has only moved the risk somewhere harder to see.

Design the failure states

Success is easy to build. What matters is what the workflow does when a dependency is unavailable, a response is unexpected, or an input is half-complete. Good failure states are explicit: stop, say what failed, leave the system in a known condition, and make the next step obvious.

Idempotency and retries

Things fail, and automation retries. A retry is only safe if repeating an action does not change the outcome beyond the first time. Designing operations to be idempotent, so doing them twice equals doing them once, turns retries from a risk into a tool.

Retries also need limits. Unbounded retrying can hide a real problem or amplify one. A bounded retry with a clear final failure is more honest.

Logging and observability

If a workflow acted, someone should be able to see what it did, what it saw and why it decided. Logging is the minimum. Observability goes further: being able to ask new questions about behaviour without changing the code.

Human approval, placed deliberately

Approval steps are not a sign of weak automation. They are a design choice about where judgement belongs. The skill is placing them where they add safety without becoming a rubber stamp. If a person approves everything without reading it, the control has failed.

Putting it together

Taken together, these properties describe automation that earns trust: validated inputs, visible confidence, defined failure behaviour, safe retries, a complete record and a human in the right place.

For security operations this is the difference between tooling that helps an analyst investigate faster and tooling that creates new things to investigate. For engineering it is simply good system design, applied to a workflow that has people inside it.