Florentin Purcea

Academy · Automation · Lesson

Automation with governance: processes that run themselves without keeping you up at night

Least privilege, approval on money and customer-facing actions, idempotency, logging and alerts, rollback, staging, secrets, API drift and a named owner. The model Universal itself runs on.

Level
Advanced
Time
6 min
Updated
Valid for
Automation, 2026

The situation

On Monday morning the reminder automation sent 340 "your invoice is overdue" messages to customers who had already paid. The cause: an export with one column shifted. Nobody saw it, because nobody was looking; there was no log, no alert, no stop button. The team spent the day on the phone apologising.

That is the difference between an automation and an automation with governance. The first saves time until the day it costs it back tenfold. The second has, built in from day one, the answers to: who is allowed, who approves, what happened, how do we stop it, how do we undo it.

Permissions: the least that is enough

Every automation has an account or token it uses to enter your systems. The rule: exactly the rights it needs, nothing more.

  • Invoice reminders read invoices and send messages. They have no right to modify invoices.
  • The lead sync creates contacts. It cannot delete.
  • Every automation has its own service account, not the administrator's. When one is compromised or misbehaves, you know which one and shut it down without stopping the rest.

Review the rights quarterly. Automations accumulate "temporary" permissions that stay.

Approval for money and for what the customer sees

Draw the line clearly: actions that move money (payments, refunds, discounts) and actions the customer sees (messages, emails, order changes) go through an approval step until they have a clean history. The approval can be light — a message with "Approve / Reject" — but it exists.

Thresholds, not absolutes: below an amount you set, automatic; above it, approval. Templated messages approved once run automatically; text generated by a model needs approval. You raise the thresholds as the log shows months without incidents.

When things fail: idempotency, logging, rollback

Idempotency and retries

External systems go down. The network goes down. The automation retries — and if it is not built right, it sends the message twice or creates the order twice.

Idempotency means: the same action run twice has the same effect as run once. In practice:

  • Every action has a unique identifier (order ID + message type). Before sending, it checks whether it already sent.
  • Retries have a limit (3–5) and increasing pauses between them. Without a limit, a system that went down receives thousands of requests when it comes back.
  • What still fails after the limit goes to a "for a human to check" queue, not silently into nothing.

Logging and alerting

Without a log there is no "what happened" — only "I think". The minimum log for every run:

  • when it started, what triggered it, what data it received;
  • every action taken in another system, with its result;
  • how it ended: success, failure, handed to a human.

Alerts are the log coming to you. Three kinds: failures (it broke), anomalies (it sent 340 messages when the average is 12), silence (it has not run at all for 48 hours — the most ignored one). Alerts go to a named person, on a channel they read, not into a "notifications" mailbox nobody opens.

Rollback: how you undo

Before launch, answer in writing: if the automation did something wrong to 500 records, how do you repair them in 30 minutes?

  • Keep the previous state at every change (the old value, in a log).
  • Prefer actions that can be undone: "marked for archiving" instead of deleted, "draft" instead of sent.
  • Have a stop button: one setting that suspends the automation immediately, without asking a developer.

If the answer is "we can't", the action must go through approval, or not be automated.

Before and after production: staging, secrets, drift

Staging before production

Every automation has two environments: a test one, with test data or a copy, and the real one. Any change — a new field, an updated connector — runs in test first, on 20–30 copied real cases. Only then in production. It costs an hour. A mistake in production costs a day.

Secrets

API keys, passwords and tokens do not live in the automation flow, in a spreadsheet, in chat or in code. They live in a secrets vault (the platform's or a dedicated one), are referenced by name, rotated periodically and revoked immediately when someone leaves the team. Every secret has an owner and an expiry date.

Drift: when APIs change

The systems you connect change without asking you: a field is renamed, a connector is deprecated, a limit shrinks. The automation does not "break" visibly — it starts sending empty data or skipping steps. Defences:

  • Validate data on the way in and on the way out: "the phone number has 10 digits", "the amount is positive". What fails stops the run and alerts.
  • A monthly regression test: the 20–30 staging cases run again and compared with the expected result.
  • Subscribe to the change announcements of the providers you depend on.

Ownership: someone is accountable

Every automation has a person's name next to it. That person receives the alerts, reads the log weekly, approves changes and decides when thresholds go up. Without an owner, an automation is an employee nobody supervises and everybody assumes is supervised.

What to remember

  • Every automation has its own account, with the least rights that are enough.
  • Money-moving and customer-facing actions go through approval until they have a clean history.
  • Idempotency and limited retries prevent duplicates; the log and alerts (including on silence) tell you what happened.
  • Rollback written before launch, staging before production, secrets in a vault.
  • A named person owns every automation; without an owner there is no governance.

Check yourself

The automation sent the same message twice after a timeout. What was missing?

Frequently asked

Is this too much for a small company?

It scales. The minimum for any automation, however small: a log, a stop button and a named owner. Approvals and staging are added when the automation touches money or customers.

How do I detect that an automation has stopped running?

With a silence alert: if it has not run within the expected window (say 48 hours), someone gets a message. Noisy failures you see; silence has to be monitored explicitly.

Where do I keep API keys?

In a secrets vault — the automation platform's or a dedicated one — referenced by name, with an owner and an expiry date. Never in the visible flow, in a spreadsheet, in chat or in repository code.

Want Valhalla to check this for your business?

Valhalla Pulse scans the public signals of your website for free, in seconds.

Analyse my business
Analyse my businessWhatsApp