Skip to main content

Infrastructure automation

Make infrastructure change reviewable before it becomes real.

Werkon turns infrastructure intent into a controlled change system. Desired state, discovered live state, the proposed plan, policy evidence, authority, execution identity, protected state, observed outcome, drift, and recovery stay connected so repeatability does not come at the cost of unsafe reach.

Automation contract

Treat code, plan, state, and live resources as different authorities.

Versioned configuration records intended state. A planning engine combines that intent with provider behavior, stored bindings, and observed resources. The apply changes reality. Keeping those boundaries explicit makes review, security, failure handling, and audit materially more useful.

Inputs

Estate and service context
Products and services, criticality, environments, accounts and subscriptions, regions, networks, compute, storage, databases, identity, security services, observability, backup, recovery, external providers, shared resources, dependencies, capacity, cost, regulations, change windows, incidents, and owners.
Current configuration and writers
Console and API changes, scripts, existing code and modules, configuration tools, images, templates, provider and tool versions, state stores, imports, ignored fields, defaults, controllers, service catalogs, scheduled jobs, break-glass work, undocumented resources, drift, and every human or automated writer.
Change and control evidence
Desired-state changes, reviews, plan artifacts, policy and security checks, destructive actions, replacements, cost estimates, dependency and impact analysis, approvals, exceptions, identities, credentials, secrets, locks, concurrency, apply logs, outputs, notifications, audit records, retention, and segregation requirements.
Failure, recovery, and operation
Partial applies, provider limits, timeouts, eventual consistency, state conflicts, failed imports, resource replacement, data-bearing resources, backups, restore tests, prior configuration, rollback and forward-repair procedures, manual intervention, drift handling, maintenance, support, upgrades, deprecation, and operating ownership.

Outputs

Resource and authority boundary
A map of managed, imported, observed, shared, externally controlled, intentionally manual, and excluded resources with source of truth, writers, dependencies, data consequence, change and recovery authority, state boundary, ownership, and known unknowns.
Reviewed desired-state modules
Readable versioned configuration with pinned tool and provider constraints, bounded modules, explicit inputs and outputs, secure defaults, validation, ownership, documentation, examples, compatibility and upgrade policy, deprecation path, and tests appropriate to the resource consequence.
Plan-to-apply control path
Automation that discovers live resources, produces a retained plan, highlights creation, mutation, replacement, deletion, security and cost consequence, validates policy, obtains risk-appropriate authority, applies the exact eligible plan with bounded identity and locking, and records outputs without exposing secrets.
Drift and recovery operating record
Scheduled or event-driven comparison, classified drift, named disposition, emergency-change reconciliation, current state protection, backup and restore evidence, partial-failure procedures, rollback or forward-repair limits, platform health, change outcomes, maintenance work, exceptions, and review cadence.

Automation path

Codify one consequential resource slice before scaling the pattern.

A bounded slice reveals import risk, provider behavior, state sensitivity, approval needs, partial failure, and operator usability without granting an immature automation path control over an entire estate.

  1. 01

    Observe real infrastructure work

    Trace a normal, emergency, failed, and out-of-band change through request, discovery, design, implementation, review, access, execution, verification, service impact, incident, recovery, documentation, audit, and later maintenance with every human and automated writer represented.

  2. 02

    Set the first authority boundary

    Choose a coherent resource and dependency slice; name desired-state, live-resource, state, secret, policy, and ownership boundaries; decide what is imported, recreated, observed, shared, excluded, or left manual; and document destructive and data-bearing consequences before code takes control.

  3. 03

    Build reviewable plan evidence

    Version configuration, pin relevant tooling, validate structure and policy, discover current resources, generate the change plan, expose create, mutate, replace, and delete actions, resolve unknown and sensitive values safely, assess service and cost impact, and route the exact plan to accountable authority.

  4. 04

    Apply with bounded authority

    Use a protected automation identity with only required scope, coordinate writers through state locking or an equivalent control, apply the eligible plan, capture provider results and outputs, stop on uncertain state, verify resource and service behavior, and retain enough evidence for reconciliation and audit.

  5. 05

    Detect drift and prove recovery

    Compare desired, stored, and live state on a defined cadence; classify expected, emergency, malicious, provider-generated, and accidental differences; reconcile through review; test state and resource recovery; measure failures and manual work; and expand automation only when the operating contract remains usable.

Automation decision

Choose adoption, reconciliation, or redesign from the actual resource state.

Existing infrastructure rarely starts clean. The safe path depends on resource criticality, import fidelity, hidden writers, data consequence, provider semantics, and whether the current shape should be preserved at all.

01The boundary is new and its dependencies can be made explicit

Codify new infrastructure

Create the desired-state, module, plan, policy, apply, evidence, drift, and recovery pattern before production resources exist. Keep the first slice small enough to inspect, delete, and recreate safely while establishing conventions from demonstrated operator tasks.

Evidence: Service need, resource and dependency model, data classification, options, desired configuration, validation, plan, policy result, identity scope, cost range, failure injection, verification, deletion and recovery test, documentation, owner, and review date.

02Existing resources are valid but management is manual or fragmented

Import and adopt

Inventory actual resources and writers, capture a recovery point, import stable identities into a bounded state, write configuration that matches intended reality, explain every proposed difference, suppress only provider-generated or deliberately external fields, and transfer change authority gradually.

Evidence: Live inventory, resource identifiers, current configuration, dependencies, writers, data and service consequence, backup, import mapping, state result, zero or explained plan, ignored-field rationale, review, verification, manual-path retirement, and owner acceptance.

03Desired, stored, and live state disagree after a known or unknown change

Reconcile drift

Determine which authority should win before applying anything. Preserve a valid emergency or provider change in reviewed configuration, revert an unauthorized change through the normal path, repair stale state carefully, or accept a deliberate exception with scope, owner, evidence, and expiry.

Evidence: Desired configuration, plan, stored bindings, live observation, writer and timestamp, change record, service and security consequence, emergency context, source authority, selected disposition, approval, apply or state action, verification, exception expiry, and recurrence prevention.

04The current resource shape is too coupled, privileged, or irreversible

Redesign before automation

Change account, network, identity, module, state, data, dependency, or recovery boundaries before expanding automation. Introduce the target through reversible slices, migrate authority and state deliberately, reconcile outputs, and retire the old path only after production evidence supports it.

Evidence: Current failure and privilege boundary, coupling, state blast radius, data authority, recovery unit, target options, migration seam, compatibility, dual-running limits, security review, tests, observability, cutover, restoration evidence, residual risk, cleanup, and ownership.

Automation controls

Reproducibility does not excuse excessive authority or blind convergence.

Infrastructure automation can create, mutate, and destroy production foundations at machine speed. These controls keep that power tied to a reviewed consequence, current coordination state, and recoverable operating decision.

Approve the plan, not just the code
Review configuration for intent and the generated plan for consequence. Make creation, replacement, deletion, permission, network, data, availability, cost, and unknown effects legible. Bind approval to the exact eligible plan, invalidate it when inputs or live state materially change, and preserve a separate emergency authority path.
Treat state as sensitive production data
Protect state confidentiality, integrity, availability, version history, access, and recovery according to the attributes and bindings it contains. Coordinate writers with locking or equivalent serialization, keep local copies out of normal workflows, never print secret values as evidence, and restrict force-unlock or state surgery to diagnosed cases with records.
Bound identities and writers
Separate plan, apply, read-only discovery, drift, and emergency roles where consequence requires it. Prefer short-lived credentials, limit account, workspace, resource, action, network, and time scope, inventory every controller and manual writer, log consequential actions, and remove obsolete access when authority moves to automation.
Reconcile before converging
Do not automatically overwrite unexplained drift. Identify the writer, reason, service effect, source authority, and data consequence first. Route the result to adopt, revert, state repair, exception, incident, or redesign, then verify live behavior and update desired state, procedures, and prevention controls.

Engagement fit

Use infrastructure automation when real resources and change authority can be bounded.

Good reason to begin

  • A coherent resource slice has identifiable service purpose, configuration, dependencies, environments, writers, state, data consequence, controls, failure modes, recovery needs, and accountable owners.
  • Platform, cloud, network, security, data, operations, support, risk, finance, and application participants can show actual work, explain policy intent, review plans, test outcomes, and maintain the resulting automation.
  • Current resources, scripts, code, state, plans, changes, identities, permissions, provider behavior, drift, incidents, backups, restores, manual work, timings, costs, and audit needs can be inspected without exposing secret values.
  • The organization can protect state and credentials, fund module and provider maintenance, resolve out-of-band writers, test recovery, redesign unsafe boundaries, support automation users, retire obsolete paths, and review exceptions and ownership over time.

Resolve before beginning

  • The desired answer is fixed as a full estate rewrite, a named tool, automatic remediation of every difference, or elimination of qualified approval regardless of resource consequence.
  • Automation is expected to copy an undocumented manual process, conceal destructive actions, broaden standing privilege, store sensitive state carelessly, or overwrite emergency and provider changes without reconciliation.
  • Resource inventory, dependencies, live access, state, identity, data, incident, backup, restore, or key operator knowledge is unavailable enough that adoption could mutate or destroy infrastructure through guesswork.
  • No accountable owner can approve the resource boundary, import, desired state, plan, policy, identity scope, destructive action, exception, drift disposition, residual risk, recovery, maintenance, or retirement decision.

Source basis

Sources behind the control model.

  • 01

    DORA

    Flexible infrastructure

    Current DORA guidance connects flexible infrastructure with on-demand capability and continuous delivery, describes version-controlled infrastructure changes and approval as a way to support control, and recommends starting small because adoption requires engineering and process change.

  • 02

    National Institute of Standards and Technology

    Guide for Security-Focused Configuration Management of Information Systems

    Final NIST Special Publication 800-128 treats security as integral to configuration management and covers management and monitoring of system configurations to support required functionality while reducing organizational risk.

  • 03

    OpenTofu

    Provisioning infrastructure with OpenTofu

    Current OpenTofu documentation distinguishes desired configuration, observed real resources, state bindings, a proposed plan, a saved plan artifact, controlled apply, and destroy operations, including application of an exact pre-approved plan.

  • 04

    OpenTofu

    Sensitive data in state and state locking

    Current OpenTofu documentation states that infrastructure state can contain sensitive resource attributes and should be protected accordingly; its related locking contract serializes supported write operations to reduce conflicting writers and state corruption.

[ WORKFLOW / SYSTEMS AUDIT ]
THE FIRST ENGAGEMENT

Start with one real workflow

A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.

Show Us the WorkflowStart with the free automation readiness checklist

OBSERVEQUANTIFYDECIDEBUILD