Continuity controls
Keep the estate operable and removable without one specialist's memory.
Established clusters accumulate hidden job dependencies, distribution patches, queue conventions, service identities, storage layouts, checkpoint locations, operational workarounds, consumer assumptions, and reasons a workload was never moved. The client record should let another qualified engineer reproduce a run, recover a failure, explain the cost, and continue or end the estate deliberately.
- Client-held estate register
- Workloads, owners, sources, consumers, schemas, dependencies, versions and distributions, code, configuration, libraries, queues, service identities, HDFS paths and policies, checkpoints, tests, releases, incidents, service limits, costs, migrations, exceptions, and retirement decisions remain current in approved client systems.
- Reproducible run and recovery chain
- Approved source fixtures or samples, build and dependency locks, configuration, submission path, query and task evidence, checkpoints, output validation, failure exercise, reconciliation, rollback, and runbook let the client reproduce representative execution and recovery without personal workarounds.
- Least-privilege cluster path
- Individual identities, service principals, keytabs and secrets, data and HDFS permissions, queues, cluster administration, hosts, networks, artifacts, deployments, telemetry, recovery, migration, support, and emergency access are approved, reviewable, and revoked through a client-owned transition path.
- Demonstrated handoff and exit
- A receiving engineer can obtain approved access, identify a workload and its consumers, reproduce a representative run, read plan and resource evidence, diagnose an interruption, reconcile outputs, operate alerts, stage a compatible change, and complete or reverse one migration or retirement step before responsibility changes.