Capability evidence
Assess one risky change from automation choice to a diagnosable CI result.
A framework demo is weak hiring evidence. Use a feature with disputed rules, several layers, one unstable dependency, shared test data, an intermittent failure and a release request that exposes whether the person can build trust rather than just scripts.
01Automation value, risk, oracle, and layer placement
Provide a product risk, existing manual tests and a request to automate everything. Ask the person to select the feedback that deserves code and explain where it belongs.
Confirm: The person starts with the user or system outcome, failure consequence, expected behavior and source authority; distinguishes a test objective from steps and a check from broader investigation; identifies preconditions, inputs, states, invariants, outputs, side effects and a usable oracle; asks whether static review or better product design can prevent the fault first; compares unit, component, contract, integration, service, browser and device layers by fault-detection value, fidelity, isolation, speed, diagnostic precision and cost; keeps many fast deterministic checks near the changed logic; uses contracts for owned interface expectations rather than schema appearance alone; reserves browsers and devices for user-visible behavior and integration that only those boundaries can reveal; avoids duplicating the same claim at every level; names manual, exploratory, usability, accessibility, security, performance and acceptance questions that should not be reduced to this suite; and defines useful coverage and exclusions without promising exhaustive testing or a target automation percentage.
02Automation architecture, testability, and maintainable implementation
Provide duplicated scripts, fragile selectors, hard-coded data and an application with hidden state. Ask for an architecture that preserves readable test intent and local ownership.
Confirm: The person treats test code as an engineered product; separates domain intent, orchestration, boundary adapters, data builders and reporting only where the distinction reduces real duplication; uses typed or explicit contracts; keeps assertions close to the behavior they verify; builds narrow helpers instead of a second application framework; chooses user-visible roles and labels for interface checks where they express the contract; avoids selectors coupled to styling and layout; relies on observable readiness instead of arbitrary sleeps; introduces stable identifiers only when no meaningful user-facing contract exists; makes clocks, randomness, queues, retries and asynchronous completion controllable; uses test doubles deliberately and records the fidelity they remove; verifies important consumer and provider expectations against real implementations; designs product seams and telemetry with developers; versions source, dependencies and configuration; reviews security of credentials and artifacts; tests the automation infrastructure itself; and writes failure messages that identify the violated claim rather than the helper that happened to throw.
03Data, environment, execution, observability, and failure diagnosis
Provide parallel workers, shared accounts, an intermittent timeout and a passing retry. Ask the person to make the original result reproducible before changing the threshold.
Confirm: The person identifies repository revision, build, package, feature flags, configuration, browser or device, runner image, service versions and time; creates the minimum state a check needs; isolates accounts and data per test or worker; uses synthetic or approved masked data rather than copying production by default; owns setup, cleanup, expiration and failure residue; controls network and dependency behavior only when the test objective permits it; gives integrated paths realistic services where fidelity matters; makes ordering assumptions explicit; runs independent checks in clean contexts; balances parallelism against shared-resource limits; preserves attempt, assertion, request, response, trace, structured log, metric, screenshot, video or state delta according to diagnostic value; correlates runner activity with system telemetry; distinguishes product defect, wrong oracle, check defect, data collision, environment drift, dependency outage and infrastructure fault; reproduces the smallest failing condition; measures intermittent frequency; avoids masking evidence with unbounded waits or retries; and quarantines only with an owner, reason, visibility, repair target and expiry.
04CI feedback, reporting, suite governance, and continuity
Provide a slow pipeline, skipped checks, old quarantines and pressure to call the release safe. Ask for a feedback and maintenance policy that developers can operate.
Confirm: The person maps fast checks to local and pull-request feedback, selects slower integrated or device work by change risk and schedule, and verifies the same authoritative build that can progress; fails explicitly when required infrastructure or evidence is unavailable instead of reporting a pass; separates passed, failed, skipped, blocked, timed-out, flaky, quarantined, not-run and unavailable results; exposes first failure and retry history; gives developers a short route from report to responsible code and system evidence; keeps branch protection and gate policy owned and reviewable; monitors duration, queue time, failure yield, false alarms, flake rate, diagnosis time, maintenance effort and escaped failures without turning one metric into a target; links important corrected defects to focused regression checks; reviews dependencies, selectors, contracts, test data and supported configurations when the product changes; deletes redundant or obsolete checks with evidence; preserves historical decisions; stores automation and runbooks in client custody; demonstrates another engineer can execute, diagnose and extend the suite; and gives bounded evidence to release owners rather than declaring quality or readiness.