Capability evidence
Assess what happens when the average stops describing the workload.
A credible assessment presents representative distributions, not only a row count and desired technology. It should reveal whether the person can reason about the difficult key, time window, join, state store, source, sink, consumer, failure, and operating constraint without hiding correctness behind successful job completion.
01Workload and architecture judgment
Ask the engineer to profile volume, rate, bursts, object and record size, skew, cardinality, joins, scans, windows, state, concurrency, retention, locality, network, growth, recovery, and cost before comparing optimized single-system, parallel batch, distributed query or storage, and streaming options.
Confirm: The person can identify the actual bottleneck, challenge premature distribution, explain the complexity budget, and select components from measured requirements rather than resume keywords.
02Data, time, and state correctness
Use a scenario with repeated, missing, out-of-order, late, corrected, deleted, and conflicting records to inspect identity, schema evolution, partitions, ordering scope, event and processing time, windows, watermarks, state, offsets, checkpoints, output modes, idempotency, and reconciliation.
Confirm: The person distinguishes transport, processing, storage, and business effects; states every guarantee's source and sink conditions; and preserves unknown, partial, duplicate, corrected, and irreconcilable outcomes.
03Performance and recovery
Review a plan for sustained and burst load, hot keys, large records, shuffle, spill, memory, serialization, garbage collection, backpressure, slow consumers, worker and coordinator loss, partial writes, checkpoint loss, replay, rebalancing, backfill, failover, and source-to-sink reconciliation.
Confirm: The person measures end-to-end latency, throughput, lag, resource use, recovery time, correctness, and unit cost by material partition and workload shape, then narrows or rolls back when the proved envelope fails.
04Operation and evolution
Ask how the person will instrument source, job, partition, state, storage, sink, consumer, quality, lineage, security, privacy, capacity, cost, and recovery signals; handle incidents; change schemas and platforms; and retire old data and components safely.
Confirm: The person connects telemetry to data and business evidence, can stage compatible change, preserves client-held recovery knowledge, and treats scale-down, simplification, migration, and retirement as valid engineering outcomes.