Skip to main content

Hire performance testers

A fast average can hide a system at its limit.

Performance testers investigate whether a service can complete the required user journeys under a realistic demand model. Werkon would assess one reproducible capacity question using controlled releases, data, dependencies and client conditions. Useful evidence includes failed work, latency distributions, resource use and recovery after stress. Product and service owners retain performance targets, production-risk, cost and release authority.

Responsibility contract

Give performance evidence an owner without letting a benchmark make the release decision.

A performance tester can model demand, run controlled experiments and explain observed limits. They cannot decide what experience users should accept, expose production to load without authority, approve unlimited cost or turn one test into a capacity promise. Define those lines before a target becomes a script.

01

Client product, production, and release authority

Named owners define the outcome, acceptable envelope, safe test boundary and consequential decision.

  • Product service operations and business owners define critical journeys, correct completion, user populations, service indicators, targets, forecast horizon, unacceptable degradation and the tradeoffs among speed, reliability, security, accessibility, cost and delivery time.
  • Engineering architecture platform data network and supplier owners identify release, topology, configurations, quotas, autoscaling, dependencies, test data, cache and connection state, observability limits and known incidents; they approve representative environments and any difference from production.
  • Production incident finance security privacy and release owners approve load, time, budget, data, access, monitoring, stop conditions and recovery; accept residual uncertainty; choose remediation or capacity changes; and make launch, release, scaling and customer-commitment decisions. The tester does not sign for them.
02

Performance tester contribution

The tester turns agreed outcomes and demand into controlled, explainable system evidence.

  • Build a versioned workload model from current telemetry and forecasts, separating arrivals from concurrency and representing journey, operation, payload, data, tenant, geography, device, network, pacing, abandonment, burst, background work and dependency conditions; state every proxy and missing population.
  • Design smoke, average-load, stress, spike, soak, breakpoint or component benchmarks for distinct questions; provision controlled accounts and data; validate correctness at low load; capture environment and release identity; warm and repeat deliberately; monitor the injector; gate hazardous runs; and preserve raw results with clocks, tool versions and configuration.
  • Analyze latency distributions, goodput, errors, dropped and timed-out work, queues, retries, throttling, resource saturation, scaling, traces, profiles, dependency behavior, recovery and cost; test bottleneck hypotheses; compare controlled runs; expose variance and limitations; and leave maintainable scripts, thresholds, dashboards and runbooks in client custody.
03

Shared measurement and capacity operating model

Performance is an end-to-end product property that crosses code, data, infrastructure and operations.

  • Product owners connect timings to meaningful outcomes; developers preserve correctness and add diagnostic instrumentation; platform and data teams expose queues limits and resources; reliability teams align test and production indicators; network browser mobile and supplier owners explain boundaries outside the service process.
  • Tools may generate traffic, capture timings, aggregate distributions, trace requests and suggest correlations, but people validate workload meaning, timing boundaries, missing data, clock alignment and causality. An automated pass cannot authorize production load, conceal dropped work or approve a release.
  • Governance keeps workloads, targets and forecasts current, reviews regression and capacity evidence with cost and risk, funds the chosen correction, verifies recovery and preserves a comparable baseline. Hiring owners confirm domain competence, safe judgment, collaboration, terms and current availability.

Capability evidence

Assess one performance question from user demand to an owned capacity decision.

A tool certificate or requests-per-second screenshot is weak hiring evidence. Use a real workflow with mixed demand, skewed data, a hidden queue, one incorrect fast response, a scaling delay and an environment difference that must be made visible.

01

User outcomes, indicators, targets, demand, and workload models

Provide product analytics, service telemetry, a forecast and a request to prove scale. Ask the person to define the outcome, measurement and workload before selecting a generator.

Confirm: The person begins with correct user or system completion and distinguishes response time, end-to-end latency, queue time, throughput, goodput, freshness, batch duration, availability and resource efficiency; defines the measured population, success criteria, exclusions, time window, percentiles and error policy; separates SLI, SLO, SLA, internal threshold and capacity forecast; avoids universal speed targets; identifies interactive, asynchronous, streaming and batch classes; derives arrival rates, concurrency, pacing, think time, abandonment, session length, journey and operation mix, payload and response size, read-write balance, tenant and key skew, data volume and growth, cache state, geography, device, network, scheduled work, retries and dependencies; models normal, burst, seasonal, growth and failure demand; distinguishes an open arrival model from a closed virtual-user model; records missing or biased telemetry; and obtains owner acceptance of proxies and target tradeoffs.

02

Reproducible scenarios, environments, generators, data, and safe execution

Provide an environment with smaller capacity, stale data, warm caches and shared neighbors, plus a script that saturates its own injector. Ask for a test system whose result can be repeated and bounded.

Confirm: The person maps the end-to-end path and differences among local, component, integration, preproduction and production conditions; records release, build, configuration, feature flags, topology, instance types, quotas, autoscaling, regions, network, dependencies and observability; designs smoke, average-load, stress, spike, soak, breakpoint and microbenchmarks only for their specific questions; creates separable safe accounts and representative data distributions; preserves cold, warm and steady states deliberately; validates functionality before load and samples correctness during it; synchronizes useful clocks and correlation identity; versions scripts, images, tools and parameters; sizes and monitors load generators, connections and network; uses distributed generation only when necessary and reconciles nodes; defines ramp, duration, warmup, steady state, repetition, cooldown and cleanup; controls cost and test data; obtains production approval; sets abort conditions for user harm, error, saturation, spend or lost observability; and restores the system after stress.

03

Distributions, correctness, bottleneck hypotheses, and causal evidence

Provide a run with a healthy mean, a poor tail, fast failures, dropped arrivals, retry amplification and rising queue depth. Ask what the system actually did and where investigation should go next.

Confirm: The person keeps successful, failed, timed-out, canceled, rejected and wrong work visible; distinguishes offered load, accepted traffic, completed throughput and correct goodput; checks dropped iterations and coordinated-omission risk; reports counts and distributions with median and relevant high percentiles rather than averaging percentiles across workers; understands histogram boundaries, estimation error, sample size and aggregation; segments by journey, operation, status, payload, tenant, region, client and release without creating unusable cardinality; aligns client timing, server timing, traces, logs, profiles and resource metrics; observes latency, traffic, explicit and implicit errors, queues, locks, pools, caches, retries, throttling, garbage collection, CPU, memory, storage, network and dependency saturation; distinguishes symptom, correlation, contention and cause; identifies the capacity knee and nonlinear behavior; changes one useful factor; predicts the result; repeats under controlled conditions; and reports counter-effects on correctness reliability security cost and other workloads.

04

Capacity envelope, regression, recovery, decisions, and continuity

Provide a proposed optimization, an autoscaling change, a prior baseline and a release deadline. Ask for a decision record that remains useful after traffic and architecture change.

Confirm: The person states capacity as a bounded curve across demand mix, latency, goodput, errors, saturation and cost rather than one maximum; separates sustained capacity, burst tolerance, breakpoint, safety margin and forecast; tests overload controls including quotas, backpressure, shedding and retry behavior without assuming they preserve correctness; verifies scaling and downscaling lag, state movement and dependency limits; observes recovery, backlog drain, cache refill and lingering resource damage after load stops; compares candidate and baseline on equivalent releases, environments, data, scenarios and repetitions; reports run-to-run variance and material environment drift; treats a threshold as one gate, not universal safety; stores raw and summarized results, queries and evidence; links regression to an owner and decision; updates continuous checks at proportionate frequency; retires stale workloads; and demonstrates that another qualified person can run, interpret and safely stop the suite.

Assessment sequence

Move from an observable outcome to a capacity envelope with stated limits.

Performance work becomes defensible when the workload, test system, measurements and decision remain connected. Start with the behavior that matters and keep failed or missing work in the denominator.

  1. 01

    Define the outcome and performance contract

    Identify correct completion, users and flows, indicators, populations, windows, percentiles, thresholds, forecasts, unacceptable behavior, tradeoffs and decision owners; record weak proxies and unavailable evidence.

  2. 02

    Model representative demand and state

    Derive arrivals, concurrency, pacing, journeys, operations, payloads, data and tenant skew, cache and connection states, locations, devices, networks, background work, dependencies, bursts, growth and failures from current evidence.

  3. 03

    Build and validate the test system

    Version the release environment data scripts tools and generator capacity, establish observability and correlation, prove correctness at low load, define warmup steady state repetition cleanup cost and abort rules, and approve any production exposure.

  4. 04

    Run, observe, and test bottleneck hypotheses

    Increase or vary demand for the stated scenario, preserve errors timeouts drops and wrong results, inspect latency distributions queues resources traces profiles dependencies scaling and cost, then change one factor and repeat to challenge the suspected constraint.

  5. 05

    Compare, recover, decide, and preserve

    Verify backlog drain and restored state, compare controlled repeated runs, state capacity headroom variance and limits, route regressions and options to owners, retain raw evidence and update or retire continuous test assets.

Performance evidence loops

Keep every fast number attached to demand, correctness, system state, and a decision.

False confidence appears when a script defines the user, averages erase the tail, generator failures reduce offered load or an optimization moves cost and failure elsewhere. These loops keep the test interpretable.

  1. 01

    Outcome and workload loop

    Does the generated demand still represent the people, systems and complete work covered by the target?

    Working evidence: Outcome, correctness rule, indicator, objective and window, decision owner, analytics and telemetry period, forecast, arrival and concurrency model, pacing, session and journey mix, operation, payload, data and key skew, tenant, geography, device, network, cache, scheduled work, dependency, retry, burst, growth, proxy and exclusion.

  2. 02

    Run identity and measurement loop

    Can another tester reproduce the offered load and account for every completed, failed, dropped or abandoned unit?

    Working evidence: Run and scenario identity, release, commit, artifact, configuration, flags, infrastructure, topology, quotas, dataset, cache and connection state, tool and script version, generator nodes, clocks, ramp, warmup, steady state, duration, repetitions, offered arrivals, active users, accepted requests, completed work, failures, timeouts, cancellations, drops and correct goodput.

  3. 03

    Distribution and constraint loop

    Which constraint explains the change in user-visible behavior as demand rises?

    Working evidence: Latency histograms and percentiles by meaningful segment, sample count and estimation limits, errors and wrong results, queue and pool depth, lock and wait state, CPU, memory, allocation, garbage collection, storage, network, cache, throttle, retry, scaling, dependency, trace, log, profile, cost, capacity knee, hypothesis, controlled factor and repeated result.

  4. 04

    Change, regression, and recovery loop

    Did the changed system improve the same workload, avoid unacceptable counter-effects and return to a healthy state?

    Working evidence: Baseline and candidate identity, equivalent conditions, run sequence, variance, threshold result, correctness, latency, goodput, errors, saturation, cost, security and reliability counter-effects, backlog drain, scale stabilization, cache recovery, leaked resources, cleanup, owner, decision, residual limit, forecast trigger, next run and receiving team.

Continuity controls

Recover without one tester, one dashboard, or a baseline whose environment disappeared.

The useful asset is a reproducible question and evidence chain. Workloads, scripts, raw results, analysis and safe operating instructions should remain with the client and change when the system or demand changes.

Versioned performance contract and workload
User outcomes, correct completion, indicators, targets, populations, windows, forecasts, journey and operation mix, arrival and concurrency models, data distributions, client conditions, dependencies, assumptions, exclusions and owner decisions remain reviewable.
Reproducible test-system manifest
Releases, configurations, environments, infrastructure, topology, quotas, datasets, cache states, scripts, tools, generators, clocks, access controls, run parameters, safety gates, cost limits, cleanup and known production differences are captured as code or durable records.
Raw run, diagnosis, and comparison record
Counts, distributions, errors, wrong results, drops, queues, resources, traces, profiles, logs, scaling, dependencies, cost, hypotheses, changed factors, repetitions, variance, environment drift, bottleneck conclusions and rejected explanations remain linked to each run.
Demonstrated handoff and suite retirement
Another qualified person can prepare, validate, run, observe, abort, clean up, analyze and compare the suite; thresholds and dashboards have owners; stale scenarios are retired; and new releases demand models or system boundaries trigger review.

Fit check

Use a performance tester when a system decision needs representative load and reproducible limits.

Good reason to begin

  • A critical workflow, release, scale forecast, bottleneck, capacity question or performance regression can be bounded, and owners can agree correct completion, relevant indicators and acceptable tradeoffs.
  • The client can provide current traffic and data evidence, identifiable builds and environments, controlled accounts, useful telemetry, safe test capacity, domain owners and authority for any production or costly run.
  • Engineering can act on findings, product and risk owners will accept uncertainty and limits, and the organization wants reusable workloads, diagnostic evidence, recovery proof and comparative decisions rather than one benchmark score.

Resolve before beginning

  • No one can define which user or system outcome matters, correctness is unknown, the target is an arbitrary number, or leadership wants a result that confirms a predetermined infrastructure purchase or launch claim.
  • The release environment data and dependencies cannot be identified, observability is too weak to account for failed or dropped work, the generator cannot be monitored, or different runs cannot be made meaningfully comparable.
  • A production test lacks accountable approval, customer protection, incident contacts, cost limits, abort conditions and recovery, or test accounts and data cannot be separated safely from real people and consequential effects.
  • There is no engineering owner or time to investigate and repair constraints, no path to verify recovery or retest a change, or the expected deliverable is only a tool script, dashboard screenshot or universal scalability guarantee.

Source basis

Sources behind the control model.

  • 01

    Google Site Reliability Engineering

    Service Level Objectives

    The chapter defines SLIs and SLOs, connects latency and throughput to user needs, warns that averages hide tail behavior and performance cliffs, and recommends explicit populations, windows and percentiles. Its Google examples are not client targets.

  • 02

    Google Site Reliability Engineering

    Monitoring Distributed Systems

    The chapter distinguishes user-visible symptoms from internal causes and organizes latency, traffic, errors and saturation as core signals. It also separates successful from failed latency and warns that systems may degrade before full utilization.

  • 03

    Google Site Reliability Engineering

    Handling Overload

    The chapter describes overload, admission control, load shedding and client-side throttling in large distributed systems. Its specific mechanisms require local validation and do not establish capacity for another architecture.

  • 04

    Google Site Reliability Engineering

    Addressing Cascading Failures

    The chapter explains how overload, queues, retries, resource exhaustion and recovery behavior can amplify failure across dependencies. It supports testing degraded paths and recovery instead of measuring healthy steady state alone.

  • 05

    Grafana Labs

    Automated performance testing

    The current k6 guidance connects scenarios, production-informed traffic, continuous comparison and pass-fail criteria, while warning that a pass status alone can create false confidence for larger tests.

  • 06

    Grafana Labs

    Load test types

    The current guide distinguishes smoke, average-load, stress, soak, spike and breakpoint tests by purpose and shape. Choosing a named type does not make its workload or target appropriate for a specific service.

  • 07

    Grafana Labs

    Thresholds

    Thresholds encode pass-fail conditions over test metrics and can support automation. Their usefulness depends on approved metric meaning, scope and workload, and they do not replace analysis of missing traffic or unsafe conditions.

  • 08

    Grafana Labs

    Metrics

    The current reference distinguishes counters, gauges, rates and trends and identifies request counts, failures and duration as complementary signals. Tool summaries still need correctness, system and environment context.

  • 09

    Grafana Labs

    Scenarios

    The current reference supports separate workload functions, scheduling, tags and open or closed execution models. It demonstrates why arrival rate, virtual users, timing and scenario identity must be chosen deliberately.

  • 10

    Apache Software Foundation

    Apache JMeter best practices

    The official manual covers resource-efficient non-GUI execution, scripting, result handling, distributed tests and injector constraints. These are tool-operating practices, not proof that a workload is representative or a result causal.

  • 11

    OpenTelemetry

    Performance Benchmark of OpenTelemetry API

    The specification defines a controlled benchmark configuration and reporting shape for throughput, CPU and memory overhead of telemetry SDK behavior. Its narrow benchmark is a useful example of explicit conditions and limits.

  • 12

    OpenTelemetry

    Semantic conventions for HTTP metrics

    The current conventions define HTTP request and response metric meaning, units and attributes across client and server observations. Consistent names help correlation but do not guarantee complete instrumentation or aligned clocks.

  • 13

    Prometheus

    Histograms and summaries

    The official guidance explains quantiles, bucket choices, estimation error and why precomputed quantiles should not be averaged across instances. Distribution design and sample coverage remain local responsibilities.

  • 14

    World Wide Web Consortium

    Navigation Timing Level 2

    The current Working Draft defines document-navigation timing from the user agent. Its draft status, cache behavior, same-origin rules and browser coverage must be preserved when interpreting client-side measurements.

  • 15

    World Wide Web Consortium

    Resource Timing

    The current Candidate Recommendation Draft defines resource fetch timing, sizes, status and delivery details available to web applications. It is a measurement interface, not an experience target or complete user journey.

  • 16

    World Wide Web Consortium

    Server Timing

    The specification defines how servers expose selected request-response metrics to user agents through the Server-Timing header. Published timing descriptions need deliberate naming, privacy review and correlation with end-to-end evidence.

  • 17

    Google web.dev

    Web Vitals

    The current guidance defines Core Web Vitals around loading, responsiveness and visual stability and distinguishes field from lab measurement. Thresholds and sampled browser evidence do not replace product-specific journey correctness.

  • 18

    Microsoft Azure Well-Architected Framework

    Architecture strategies for performance testing

    The current guidance recommends measurable goals, early and repeated tests, realistic conditions, hypothesis-driven comparison, multiple test types and performance-informed design decisions. Azure examples do not establish provider-neutral capacity.

  • 19

    Amazon Web Services

    AWS Well-Architected Performance Efficiency pillar

    The current pillar organizes selection, review, monitoring and tradeoffs for efficient cloud resources and encourages experimentation and load testing. Recommendations are provider-specific inputs, not measured proof for a client workload.

  • 20

    Google Cloud Architecture Center

    Well-Architected performance optimization pillar

    The current pillar connects requirements, architecture, observability, load testing, capacity planning and continuous optimization. Its provider guidance must be tested against the identified release, configuration and demand.

[ WORKFLOW / SYSTEMS AUDIT ]
THE FIRST ENGAGEMENT

Start with one real workflow

A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.

Show Us the WorkflowStart with the free automation readiness checklist

OBSERVEQUANTIFYDECIDEBUILD