Insecure design: flaws that patches cannot fix
Insecure design describes missing or ineffective controls in the workflow itself, even when the implementation has no obvious coding bug.
Scope
For product, architecture, and engineering teams designing sensitive workflows and trust boundaries.
Start with the situation, not the slogan
A vulnerability is not made useful by giving it a dramatic name. It becomes useful when a team can identify the affected trust boundary, reproduce the unsafe behaviour safely, explain the consequence, and verify that the correction changes the result. Insecure design: flaws that patches cannot fix is therefore treated here as an engineering condition rather than a badge for a dashboard.
Severity depends on exposure, reachable functionality, data, privilege and compensating controls. A scanner can point at a door; it cannot reliably tell you who can reach it, what is behind it, or whether the hinges are decorative. Start with the system’s intended rule, then compare that rule with observed behaviour.
If password recovery, approvals, payments, tenancy, or rate limits are unsafe by design, input validation and patching will not supply the missing security property.
Composite scenario
How this usually reaches the desk
A useful review begins when a tester can state an expectation and an observation in the same sentence. The expectation is the documented boundary; the observation is recorded as “Sensitive actions have no abuse cases, trust-boundary diagram, or explicit security requirements.” The next task is not to launch a larger bag of payloads. It is to reproduce the smallest authorised case, capture the request and result, and identify where the missing decision should have been made.
This scenario combines common operational patterns; it is not presented as a report of one named incident.
What to look for
Begin with preserved, comparable evidence. One signal is rarely proof; use independent observations and a reliable timeline before declaring scope or intent.
Sensitive actions have no abuse cases, trust-boundary diagram, or explicit security requirements.
Translate this clue into a testable condition: actor, resource, operation and expected denial or safe handling. Preserve the smallest request and response that demonstrate the difference.
A single user, service, or step can approve and execute a high-impact action without independent control.
Check whether the behaviour exists in source, configuration, generated artefacts and the deployed service. Many splendid fixes have lived only in a branch while production carried on with its own arrangements.
The design assumes clients, networks, integrations, or support staff will always behave honestly.
Measure reach and consequence. Determine which roles, tenants, records, secrets or processes are exposed, and whether an existing control genuinely prevents abuse or merely makes it less convenient.
Working examples
Run it, read it, decide what changes
These examples use documentation addresses, test identities and bounded targets. Replace placeholders only inside systems you own or are explicitly authorised to operate. Read the expected result and next action before running the command; a successful command is evidence, not yet a conclusion.
Add an upstream timeout
- Prerequisites
- An owned service calling a dependency.
const response = await fetch('https://dependency.internal/health', {
signal: AbortSignal.timeout(3000),
headers: {'X-Request-ID': req.id}
});The request either completes or aborts after about three seconds instead of waiting indefinitely.
A timeout without error handling may still produce a 500 storm. Three seconds is an example, not a universal value.
Handle timeout separately, return a safe degraded response, emit a metric and test dependency slowness in staging.
Make a write request idempotent
- Prerequisites
- A staging endpoint supporting idempotency keys.
key="nmf-test-$(date +%s)"
curl -i -X POST https://api.test.local/payments -H "Idempotency-Key: $key" -H 'Content-Type: application/json' -d '{"amount":100,"currency":"EUR"}'
curl -i -X POST https://api.test.local/payments -H "Idempotency-Key: $key" -H 'Content-Type: application/json' -d '{"amount":100,"currency":"EUR"}'Both requests return the same payment identity and only one business operation is committed.
A key stored only in memory fails across replicas or restarts; keys also need a scope and expiry.
Test concurrency and restart behaviour, persist the result transactionally, and alert on key/payload conflicts.
Exercise a dependency failure
- Prerequisites
- A non-production dependency service and a documented restart command.
docker compose stop dependency
curl -i https://app.test.local/health
docker compose start dependencyThe application returns the designed degraded/health response and recovers after the dependency starts.
Stopping a production dependency is not an acceptable first test. A green health endpoint can miss broken user flows.
Measure error budget, queue behaviour and data consistency, then verify a real read and write after recovery.
What to do
Read the whole sequence before starting. Several workstreams may run in parallel, but their evidence, authority and expected outcomes still need to be explicit. Every step below points back to a concrete example; use the example as implementation evidence, not as permission to operate outside the stated scope.
Operational judgement
Prioritisation should combine likelihood, consequence and exposure. Internet reachability, valuable data, privileged execution and a reliable abuse path all raise urgency. Strong isolation, narrow permissions or a disabled feature may lower immediate risk, but document those assumptions and test them. A numerical score is useful shorthand; it is not a substitute for knowing which business process can be harmed.
A durable remediation usually has three layers: remove the immediate unsafe path, improve the design or default that allowed it, and add a signal that reveals recurrence. For OWASP A06, Threat modeling and Design, this often means aligning application behaviour, deployment configuration and operational monitoring rather than asking one patch to perform a small miracle.
Handover
Make the result useful to the next person
Write the finding so an engineer can act without translating theatre into requirements. Include affected component and version, actor and preconditions, smallest safe reproduction, expected rule, observed result, consequence, evidence, and the owner of remediation. Keep secret values and personal data out of the ordinary ticket. If disclosure to a vendor or maintainer is required, use their published security channel and agree on what may be shared before placing proof in a public issue tracker.
The handover to remediation should begin with “Model assets, actors, entry points, trust boundaries, abuse cases, and business impact before implementation.” and preserve the exact case needed to prove “Revisit the threat model after architecture, integration, or business-process changes.” Link code and configuration changes to that test. Record whether the remedy eliminates the unsafe state or merely narrows exposure, and name any compensating control. A reviewer should be able to answer three questions without calling the original tester: what was wrong, why this change is sufficient, and how production will demonstrate the corrected behaviour.
Validate before you close
Run abuse-case reviews and tabletop exercises with product, engineering, operations, and business owners; confirm the implemented workflow enforces the intended controls.
Capture the test, the expected result and the observed result. Where a person or business owner must accept restored service, name them in the record. A green dashboard can confirm that a component is answering; it cannot confirm that invoices, identities or restored data are trustworthy.
Finish with a compact closure note: the original trigger, confirmed scope, evidence retained, controls changed, tests passed, known gaps, residual risk, and the people responsible for the remaining work. Schedule a review while the timeline is still fresh enough to challenge. The purpose is not to find a person to blame; computers already perform blame with admirable efficiency. The purpose is to make the next response faster, safer and less dependent on one person remembering where the useful log was hidden.
Common mistakes
- Treating a scanner result as proof without confirming the affected path and version.
- Applying a tactical patch while leaving the underlying design weakness in place.
- Testing production with exploit code before establishing an authorised, isolated validation plan.
These errors usually come from haste, unclear ownership or misplaced confidence. Build the safeguard into the runbook: a required evidence field, a second-person review, a rollback test or a specific exit criterion.
Questions people ask when the clock is running
Is a scanner result enough to open a critical incident?
It is enough to triage. Confirm the affected version and reachable path, reproduce safely, and measure privilege and data impact. A false positive and a missed exposure are both easier to manage when the evidence is explicit.
Can a compensating control count as the fix?
Sometimes, for a defined period. It must be enforced, monitored, owned and tested against the same abuse case. Write down its expiry or review date so “temporary” does not become a geological era.
What proves remediation?
Run abuse-case reviews and tabletop exercises with product, engineering, operations, and business owners; confirm the implemented workflow enforces the intended controls. Keep the original test as a regression case and verify the deployed system, not merely the ticket status.
Safety boundary
Use these steps only on systems you own or are explicitly authorised to assess. Preserve evidence, follow your organisation’s legal and regulatory obligations, and prefer reversible actions when the situation is not yet understood.
Primary references
- OWASP Top 10:2025OWASP
- Threat Modeling Cheat SheetOWASP
- Docker SecurityDocker Docs
Editorial status: first edition. Review the linked vendor documentation for product- and version-specific changes before acting.