——
HighVULNERABILITIES

Logging and alerting failures: evidence you cannot use

Security logging fails when important events are absent, ambiguous, delayed, unprotected, or impossible to turn into a timely decision.

For application, platform, and detection teams defining operational security telemetry.

OWASP A09LoggingDetection

Start with the situation, not the slogan

A vulnerability is not made useful by giving it a dramatic name. It becomes useful when a team can identify the affected trust boundary, reproduce the unsafe behaviour safely, explain the consequence, and verify that the correction changes the result. Logging and alerting failures: evidence you cannot use is therefore treated here as an engineering condition rather than a badge for a dashboard.

Severity depends on exposure, reachable functionality, data, privilege and compensating controls. A scanner can point at a door; it cannot reliably tell you who can reach it, what is behind it, or whether the hinges are decorative. Start with the system’s intended rule, then compare that rule with observed behaviour.

Logs are useful only when they identify actor, action, object, result, time, and context, reach a protected location, and trigger an owned response before the damage is complete.

How this usually reaches the desk

A useful review begins when a tester can state an expectation and an observation in the same sentence. The expectation is the documented boundary; the observation is recorded as “Authentication, authorisation, administrative, export, and data-access events cannot be connected to a stable identity.” The next task is not to launch a larger bag of payloads. It is to reproduce the smallest authorised case, capture the request and result, and identify where the missing decision should have been made.

This scenario combines common operational patterns; it is not presented as a report of one named incident.

What to look for

Begin with preserved, comparable evidence. One signal is rarely proof; use independent observations and a reliable timeline before declaring scope or intent.

01

Authentication, authorisation, administrative, export, and data-access events cannot be connected to a stable identity.

Translate this clue into a testable condition: actor, resource, operation and expected denial or safe handling. Preserve the smallest request and response that demonstrate the difference.

02

Critical failures appear only in local files that attackers or system loss can remove.

Check whether the behaviour exists in source, configuration, generated artefacts and the deployed service. Many splendid fixes have lived only in a branch while production carried on with its own arrangements.

03

Alerts have no owner, threshold rationale, runbook, or measure of whether anyone responded.

Measure reach and consequence. Determine which roles, tenants, records, secrets or processes are exposed, and whether an existing control genuinely prevents abuse or merely makes it less convenient.

Run it, read it, decide what changes

These examples use documentation addresses, test identities and bounded targets. Replace placeholders only inside systems you own or are explicitly authorised to operate. Read the expected result and next action before running the command; a successful command is evidence, not yet a conclusion.

Example 01

Emit a structured security event

Application loggerNode.js
Prerequisites
A logger that serialises objects as JSON and a request correlation ID.
javascript
logger.warn({
  event: 'authorisation_denied',
  actor_id: req.user.id,
  resource_type: 'invoice',
  resource_id: req.params.id,
  request_id: req.id,
  outcome: 'deny'
});
Expected result

One parseable event contains stable names and the decision context without passwords or tokens.

How to interpret it

Logging raw request headers or bodies can leak credentials and personal data. A log call does not prove delivery.

Next action

Generate one denial, find it in the central platform, verify time and fields, then alert on a bounded meaningful pattern.

Example 02

Verify end-to-end collection

logger and central searchLinux
Prerequisites
A host whose syslog is meant to reach the central platform.
shell
marker="NMF_LOG_TEST_$(date -u +%Y%m%dT%H%M%SZ)"
logger -p auth.notice -t nmf-test "$marker"
printf '%s
' "$marker"
Expected result

A unique marker is emitted locally and printed for pasting into the central search.

How to interpret it

Local command success proves only submission to the local logging path.

Next action

Search centrally, record arrival delay and parsed host/source fields, then alert when this scheduled canary stops arriving.

Example 03

Check journal retention boundaries

journalctlsystemd Linux
Prerequisites
Read access to the journal.
shell
journalctl --disk-usage
journalctl --list-boots
journalctl --since "7 days ago" --until "6 days ago" -n 3 --no-pager
Expected result

Disk use, retained boot ranges and sample events from an older window are shown.

How to interpret it

Empty historical output can mean low activity, filters or insufficient retention.

Next action

Set retention from business and legal needs, monitor rollover, and test recovery from the central archive.

What to do

Read the whole sequence before starting. Several workstreams may run in parallel, but their evidence, authority and expected outcomes still need to be explicit. Every step below points back to a concrete example; use the example as implementation evidence, not as permission to operate outside the stated scope.

  1. 01

    Define the decisions responders must make, then identify the minimum events and fields needed for each decision.

    Define the security property in plain language before changing code. If the team cannot say what must always be true, it cannot write a convincing regression test for it.

    Working example 01: Emit a structured security event — Application logger on Node.js.

  2. 02

    Use consistent time, identity, request, tenant, object, outcome, and source fields across services.

    Reproduce with the least dangerous input in an isolated or explicitly authorised environment. Record version, configuration, identity and expected result so another engineer can verify the finding.

    Working example 02: Verify end-to-end collection — logger and central search on Linux.

  3. 03

    Forward security events to protected storage with retention, access control, integrity, and health monitoring.

    Fix the decision at the authoritative layer. Client-side checks, hidden buttons and polite documentation are helpful interface features, but they are not enforcement boundaries.

    Working example 03: Check journal retention boundaries — journalctl on systemd Linux.

  4. 04

    Create alerts for high-confidence abuse paths and test them with safe simulations.

    Add tests for allowed, denied and malformed cases. Include adjacent roles and objects, because vulnerabilities are sociable creatures and seldom respect the exact example in the original ticket.

    Working example 01: Emit a structured security event — Application logger on Node.js.

  5. 05

    Measure collection gaps, processing delay, alert acknowledgement, investigation quality, and false-negative findings.

    Deploy with monitoring and a rollback plan, then repeat the original case against the actual release. Evidence from a local build does not certify the service your users can reach.

    Working example 02: Verify end-to-end collection — logger and central search on Linux.

Operational judgement

Prioritisation should combine likelihood, consequence and exposure. Internet reachability, valuable data, privileged execution and a reliable abuse path all raise urgency. Strong isolation, narrow permissions or a disabled feature may lower immediate risk, but document those assumptions and test them. A numerical score is useful shorthand; it is not a substitute for knowing which business process can be harmed.

A durable remediation usually has three layers: remove the immediate unsafe path, improve the design or default that allowed it, and add a signal that reveals recurrence. For OWASP A09, Logging and Detection, this often means aligning application behaviour, deployment configuration and operational monitoring rather than asking one patch to perform a small miracle.

Make the result useful to the next person

Write the finding so an engineer can act without translating theatre into requirements. Include affected component and version, actor and preconditions, smallest safe reproduction, expected rule, observed result, consequence, evidence, and the owner of remediation. Keep secret values and personal data out of the ordinary ticket. If disclosure to a vendor or maintainer is required, use their published security channel and agree on what may be shared before placing proof in a public issue tracker.

The handover to remediation should begin with “Define the decisions responders must make, then identify the minimum events and fields needed for each decision.” and preserve the exact case needed to prove “Measure collection gaps, processing delay, alert acknowledgement, investigation quality, and false-negative findings.” Link code and configuration changes to that test. Record whether the remedy eliminates the unsafe state or merely narrows exposure, and name any compensating control. A reviewer should be able to answer three questions without calling the original tester: what was wrong, why this change is sufficient, and how production will demonstrate the corrected behaviour.

Validate before you close

Prove that a test event is generated, transported, parsed, correlated, alerted, assigned, investigated, and retained with enough context to explain the outcome.

Capture the test, the expected result and the observed result. Where a person or business owner must accept restored service, name them in the record. A green dashboard can confirm that a component is answering; it cannot confirm that invoices, identities or restored data are trustworthy.

Finish with a compact closure note: the original trigger, confirmed scope, evidence retained, controls changed, tests passed, known gaps, residual risk, and the people responsible for the remaining work. Schedule a review while the timeline is still fresh enough to challenge. The purpose is not to find a person to blame; computers already perform blame with admirable efficiency. The purpose is to make the next response faster, safer and less dependent on one person remembering where the useful log was hidden.

Common mistakes

  • Treating a scanner result as proof without confirming the affected path and version.
  • Applying a tactical patch while leaving the underlying design weakness in place.
  • Testing production with exploit code before establishing an authorised, isolated validation plan.

These errors usually come from haste, unclear ownership or misplaced confidence. Build the safeguard into the runbook: a required evidence field, a second-person review, a rollback test or a specific exit criterion.

Questions people ask when the clock is running

Is a scanner result enough to open a critical incident?

It is enough to triage. Confirm the affected version and reachable path, reproduce safely, and measure privilege and data impact. A false positive and a missed exposure are both easier to manage when the evidence is explicit.

Can a compensating control count as the fix?

Sometimes, for a defined period. It must be enforced, monitored, owned and tested against the same abuse case. Write down its expiry or review date so “temporary” does not become a geological era.

What proves remediation?

Prove that a test event is generated, transported, parsed, correlated, alerted, assigned, investigated, and retained with enough context to explain the outcome. Keep the original test as a regression case and verify the deployed system, not merely the ticket status.

Safety boundary

Use these steps only on systems you own or are explicitly authorised to assess. Preserve evidence, follow your organisation’s legal and regulatory obligations, and prefer reversible actions when the situation is not yet understood.

Primary references

  1. OWASP Top 10:2025OWASP
  2. Logging Cheat SheetOWASP
  3. SP 800-61 Rev. 3: Incident Response Recommendations and ConsiderationsNIST

Editorial status: first edition. Review the linked vendor documentation for product- and version-specific changes before acting.