——
GuidePLAYBOOKS

How to Detect Data Injection: a practical application and SOC playbook

Combine code review, safe staging tests, structured application events, web telemetry and endpoint evidence to find injection flaws and investigate attempted exploitation.

Combine code review, safe staging tests, structured application events, web telemetry and endpoint evidence to find injection flaws and investigate attempted exploitation.

Data injectionDetection engineeringOWASP ZAPSASTIncident response

Was I hacked?

A WAF alert, SQL error or injection-shaped URL is not enough to answer that question. It tells you that a request looked interesting or that a parser became unhappy. Compromise requires a connection between the input, the vulnerable sink and an unauthorised result.

Escalate when you see two or more independent signals: a suspicious request followed by a successful response, abnormal database activity, a web process spawning an interpreter, unexpected files, new identities, unusual exports or outbound connections from a service that normally has none.

Preserve the original event, but minimise sensitive data. Store a hash or rule identifier for a rejected payload when the body may contain credentials or personal data. Logs are evidence, not a second copy of the attacker’s luggage.

Detection needs three views

Use three complementary views:

  1. Code view: can untrusted data reach an interpreter without a safe boundary?
  2. Behaviour view: can a controlled staging input alter the downstream operation?
  3. Operational view: did production telemetry show an attempt and a meaningful effect?

No single scanner covers all three.

Step 1: build a sink inventory

Search for operations that invoke interpreters. Adapt extensions and paths to your stack:

rg -n --glob '!node_modules/**' --glob '!vendor/**' \
  'exec\(|execFile\(|spawn\(|system\(|shell=True|eval\(|query\(|execute\(|render_template_string|ctx\.search\(' \
  src app lib

Classify each match:

Sink Questions
Database query Is statement structure fixed? Are values bound separately?
Command runner Is a shell involved? Is the executable fixed? Are arguments separate and validated?
NoSQL query Can the client provide operators or a complete query object?
LDAP Is the value encoded for the exact filter or DN context?
Template Is user input template data or template source?
Logger/exporter Can delimiters, newlines or formula syntax change record meaning?

For every reachable sink, trace backwards to requests, queues, imports and stored values. A result from rg is a review queue, not a vulnerability report.

Step 2: add static analysis

Run the project’s established SAST tool in CI. If GitHub CodeQL is already enabled, use the security-extended query suite for broader security coverage and review language-specific injection queries in the CodeQL query documentation.

An example workflow configuration fragment is:

- name: Initialise CodeQL
  uses: github/codeql-action/init@v3
  with:
    languages: javascript-typescript
    queries: security-extended

Pin third-party actions according to your organisation’s supply-chain policy. Confirm the language list matches the repository and triage data-flow results with a developer who understands the framework.

Static analysis is strongest when it can see the source, transformations and sink. Generated queries, reflection, custom wrappers and background jobs may need framework models or manual review.

Step 3: write construction tests

An HTTP response test can miss an injection flaw that happens to return no records. Test the call made to the interpreter.

it('binds the search value instead of joining SQL text', async () => {
  const pool = { query: vi.fn().mockResolvedValue({ rows: [] }) };
  await searchUsers(pool, "O'Brien");

  expect(pool.query).toHaveBeenCalledWith(
    'SELECT id, email FROM users WHERE surname = $1',
    ["O'Brien"]
  );
});

For command execution, mock execFile or spawn and assert:

  • the executable path is fixed;
  • shell is false;
  • each value occupies one argument;
  • an end-of-options delimiter is present when supported;
  • a timeout and bounded output exist.

For NoSQL, assert that an object supplied where a string is expected is rejected before the driver call.

Step 4: perform a passive web review

OWASP ZAP’s Baseline Scan spiders for a limited time and performs passive checks. It does not actively attack the target, but crawling can still create ordinary GET requests, trigger analytics and reach logout or state-changing endpoints on poorly designed applications. Use a staging environment and an approved test account.

mkdir -p zap-reports

docker run --rm \
  -v "$PWD/zap-reports:/zap/wrk/:rw" \
  -t zaproxy/zap-stable \
  zap-baseline.py \
  -t 'https://staging.example.test' \
  -r zap-baseline.html

Open zap-reports/zap-baseline.html, record each route and separate configuration findings from injection evidence.

ZAP’s Full Scan and API Scan perform active testing. Use them only against an isolated or explicitly authorised environment with backups, rate limits and a named owner. Do not point a full scan at production because the calendar looked quiet.

Step 5: add bounded dynamic regression cases

Use harmless syntax-shaped data that should remain literal. Examples include a real surname with an apostrophe, {{ 7 * 7 }} in a message field and a filename beginning with a hyphen.

curl -i --get 'https://staging.example.test/search' \
  --data-urlencode "q=O'Brien" \
  -H 'X-Test-Case: injection-boundary-001'

Expected result:

  • a normal empty or matching result;
  • no database or template error;
  • no broadened data set;
  • the test-case identifier in application telemetry;
  • a bound value in database instrumentation;
  • no unexpected child process.

Do not infer safety from one input. The purpose is to verify a known repair, not to replace an application-security assessment.

Step 6: instrument application decisions

Web-server logs show what arrived; application logs should show what the application decided. Emit a structured event for validation failures without storing the full body:

logger.warn({
  event: 'input_validation_failure',
  ruleId: 'SEARCH_TERM_LENGTH',
  parameter: 'q',
  requestId: req.id,
  actorId: req.user?.id ?? null,
  sourceIp: req.ip,
  route: req.route.path,
  outcome: 'rejected'
});

Useful fields include timestamp with timezone, service, environment, route, parameter name, validation rule, result, request ID, actor, source address and a confidence label. Exclude passwords, tokens, cookies and complete query bodies.

OWASP’s Logging Cheat Sheet recommends logging input-validation failures and treating event data from other trust zones as untrusted. Sanitise carriage returns, line feeds and delimiters before they reach a line-oriented sink.

Step 7: collect downstream evidence

Database

Collect authentication, permission changes, schema changes, slow-query information and approved audit events. Avoid logging every parameter value when it may contain secrets. Alert on service accounts performing operations outside their expected tables or time windows.

Endpoint

Collect process creation with parent image, command line, user, hashes and container or host identity. High-value signals include a web worker spawning a shell, scripting engine, network utility or account-management tool.

Filesystem and network

Watch application directories for unexpected executables, scripts and configuration changes. Baseline the outbound destinations required by the service and alert on new destinations, especially after an injection alert.

Directory and application services

Record unusually broad LDAP searches, repeated filter errors, privilege changes and exports. A directory result count that grows from one expected object to thousands deserves investigation even when the HTTP status remains 200.

Step 8: correlate, do not merely count strings

A practical alert sequence is:

injection-shaped request or validation failure
  + successful/non-error response
  + abnormal database, process or file event
  = high-priority investigation

Tune known security scanners by approved source address, user agent and maintenance window, but do not discard their downstream effects. A scanner should not cause a web process to launch PowerShell merely because it has a change ticket.

Use the companion Sigma rules article to implement portable web and endpoint detections.

Step 9: triage a positive signal

  1. Confirm the target route, deployed version and feature flag.
  2. Identify the authenticated and runtime identities.
  3. Preserve proxy, application, database and endpoint events in a common timeline.
  4. Determine whether the request reached the interpreter.
  5. Identify the effect: data read, data changed, process created, file written or external connection.
  6. Restrict the route or service at a safe upstream point.
  7. Rotate credentials the process could expose, from a trusted system.
  8. Fix the boundary and add a regression test.
  9. Deploy, repeat the authorised test and monitor for recurrence.

If process execution or privileged data access occurred, treat the host and exposed credentials as potentially compromised. Rebuilding may be safer than attempting to clean an unknown execution chain.

Common detection failures

  • Only searching for apostrophes. Legitimate names create noise and other injection grammars remain invisible.
  • Only keeping WAF logs. Internal services, queues and stored values may bypass the edge.
  • Logging full bodies. This creates privacy, secret-handling and log-injection problems.
  • Treating 403 as prevention proof. The same sink may exist on another route or encoding path.
  • Ignoring successful responses. A low-error attack can be more important than a noisy failed probe.
  • No request ID. Without correlation, the SOC owns a pile of clocks rather than a timeline.

Closure evidence

Close the finding only when you have:

  • the vulnerable source-to-sink path documented;
  • a structured or context-correct fix in code;
  • unit and integration regression tests;
  • evidence from the deployed environment;
  • least-privilege review for the runtime identity;
  • monitoring for attempted recurrence;
  • an incident decision based on historical telemetry.

Prevention guidance is in Data Injection Prevention, and copyable safe patterns are in Data Injection Examples.

References

Reviewed 5 September 2026.