How Data Injection Attacks Work: from harmless input to unintended instructions
Follow an injection attack through decoding, string construction, interpretation and impact, including second-order injection and background-job trust boundaries.
Scope
Follow an injection attack through decoding, string construction, interpretation and impact, including second-order injection and background-job trust boundaries.
Was I hacked?
An injection-shaped request is not proof that the interpreter obeyed it. Begin with a timeline: the inbound event, application handling, downstream query or process, response and any subsequent activity. Evidence becomes stronger when the same request ID, account, process tree or database session connects those stages.
Escalate to incident response if a suspicious request is followed by data outside the requester’s scope, a web-service child process, an unexpected file, a new account, altered configuration or unexplained outbound traffic. Preserve logs and volatile evidence before restarting the service. A restart can close the attacker’s session and also close the detective’s notebook.
The attack is a parser disagreement
Applications and interpreters often disagree about where data ends and instructions begin. The application believes it is forwarding a name, search term or filename. The receiving component sees syntax.
The generic path is:
source
↓
decode and transform
↓
combine with an instruction
↓
interpreter parses the result
↓
runtime identity performs the operation
↓
response, delay or side effect reveals the outcome
CWE-74 describes the weakness at the downstream boundary. This is important: validation at the browser is not the boundary when a queue consumer, database or shell makes the authoritative decision later.
Stage 1: input enters through an ordinary feature
The source may be an HTTP request, but it can also be a spreadsheet import, webhook, message queue, synchronisation feed, stored profile field or administrator console. Injection is not obliged to arrive wearing a suspicious user-agent string.
Consider an export feature:
GET /reports?sort=newest&format=csv
Both values look harmless. The risk appears only when the application translates them into other languages. sort might become SQL structure; a cell in the export might become spreadsheet formula syntax; format might become a command-line argument.
The defensive question is not “does the request contain a quote?” It is “where will each value be interpreted next?”
Stage 2: decoding changes the representation
Web frameworks decode URL encoding, proxies may normalise paths, JSON parsers create objects and applications may decode a value more than once. If validation sees one representation and execution sees another, an apparently blocked token can reappear later.
Use one documented canonicalisation path:
const raw = String(req.query.sort ?? 'newest');
// The framework has already URL-decoded the query parameter.
// Do not decode it again.
const allowed = new Set(['newest', 'oldest']);
if (!allowed.has(raw)) return res.status(400).send('Invalid sort option');
Repeated decoding is not a clever compatibility feature. It is an ambiguity subscription.
Stage 3: data becomes structure
The decisive mistake is usually string construction or accepting a client-owned object as executable structure.
Unsafe SQL construction:
const statement = `SELECT id, title FROM posts WHERE title LIKE '%${term}%'`;
await pool.query(statement);
Safe value binding:
await pool.query(
'SELECT id, title FROM posts WHERE title ILIKE $1',
[`%${term}%`]
);
Unsafe NoSQL structure:
await users.findOne(req.body);
Safer server-owned structure:
const username = String(req.body.username ?? '');
if (!/^[a-z0-9._-]{3,40}$/i.test(username)) {
return res.status(400).json({ error: 'Invalid username' });
}
await users.findOne({ username });
The second version does not “sanitise MongoDB”. It refuses to let the client define the query object.
Stage 4: the interpreter assigns meaning
At this point the SQL engine, shell, LDAP server or template compiler follows its own grammar. The same character has different meaning in each context. This is why a generic function named sanitize() should make reviewers nervous unless its contract names the exact destination.
For OS commands, even escaping shell separators may leave argument injection. A user-controlled value beginning with -- can change how a fixed executable behaves. OWASP recommends avoiding OS commands, using parameterisation and validation, and inserting -- where the called utility supports an end-of-options delimiter.
execFile('/usr/bin/file', ['--brief', '--', absolutePath], {
shell: false,
timeout: 5000
}, callback);
The executable is fixed, the shell is absent, options end before the user-derived path and the process has a timeout. Each control removes one interpretation opportunity.
Stage 5: privilege determines the blast radius
Injection does not manufacture privileges from nothing. It borrows the authority of the vulnerable component.
A database connection limited to selected views may expose fewer records than a schema owner. A web service confined by AppArmor or SELinux and denied outbound networking is less useful to an attacker than a service running as root. A directory bind account with read-only access limits modifications.
Record the effective identities during assessment:
web process: app-web
database role: app_reader
filesystem access: /srv/app/uploads only
outbound network: package proxy and mail relay only
Least privilege is not a substitute for parameterisation. It is the seat belt worn while the engineering team repairs the brakes.
Direct, blind and side-effect evidence
An application may return data directly, report an error, behave differently, pause or create an observable side effect. Defenders should not assume that a normal-looking page means the sink was safe.
Useful defensive evidence includes:
- a mismatch between expected and returned row counts;
- database parse errors tied to a request ID;
- repeated requests with small input changes and different latency;
- unexpected child processes from a service;
- DNS or HTTP connections from a server that normally makes none;
- validation failures followed by a successful request from the same identity;
- configuration or file changes immediately after the request.
Do not reproduce dangerous side effects in production. Build a local or isolated staging test that replaces the downstream client with a recorder.
A safe local data-flow test
The following test does not connect to a database. It verifies that a repository method sends fixed SQL and a separate parameter array:
import assert from 'node:assert/strict';
const calls = [];
const pool = { query: async (...args) => { calls.push(args); return { rows: [] }; } };
async function searchPosts(term) {
return pool.query(
'SELECT id, title FROM posts WHERE title ILIKE $1',
[`%${term}%`]
);
}
await searchPosts("O'Brien");
assert.equal(calls[0][0], 'SELECT id, title FROM posts WHERE title ILIKE $1');
assert.deepEqual(calls[0][1], ["%O'Brien%"]);
Run it with the Node test runner after placing it in a test file:
node --test test/search-posts.test.mjs
The assertion proves the data boundary in this method. Add an integration test against an isolated database to prove the actual driver and deployed configuration behave as expected.
Second-order injection
Sometimes the first component stores the value safely, but a later component retrieves it and builds an instruction unsafely. This is second-order injection.
Example path:
profile display name
→ safely stored as data
→ nightly export job reads it
→ export filename is concatenated into a shell command
→ command runner interprets it
The input may have lived peacefully in the database for months. Validating only the original HTTP route misses the later sink. Treat stored and internal values as untrusted when they cross into a new interpreter.
Background jobs and service-to-service trust
Queue consumers are common injection boundaries because producers and consumers are owned by different teams. Define a message schema and reject additional properties:
const reportSchema = {
type: 'object',
additionalProperties: false,
required: ['accountId', 'format'],
properties: {
accountId: { type: 'string', pattern: '^[a-f0-9-]{36}$' },
format: { enum: ['pdf', 'csv'] }
}
};
Schema validation makes the message shape explicit. The consumer must still use safe database, template and process APIs. A valid accountId does not authorise access to every account.
Where scanners help — and stop
Static analysis can trace data from sources to sinks. Dynamic tools can send representative inputs and compare responses. Both are useful, but neither understands every business rule, generated query or background path.
Start with a sink inventory:
rg -n --glob '!vendor/**' --glob '!node_modules/**' \
'exec\(|system\(|shell=True|eval\(|query\(|execute\(|search\(' \
src app
Review each match for data flow. Do not close a finding merely because a WAF blocks one request. The vulnerable path may remain reachable internally or through a different encoding.
If exploitation is suspected
- Restrict the vulnerable route or service at the safest upstream control.
- Preserve application, proxy, database and endpoint events with consistent time zones.
- Identify the exact runtime and data-service identities.
- Search for effects: broad reads, child processes, files, accounts, tokens and outbound traffic.
- Rotate exposed credentials from a trusted system when evidence supports exposure.
- Fix the data/instruction boundary, add regression tests and redeploy.
- Repeat the original authorised test against the deployed release.
Continue with How to Detect Data Injection for a complete evidence workflow and Data Injection Prevention for implementation controls.
References
- MITRE CWE-74: Injection
- OWASP Injection Prevention Cheat Sheet
- OWASP Input Validation Cheat Sheet
- OWASP OS Command Injection Defense Cheat Sheet
- OWASP NoSQL Security Cheat Sheet
- OWASP Web Security Testing Guide: Input Validation Testing
Reviewed 5 September 2026.