What is Data Injection? A practical definition for defenders and developers
Data injection happens when untrusted data changes the meaning of a query, command, template or downstream protocol. Learn the trust boundary, common forms and first checks.
Scope
Data injection happens when untrusted data changes the meaning of a query, command, template or downstream protocol. Learn the trust boundary, common forms and first checks.
Was I hacked?
Not merely because a scanner found the word injection. A vulnerability is a condition; compromise is an event. Treat the system as a possible incident when injection-shaped requests coincide with evidence such as unexpected database reads, new files, child processes from a web service, unexplained configuration changes, abnormal exports or access by identities that should not exist.
Preserve the request identifier, source address, authenticated identity, route, response status and relevant application, database and endpoint events. Do not paste an entire hostile request into a ticket or chat without checking it for credentials and personal data. Untrusted data has already caused enough excitement for one day.
If there is no evidence of exploitation, continue as a vulnerability investigation. The first job is to identify the interpreter that received the data.
Data injection in one sentence
Data injection happens when information that should remain data changes the instructions processed by a downstream component.
That component might be a SQL database, an operating-system shell, an LDAP directory, a template engine, a NoSQL query evaluator, a log parser, a spreadsheet or an AI system. The syntax changes, but the engineering failure is recognisable: an application combines an untrusted value with an instruction and sends the mixture to something that interprets both.
CWE-74 names the broad weakness “Improper Neutralization of Special Elements in Output Used by a Downstream Component”. OWASP expresses the same idea more plainly: injection flaws occur when an application sends untrusted data to an interpreter.
The four parts of an injection problem
Most useful investigations can be drawn as four boxes:
untrusted source → transformation → interpreter sink → security impact
1. Untrusted source
“Untrusted” does not mean “came from the public internet”. It means the receiving component cannot rely on the value having the intended structure or authority. Sources include:
- URL parameters, headers, cookies and request bodies;
- uploaded files and filenames;
- messages from queues, webhooks and partner APIs;
- database values originally supplied by users;
- administrator-entered fields;
- telemetry, logs and imported CSV files;
- model output passed into another tool.
An authenticated administrator can supply untrusted data. Authentication answers who sent it, not whether it is safe to interpret as code.
2. Transformation
The application may URL-decode, normalise Unicode, parse JSON, build a string, store the value, retrieve it later or forward it to another service. Bugs often hide here because developers validate one representation and execute another.
3. Interpreter sink
A sink is the operation that gives syntax meaning. Examples include pool.query(), exec(), an LDAP search filter, a template compilation method or a spreadsheet opening a cell as a formula.
4. Impact
Impact depends on what the interpreter can do and which identity it uses. A read-only database account limits different damage from a database owner. A shell running inside a restricted container is different from a service running as root. Least privilege does not repair injection, but it can turn a catastrophe into a bounded incident.
Common forms of data injection
| Family | Interpreter | Typical unsafe construction | Better boundary |
|---|---|---|---|
| SQL injection | SQL database | Query text plus user input | Prepared statement with bound values |
| NoSQL injection | Query object or expression engine | Client-controlled operators | Construct a typed server-side filter |
| Command injection | Shell or executable | Command string concatenation | Avoid the command; otherwise use a fixed executable and argument array |
| LDAP injection | LDAP filter or DN | Filter string concatenation | Correct LDAP context encoding or parameterised framework API |
| Template injection | Template compiler | User input treated as template source | Fixed template with escaped variables |
| Log injection | Log parser or human reader | Raw CR/LF and delimiters | Structured events, length limits and contextual encoding |
| CSV/formula injection | Spreadsheet | Cell beginning with formula syntax | Export policy that neutralises formula-capable cells |
SQL injection is therefore not a synonym for data injection. It is one specific member of the family. The dedicated comparison is in Data Injection vs SQL Injection.
A small SQL example
The unsafe version places the value inside SQL text:
const sql = `SELECT id, email FROM users WHERE email = '${req.query.email}'`;
const result = await pool.query(sql);
The safe shape keeps the SQL structure fixed and binds the value separately:
const sql = 'SELECT id, email FROM users WHERE email = $1';
const result = await pool.query(sql, [req.query.email]);
The database driver now knows which part is instruction and which part is data. OWASP’s SQL Injection Prevention guidance recommends prepared statements with parameterised queries as the primary defence.
Parameters are not a universal wand. They normally bind values, not identifiers such as a table name, column name or sort direction. Map those choices from a small server-side allowlist:
const allowedSort = {
newest: 'created_at DESC',
oldest: 'created_at ASC'
};
const orderBy = allowedSort[req.query.sort] ?? allowedSort.newest;
const result = await pool.query(
`SELECT id, title FROM articles ORDER BY ${orderBy} LIMIT $1`,
[25]
);
The user selects newest; the server selects SQL text. That is a decision boundary, not character filtering.
A small command example
This version asks a shell to interpret one combined string:
exec(`dig +short A ${req.query.hostname}`, callback);
A safer design uses a DNS library. If an external command is genuinely necessary, choose the executable in code, pass arguments separately, validate the hostname and run with a low-privilege identity:
import { execFile } from 'node:child_process';
const hostname = String(req.query.hostname ?? '');
if (!/^(?=.{1,253}$)([a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?\.)+[a-z]{2,63}$/i.test(hostname)) {
return res.status(400).send('Invalid hostname');
}
execFile('/usr/bin/dig', ['+short', 'A', hostname], {
timeout: 5000,
shell: false
}, callback);
OWASP’s OS Command Injection Defense guidance recommends avoiding OS commands first, then combining parameterisation with strict input validation when a command cannot be avoided.
What data injection is not
An apostrophe in a surname is not an attack. A validation error is not proof of exploitation. A 500 response after unusual input is evidence of unsafe error handling, not automatic proof that an interpreter executed hostile instructions.
Likewise, rejecting a list of famous payload strings does not establish safety. OWASP’s input-validation guidance warns that denylists are easy to bypass and frequently reject legitimate data. Input validation should describe what the business value may be; parameterisation or a safe API must still protect the interpreter boundary.
First checks for a code owner
Search the codebase for dangerous construction points. Adapt paths to your repository:
rg -n --glob '!node_modules/**' \
'exec\(|spawn\(|system\(|eval\(|query\(|execute\(|render_template_string' \
src app lib
This is a sink inventory, not a vulnerability verdict. For every match, record:
- the value’s original source;
- every decode or transformation;
- the exact API receiving it;
- whether structure and values are separate;
- the runtime identity and reachable data;
- an automated test proving the value remains data.
A simple regression test can mock the database client and check that the user value appears only in the parameters array:
expect(pool.query).toHaveBeenCalledWith(
'SELECT id, email FROM users WHERE email = $1',
['alice@example.test']
);
First checks for a defender
Look for layers of evidence rather than one regular expression:
- repeated validation failures on a normally finite field;
- database syntax errors or unusual query latency linked to a request ID;
- large or cross-tenant result sets;
- a web process spawning a shell, interpreter or discovery utility;
- unexpected file creation by a web-service identity;
- a sequence of probing requests followed by a successful response and post-exploitation activity.
Do not log complete secrets or hostile bodies merely to improve detection. OWASP’s Logging Cheat Sheet recommends recording input-validation failures while treating event data itself as untrusted and sanitising it against log injection.
The practical rule
When reviewing any feature, ask: Which component will interpret this value next? Then use that component’s structured API, its correct contextual encoding and a deliberately restricted runtime identity.
Continue through the cluster:
- How Data Injection Attacks Work
- Data Injection Examples
- How to Detect Data Injection
- Data Injection Prevention
- Sigma Rules for Detecting Injection Attacks
References
- MITRE CWE-74: Improper Neutralization of Special Elements in Output Used by a Downstream Component
- OWASP Injection Prevention Cheat Sheet
- OWASP SQL Injection Prevention Cheat Sheet
- OWASP OS Command Injection Defense Cheat Sheet
- OWASP Input Validation Cheat Sheet
- OWASP Logging Cheat Sheet
Reviewed 5 September 2026.