A Linux server may be compromised
A careful triage path for suspicious processes, new persistence, unexpected network traffic, and altered accounts on a Linux host.
Scope
For administrators investigating a Linux server that shows unexplained CPU use, access, processes, files, or outbound connections.
Start with the situation, not the slogan
An incident rarely arrives with a neat label. It arrives as a forwarded screenshot, a worried telephone call, or a monitoring alert written by a machine with no sense of occasion. In this case “A Linux server may be compromised” is the working heading, but the label is only a starting hypothesis. The useful work is to establish what happened, what can still happen, and which decision cannot safely wait.
Keep three clocks in view: the attacker’s opportunity, the business interruption, and the lifetime of the evidence. They do not run at the same speed. A hasty change may interrupt access but erase context; a perfect investigation conducted at geological pace may leave the organisation exposed. Good response is the slightly unglamorous art of making the next reversible decision with the best evidence currently available.
Live response can alter timestamps, process state, and logs. Rebooting may remove volatile evidence but may also be necessary to stop active harm. The decision should follow service criticality, attacker activity, and evidence needs.
Composite scenario
How this usually reaches the desk
Imagine the report begins with this observation: “Unexpected users, SSH keys, sudoers changes, services, timers, cron entries, kernel modules, or startup scripts.” A second check returns another clue: “Processes with deleted executables, unusual parent-child relationships, listening ports, or connections to unknown infrastructure.” Neither fact alone tells the whole story. Together they justify a documented incident, a defined owner, and a deliberate containment decision. This is where a timeline beats a collection of heroic memories; memory is an excellent storyteller and a dreadful audit log.
This scenario combines common operational patterns; it is not presented as a report of one named incident.
What to look for
Begin with preserved, comparable evidence. One signal is rarely proof; use independent observations and a reliable timeline before declaring scope or intent.
Unexpected users, SSH keys, sudoers changes, services, timers, cron entries, kernel modules, or startup scripts.
Record the exact time, source, identity, system and time zone. Compare it with a known-good baseline and with what the user or service owner expected. A surprising event is a lead, not a conviction.
Processes with deleted executables, unusual parent-child relationships, listening ports, or connections to unknown infrastructure.
Look for the control-plane event that made the visible activity possible: a changed credential, permission, rule, route, token or trusted device. Persistence often looks administrative because, technically, it is.
Authentication and audit log gaps, modified binaries, new setuid files, or security tooling disabled without a change record.
Correlate the report with independent telemetry before deciding scope. User testimony, identity logs, endpoint evidence and service audit records are strongest when they agree on sequence rather than merely on mood.
Working examples
Run it, read it, decide what changes
These examples use documentation addresses, test identities and bounded targets. Replace placeholders only inside systems you own or are explicitly authorised to operate. Read the expected result and next action before running the command; a successful command is evidence, not yet a conclusion.
Capture processes, listeners and sessions
- Prerequisites
- Root or sudo access and a protected output directory.
sudo sh -c 'date -u; ps auxww; ss -plantu; who -a; last -Fai | head -100' > incident-evidence/live-state.txt
sha256sum incident-evidence/live-state.txtOne text file records collection time, processes, sockets and recent sessions, followed by its hash.
The snapshot is volatile and incomplete. Rootkits can lie to local tools, while containers and namespaces can hide additional processes.
Preserve relevant journals and cloud/EDR telemetry, then compare unknown processes with packages, executable hashes and service definitions.
Review service and SSH changes in a time window
- Prerequisites
- A known UTC start time and sudo access.
sudo journalctl --since "2026-08-18 08:00:00 UTC" --until "2026-08-18 10:00:00 UTC" -u ssh --no-pager > incident-evidence/ssh-journal.txt
sudo find /etc/systemd /etc/ssh /root/.ssh /home -xdev -newermt "2026-08-18 08:00 UTC" -ls > incident-evidence/changed-control-files.txtSSH events and recently changed control files are saved.
File modification time can be altered and package updates create legitimate changes. Correlate with package, configuration-management and identity records.
Inspect unexpected keys, units and drop-ins on a copy; do not delete them before recording ownership, content and timestamps.
Verify packaged files
- Prerequisites
- Root access and knowledge of the installed package manager.
# Debian/Ubuntu: reports changed packaged files
sudo dpkg -V
# RHEL/Fedora/Rocky: reports changed packaged files
sudo rpm -VaNo output usually means no packaged-file differences; lines identify files whose recorded attributes differ.
Administrators legitimately edit configuration files, and attacker files outside packages are invisible to this test.
Compare deviations with configuration management and backups, then rebuild from trusted media when integrity cannot be justified.
What to do
Read the whole sequence before starting. Several workstreams may run in parallel, but their evidence, authority and expected outcomes still need to be explicit. Every step below points back to a concrete example; use the example as implementation evidence, not as permission to operate outside the stated scope.
Operational judgement
Containment and recovery are different verbs. Containment limits the next harmful action; recovery returns a service to trustworthy operation. Between them sits eradication: removing the access path and persistence that would make the freshly restored service merely a cleaner target. For Linux, Server and Forensics, keep those decisions separate in the timeline even if a small team performs them minutes apart.
Communication is also a control. Tell affected people what is known, what remains uncertain, what they must do and when the next update will arrive. Avoid both melodrama and false reassurance. “We are investigating” is useful only when followed by an owner and a time. The goal is to reduce secondary harm without teaching a possible attacker exactly what the team has discovered.
Handover
Make the result useful to the next person
Maintain two views of the incident. The working timeline should contain detailed events, evidence locations, hypotheses and technical actions. The stakeholder update should contain confirmed impact, current containment, material uncertainty, decisions required and the next reporting time. Do not copy speculative indicators into executive statements. Equally, do not polish away uncertainty merely because it looks untidy. Both records should use absolute times with a declared time zone and should identify the source of each important fact.
At shift change, hand over the current scope, trusted administration path, preservation status, active controls, failed actions, business priorities and the next three decisions. Read back the most consequential assumptions. For this case, ensure the record begins with “Declare scope and capture time, host identity, service role, network context, current users, and observable business impact.” and does not finish until the team has addressed “Hunt for the same indicators and access path elsewhere. Rebuild from a trusted source when system integrity cannot be established.” A concise, accurate handover prevents the incoming team from repeating disruptive work or mistaking a quiet telemetry gap for successful containment.
Validate before you close
Return the service only after credentials are rotated, initial access and persistence are addressed, trusted packages or images are used, and monitoring shows normal behaviour under load.
Capture the test, the expected result and the observed result. Where a person or business owner must accept restored service, name them in the record. A green dashboard can confirm that a component is answering; it cannot confirm that invoices, identities or restored data are trustworthy.
Finish with a compact closure note: the original trigger, confirmed scope, evidence retained, controls changed, tests passed, known gaps, residual risk, and the people responsible for the remaining work. Schedule a review while the timeline is still fresh enough to challenge. The purpose is not to find a person to blame; computers already perform blame with admirable efficiency. The purpose is to make the next response faster, safer and less dependent on one person remembering where the useful log was hidden.
Common mistakes
- Resetting systems before preserving volatile evidence and audit logs.
- Treating the first visible symptom as the complete scope of the incident.
- Restoring service without verifying that the attacker’s access path is closed.
These errors usually come from haste, unclear ownership or misplaced confidence. Build the safeguard into the runbook: a required evidence field, a second-person review, a rollback test or a specific exit criterion.
Questions people ask when the clock is running
Does one suspicious event prove compromise?
No. Treat “Unexpected users, SSH keys, sudoers changes, services, timers, cron entries, kernel modules, or startup scripts.” as a reason to investigate and preserve evidence. Confidence should rise when independent identity, service, endpoint or network records support the same sequence.
Should we reset everything immediately?
Reset or revoke what the evidence and risk justify, but preserve the state you will need to understand the incident. Broad, undocumented resets can interrupt the attacker, the business and the investigation in one impressively efficient stroke.
When can the incident be closed?
Return the service only after credentials are rotated, initial access and persistence are addressed, trusted packages or images are used, and monitoring shows normal behaviour under load. Closure also requires named owners for residual risk and follow-up work; “it seems quiet now” is an observation, not an exit criterion.
Safety boundary
Use these steps only on systems you own or are explicitly authorised to assess. Preserve evidence, follow your organisation’s legal and regulatory obligations, and prefer reversible actions when the situation is not yet understood.
Primary references
- SP 800-61 Rev. 3: Incident Response Recommendations and ConsiderationsNIST
- MITRE ATT&CK Enterprise MatrixMITRE
Editorial status: first edition. Review the linked vendor documentation for product- and version-specific changes before acting.