← All field notes

Our Server Got a Second Job: The Fourth of July Cryptominer

How a seven-month-old framework bug turned 12.5 of our CPU cores into someone else's Monero rig, how we tracked it down without root, and why starting over felt more manageable with AI.


On the Fourth of July 2026, one of our servers was celebrating independence in its own way: it had stopped working for us.

The slightly embarrassing part: I'm a security engineer. I work on vulnerability management, help teams deal with security findings, and spend plenty of time thinking about how attackers get into applications. Apparently, none of that stopped my own server from picking up a side hustle mining Monero for a stranger. My job title was not an effective security control.

So yes, even a security engineer can get pwned. Knowing that you should patch something and making sure every side project actually gets patched are two different things. This is a story about that gap, what we could learn from the evidence, and the uncomfortable questions that remained once we found the miner.

The complaint that started it all was boring. "coolify4 is at 100% CPU." That box is a 32-vCPU Coolify host running about 24 app containers, so a busy afternoon is normal. A load average of 70 is not.

The host has 32 vCPUs and an observed load average around 70. Load average includes runnable and uninterruptible tasks; it is not a CPU percentage.
The host has 32 vCPUs and an observed load average around 70. Load average includes runnable and uninterruptible tasks; it is not a CPU percentage. Scroll sideways on smaller screens. Open full-size chart.

The load average was more than twice the vCPU count, which was a reason to investigate—not a measurement of CPU utilization. Linux load average includes runnable tasks and tasks in uninterruptible sleep, often waiting on I/O. Something was putting serious pressure on this machine. (Linux load-average documentation)

Finding it, with no sudo

The fun constraint: the account we logged in with had no sudo and wasn't in the docker group. No docker ps, no docker exec, no peeking into container filesystems. Just a regular user and whatever Linux shows a regular user.

That turned out to be plenty. A plain process listing had one line that stood out:

/tmp/batch5 -o xmproxy.scrap-transport-musical-hospital-brainstorm[.]com:10031 -u batch3 -k -p menudo

CPU usage: about 1,250%. That is twelve and a half cores, flat out.

If you have never met this command line before, here is the translation. A binary in /tmp with -o (pool address), -u (user), -p (password) and -k (keepalive) is consistent with XMRig, a Monero miner commonly repurposed for cryptojacking. Those flags are a clue, not a substitute for identifying the binary. The pool hostname looks like someone rolled four words on a passphrase generator. The password is "menudo". We have questions, but not for this post.

Observed miner usage was equivalent to 12.5 of 32 cores. The other 19.5 core-equivalents represent remaining capacity, not a measurement of idle CPU.
Observed miner usage was equivalent to 12.5 of 32 cores. The other 19.5 core-equivalents represent remaining capacity, not a measurement of idle CPU. Scroll sideways on smaller screens. Open full-size chart.

Which app was it?

Three clues, all readable by an unprivileged user:

  1. The cgroup. The miner's cgroup path contained a Docker container ID, so the process lived inside a container and we knew which one.
  2. The container's IP. We curled it over the Docker network. The response had X-Powered-By: Next.js and a page title that matched one of our apps, a scheduling tool.
  3. The parent process. The miner's parent PID was next-server. The app's own Node process had spawned it.

That third clue matters most: it ties the miner to the application process. Combined with the vulnerable framework version and the command output in the logs, it points toward application-level code execution. The parent process alone does not rule out every other route into the container.

The likely entry point: React2Shell

The app was pinned to next@15.1.4, an affected App Router version in the React2Shell advisory. React tracks the RCE as CVE-2025-55182; the Next.js advisory describes its downstream impact under CVE-2025-66478. The vulnerability was rated CVSS 10.0:

  • It is a deserialization bug in the "Flight" protocol that React Server Components use.
  • Affected React Server Components applications can be vulnerable even without explicitly implementing their own Server Function endpoints.
  • An unauthenticated attacker can use a crafted request to trigger remote code execution in an affected deployment.

We checked whether our own code had helped. A repo-wide search for child_process, exec, spawn, eval, new Function and friends returned zero matches. We found no direct shell-execution calls in that search. That does not prove our code was flawless; it shows why searching only for dangerous calls in our own code would have missed this framework-level exposure.

The app's logs told the rest of the story. Mixed in with normal output were the results of commands we never wrote: id (answer: uid=1001(nextjs)), then recursive ls sweeps across /apps, /workspace, /usr/src/app, /var/www and /home. Around them sat the mess a malformed Flight payload leaves behind: TypeError: Cannot read properties of undefined (reading 'r'), redirect errors with shell output stuffed inside, and a flood of ERR_HTTP_HEADERS_SENT.

The vulnerable Next.js application and command output point toward React2Shell. The diagram distinguishes that inferred entry point from observed miner activity and the need to treat readable secrets as exposed.
The vulnerable Next.js application and command output point toward React2Shell. The diagram distinguishes that inferred entry point from observed miner activity and the need to treat readable secrets as exposed. Scroll sideways on smaller screens. Open full-size chart.

The timeline

Container startup July 2 at 02:01; miner starts at 03:17; second binary appears July 3 at 22:58; investigation begins July 4 at 12:54. Times are as recorded in the incident notes; their timezone was not specified.
Container startup July 2 at 02:01; miner starts at 03:17; second binary appears July 3 at 22:58; investigation begins July 4 at 12:54. Times are as recorded in the incident notes; their timezone was not specified. Scroll sideways on smaller screens. Open full-size chart.
When What
Thu Jul 2, 02:01 Server reboots. All ~24 containers come back up.
Thu Jul 2, 03:17 Miner starts. That is 76 minutes after boot.
Fri Jul 3, 22:58 A second binary lands in /tmp. Its name is crude enough that we are calling it /tmp/[redacted].
Sat Jul 4, 12:54 Someone asks why the CPU is at 100%.

Two numbers from that table deserve a second look.

76 minutes. That was the gap between the recorded container startup and the miner process starting. It is consistent with opportunistic exploitation, but the timestamps alone do not prove who targeted us or when the initial compromise occurred. Existing persistence is another possibility we had to investigate.

About 57.6 hours. That is the interval between the recorded miner start and discovery. If it sustained the observed 12.5-core usage throughout, that would be roughly 720 core-hours of compute for a stranger. That is an estimate, not a continuous CPU measurement.

The miner was the good news

A cryptominer is the loudest thing an attacker can run. It is the reason we noticed at all. The quieter findings were worse:

  • The second-stage binary. It appeared almost two days after the miner, suggesting continued malicious activity or persistence. The timestamp alone does not establish whether an attacker returned interactively.
  • Attacker tooling in the app's home directory, files named safenet-client-alpine-amd64 and safenet_vvz.
  • The secrets. Anything the app could read, the attacker could read: database passwords, JWT secrets, service-role keys, SMTP credentials, API keys. Several apps on that host shared secrets, so the blast radius was bigger than one container.

And because we had no root during the investigation, some questions stayed open: was there persistence on the host itself, and had anything moved sideways into other containers? When you cannot rule those out, the honest answer is to treat the whole host as untrusted.

Bonus round: things we found while looking

Incident response has a way of turning over other rocks. In the same app's repo we found:

  • a GitHub personal access token committed inside package.json, in the repository URL,
  • production Postgres credentials hardcoded in a migration script, with SSL turned off,
  • a CI workflow on a self-hosted runner that inherits secrets.

We did not establish these as the entry point for this incident. The committed credentials were additional exposures, and the runner workflow deserved a separate review of which jobs could reach its secrets.

The response plan

These are the response steps the findings called for. For the most part, we chose to start over and rebuild; I describe that below. This checklist is the plan, not a claim that every item was independently verified as complete.

  1. Contain and preserve evidence. Isolate the workload and collect relevant logs, process details, and volatile evidence where feasible, then stop the compromised container. Do not delete the container or its volumes before deciding what to preserve. Both docker stop and docker kill signal processes; the choice alone does not preserve evidence.
  2. Do not just redeploy. The same vulnerable image carries the same bug. Restarting it does not close the entry point, and the 76-minute observation is not a guaranteed window of safety.
  3. Patch and rebuild from trusted sources. Next.js 15.1.9 was the original RCE fix on the 15.1 line; 15.1.11 addressed the December follow-up issues. Those are historical fixes, not a current upgrade target. Use a currently supported release with the latest applicable security patches and rebuild from a clean image.
  4. Rotate everything the container could see, plus anything shared with other apps, plus the committed credentials from the bonus round.
  5. Hunt for persistence: crontabs, systemd units and timers, authorized_keys, /tmp, /dev/shm, LD_PRELOAD.
  6. Block the pool at the egress, then review what every other container is allowed to talk to.
  7. Audit the neighbours. Plenty of other apps on that host run Next 15.x and 16.x.
  8. If host compromise cannot be ruled out, rebuild the host.

What happened afterwards: mostly, we started over

For the most part, we started over and rebuilt things. Using AI made that process feel less difficult than it would have been previously. Rebuilding was still work, but it felt much more manageable with help from the tools I'd already been using for my side projects.

That is another interesting part of this experience. The missed patch got us into trouble, and AI didn't change that. But having help with the rebuilding process made starting over feel like a more practical option. Even after an incident, the amount of work I can reasonably take on is changing.

What we learned

1. "100% CPU" is a symptom, never a diagnosis. The first question is "which process?", and it takes one command to answer. We could have answered it 57 hours earlier.

2. Your framework is your attack surface. We found no direct shell-execution calls in our code, but that did not remove the framework exposure. The fix had been public for about seven months when we got hit, and one pinned version in one package.json was enough.

3. Being small does not make you invisible. There is no grace period for an unpatched public endpoint, and being small does not hide you. Scanners do not check your traffic numbers first.

4. A restart is not a fix. Killing the miner or bouncing the container treats the symptom. Until the version changes, the door is open.

5. Be grateful for noisy attackers, and assume a quiet one came too. The miner got caught because it was greedy. The same access could have been used to copy our secrets and leave, and we would have learned nothing from a CPU graph.

6. Containers are a boundary, not proof of containment. The app ran as a non-root user inside a container, which reduced its privileges. That did not protect the secrets the app could read, and it did not establish that the host or neighboring containers were untouched.

7. Outbound traffic deserves rules too. A web app has no business opening a connection to a mining proxy on port 10031. Default-deny egress would have left the miner with nobody to talk to.

8. Alert on sustained CPU. A box pinned for two days should page someone long before a human gets curious.

9. Know what you run. With 24 containers on one host, "which of these are on a vulnerable Next.js?" should be a query, not a research project.

10. Secrets do not belong in git. Not in package.json, not in a migration script, not "just for now".

11. You can do real forensics without root. Process listings, cgroup paths, a curl to a container IP and the app's own logs were enough to identify the container, the app, the parent process, and a likely vulnerability to investigate.

Yes, the security engineer got pwned

There is some humor in spending your working day helping other people manage vulnerabilities, then discovering that your own infrastructure has been donating compute to an attacker. I can laugh at the irony. The exposed secrets and unanswered questions about the scope of the compromise are harder to laugh off.

The useful part of telling this story is owning both. Security experience helped us investigate, but it didn't make the missed patch harmless or the detection fast enough. My side projects still need the same follow-through I would ask of any other team: know what's running, keep it patched, limit what it can access, and notice when it starts behaving strangely. Being the security person doesn't let me skip any of that.

By the numbers

Metric Observation
vCPUs on the host 32
Load average at discovery ~70
Cores taken by the miner ~12.5 (39%)
Recorded container start to miner start 76 minutes
Miner start to discovery ~57.6 hours
Estimated compute at sustained observed usage ~720 core-hours
Age of the fix when we got hit ~7 months
Vulnerable framework version found Next.js 15.1.4

Indicators, if you want to check your own boxes

  • A next-server process with children running out of /tmp
  • Binaries in /tmp taking XMRig-style flags (-o, -u, -p, -k)
  • Outbound connections to xmproxy.scrap-transport-musical-hospital-brainstorm[.]com:10031
  • Unexpected files in the app user's home directory, such as safenet-client-alpine-amd64
  • App logs with NEXT_REDIRECT digests containing shell output, TypeError: Cannot read properties of undefined (reading 'r') or (reading 'd'), and bursts of ERR_HTTP_HEADERS_SENT

These are investigation leads, not standalone proof: some errors can also occur in ordinary application failures. If the process behavior and logs look familiar, investigate the workload and check its framework version before you finish your coffee.

Thanks for reading. These are my personal notes and opinions.