Your OpenClaw agent's server stopped on its own? Why, and the fix
A server your OpenClaw agent starts with & or nohup stops with its command, or at the exec tool's 30-minute limit. What we saw, and how to keep one running.
OpenClaw’s exec tool owns every process a command starts, and ends them with the command. A server your agent starts with &, with or without nohup, is stopped the moment its command finishes, since OpenClaw 2026.9.5. If it keeps the command’s output open, the command is moved to the background after 10 seconds instead, and everything it started is ended at the tool’s default limit of 30 minutes. To keep a server up, have the agent start it as its own background command: background: true, timeoutSeconds: 0, and the server itself as the command. That lasts until the Gateway restarts. To survive restarts, make it a service of the machine or a line in OpenClaw’s BOOT.md.
What we watched happen
On September 23, 2026, on OpenClaw 2026.9.4, we asked two agents on two machines for a one-page website and its link. Each built the page, served it with Python’s http.server on port 8080, checked it, and told its owner the site was live. Both answered at 12:41 UTC. At 13:53 one of the addresses returned 502 Bad Gateway, and it never came back. That machine had not rebooted, its Gateway had not restarted, and nothing had run out of memory. The program behind the address was gone.
That agent did everything in one exec call: write a start script, run it (the script started the server with nohup … &), wait two seconds, then fetch the page locally and from outside. Both fetches returned 200. Ten seconds after the call began, the tool’s result came back:
Command still running (session briny-falcon, pid 809)A subshell the script left behind, now the server’s parent, was still holding the command’s output open, so OpenClaw could not treat the command as finished and moved it into the background. A background command inherits the tool’s default limit of 1,800 seconds, and when it expires OpenClaw ends the command’s whole process group. The server was in that group. Its log holds requests up to 12:41 and no clean exit. The limit put the end at 13:04; nobody visited between 12:41 and 13:53, so the exact minute was not seen.
The other agent started its server with one line at the top level of its command, every stream redirected:
nohup ./start.sh >> start.log 2>&1 < /dev/null &Nothing held the output, the tool returned at once with launched pid=834, no timer applied, and a check from a separate command found the server running. That site kept serving. It survived on 2026.9.4’s leniency, which has since gone.
The next day we reproduced the first case on the same version with a one-minute limit: a sleep 600 started with nohup … & from inside a script, run through exec with timeoutSeconds set to 60. The tool reported it still running. The sleep was alive at 49 seconds and gone at 76.
Since 2026.9.5, a trailing & stops with the command
OpenClaw 2026.9.5, released September 19, changed shell backgrounding. After a command on the Gateway’s host finishes, OpenClaw now releases the command’s process group before it reports the result, and, in the words of the background-process docs, “Children left behind by shell backgrounding (&) are stopped with that group.” The pull request that made the change calls it intentional and says to keep long-running work in background: true mode “with the long-running command as the root.” It is unchanged in 2026.9.6, the current release on September 28.
So the second agent’s line, which kept a site up on 2026.9.4, now stops the server the instant its command returns. A check inside the same command still passes, because it runs before the command finishes, which is how an agent can report in good faith that a site is live when it is already gone.
Which one you hit
- Gone the moment the agent finished, even though it said it had checked: the server was started with
&, with or withoutnohup, in a command that then returned. That is 2026.9.5 and later. - Gone about half an hour after it started: the command never finished in OpenClaw’s eyes. Either something it started held its output open, as in our run on 2026.9.4, or the agent used
background: trueand left the timer on. Backgrounded commands inherittools.exec.timeoutSeconds, 1,800 seconds by default, and the exec docs say expiry “terminates the process even afterbackgroundoryieldMsreturns a session ID.” - Gone after an update, a restart or a reboot: background sessions are held in the Gateway’s memory and “lost on process restart”, and an update restarts the Gateway.
If your agent’s commands run in OpenClaw’s sandbox, which is off by default, the sandbox backend decides how long they live instead.
Why nohup does not help
nohup makes a program ignore one signal: the hangup a terminal sends when it closes. OpenClaw does not stop a command by closing a terminal. It signals the command’s whole process group, and nohup changes nothing about which group a program is in. Both servers in our run were started with nohup, and the one OpenClaw ended left no clean exit in its log. The exec tool’s own timeout message, in the 2026.9.6 source, says it directly: “Do not rely on shell backgrounding with a trailing &.”
Keep it running: start it as its own background command
The documented pattern for “a persistent service on the gateway” is an exec call with background: true and timeoutSeconds: 0, where the server itself is the command, running in the foreground of that call. The server is then the process OpenClaw tracks. It keeps running after the turn ends, the agent can see it with process list and stop it with process kill, and if it crashes, OpenClaw’s completion notice reaches the agent (in our run, at its next heartbeat; more on that below). Tell your agent in so many words, because its habit is the shell’s:
Start the site as its own background command: exec with background set to true and timeoutSeconds set to 0, with the server itself as the command, running in the foreground. No &, no nohup, no script that backgrounds it. Then check the address from a separate command a minute later.The call it should make looks like this:
{
"tool": "exec",
"command": "cd ~/site && python3 -m http.server 8080",
"background": true,
"timeoutSeconds": 0
}Set the timeout per call, not by raising tools.exec.timeoutSeconds for everything: the default limit is what ends commands that hang. And run the check as a separate command after the first has returned, since a check inside the starting command passes before OpenClaw has cleaned up.
This lasts as long as the Gateway does. The exec docs say that disabling the timeout “does not make the process survive its host or worker shutting down,” and a Gateway restart loses the session.
Make it come back after a restart
A server that should still be there tomorrow needs something outside the agent’s command to start it. There are two ways, and which one fits depends on how OpenClaw runs.
A service of the machine
This is how OpenClaw keeps its own Gateway running: a user systemd unit on Linux, a LaunchAgent on a Mac. A program started this way is started by the service manager, not by any agent command, so no exec cleanup reaches it. On Linux, a user unit for the site, written by you or by the agent:
# ~/.config/systemd/user/site.service
[Unit]
Description=Agent's website
[Service]
WorkingDirectory=%h/site
ExecStart=/usr/bin/python3 -m http.server 8080
Restart=always
[Install]
WantedBy=default.targetsystemctl --user daemon-reload
systemctl --user enable --now site.serviceRestart=always brings it back if it crashes, and enable starts it at boot. On a server nobody logs into, user services also need lingering, which loginctl enable-linger turns on so that “a user manager is spawned for the user at boot and kept around after logouts.” OpenClaw’s onboarding already attempts that for the user its Gateway runs as. On a Mac the equivalent is a LaunchAgent with KeepAlive. This is the ordinary way a computer keeps a program up. If OpenClaw runs in a Docker container, there is usually no service manager inside it, and the lasting place for a server is a service of its own in your Compose file, or the next route.
OpenClaw’s startup hook
The bundled boot-md hook runs a workspace’s BOOT.md as an agent turn rather than a script, and, per the BOOT.md reference, “once per agent workspace every time the gateway starts.” It ships disabled, so enable it first with openclaw hooks enable boot-md. A line in BOOT.md then has the agent start the server the right way on every start:
If nothing answers on port 8080, start the site with exec: background true, timeoutSeconds 0, command "cd ~/site && python3 -m http.server 8080".Each Gateway start then costs one model turn, and the bundled-hooks reference asks for boot instructions that are “short and safe to repeat on every restart.” Nothing in it depends on a service manager, so it serves a Gateway in a container too.
Why nobody told you
OpenClaw did tell the agent. In our run, the harness’s completion notice, a message marked [OpenClaw exec completion], reached the agent at its next heartbeat, at 13:21. The agent read it, judged there was nothing its owner needed to hear, and gave the heartbeat’s silent reply, NO_REPLY. The owner found out by opening the link. In the 2026.9.6 source, the notice for a killed command reads like Exec failed (<id>, signal SIGTERM) followed by the end of its output, the same line another owner reported in February, when their background jobs kept dying at the half-hour mark.
The evidence fades quickly. The process list keeps a finished session for 30 minutes by default (tools.exec.cleanupMs), in memory only, so by the time you notice, process list may show nothing. Ask the agent how it started the server and whether it received an exec completion notice for it; the command and the notice are usually still in its conversation. If you want to hear next time without relying on how the agent weighs a notice, give it a scheduled job that fetches the address and messages you when it does not answer.