Reading Container Logs Like an SRE: The Ten Failures Behind Most Failed First Deploys
Most failed first deploys are one of a handful of problems, and the log says which one in the first or last few lines. This guide is how to read the two logs a container host gives you, what each of the common failure messages means, and the fix for each. Every example here is a real category we see on SnapDeploy; the fixes are the same on any container platform.
Two logs, read in opposite directions
- The build log is your Dockerfile executing, step by step. Read it top-down and stop at the first line that says
ERRORor a non-zero exit. Everything after it is fallout. - The runtime log is your process's stdout and stderr once the container starts. Read it bottom-up: the last lines before the exit are the cause. The exit code is part of the story:
0means the process finished on its own,1is an unhandled error,137is the kernel killing it for memory.
When a deployment fails, SnapDeploy puts a plain-language reason at the top of it, based on what it found in those logs. The rest of this post is that list, with the fix for each.
1. "Your app is listening on the wrong port"
The container started, the process is running, but nothing answers on the port the platform routes to. Either the app hard-codes a port that differs from the container's setting, or it binds to localhost/127.0.0.1 instead of 0.0.0.0, or it listens on IPv6 only. The message names both ports when it can. Fix: read PORT from the environment, bind to 0.0.0.0, and make EXPOSE in the Dockerfile match. A local docker run -p 8080:8080 followed by curl localhost:8080 reproduces it in ten seconds.
2. "Built for a different CPU architecture"
The runtime log shows exec format error. You deployed a prebuilt image that only exists for x86, and the container runs on ARM64. Fix: use the image's arm64 tag if it has one, build a multi-architecture image with docker buildx build --platform linux/amd64,linux/arm64, or deploy the repository instead so the platform builds it for the right architecture. Official base images are multi-architecture already; this bites with smaller third-party images.
3. "Shared library not found"
A line like error while loading shared libraries: libssl.so.1.1: cannot open shared object file. A native dependency was compiled against a library the runtime image does not have, or has in a different version. The classic case is Prisma's engine on an Alpine image expecting OpenSSL. Fix: use a Debian-slim base (node:22-slim, python:3.12-slim) for anything with native modules, or install the library in the runtime stage (apk add --no-cache openssl). Multi-stage builds must copy native modules from a stage with the same base.
4. "Permission denied" on the entrypoint
The start command points at a script that is not executable, or the USER line switched to a user who cannot read the file. Fix: chmod +x entrypoint.sh before committing (Git tracks the bit), or call it through the interpreter (CMD ["sh", "entrypoint.sh"]), and make sure files copied before USER are owned or readable by that user (COPY --chown=app:app).
5. "The start command points at a file that does not exist"
The most common wrong-start-command failure: python: can't open file 'app.py', Cannot find module '/app/index.js', no main manifest attribute. A generated Dockerfile guessed the entry point and guessed wrong, or your own CMD has a path that was true on your laptop. Fix: add a Dockerfile with the correct CMD (or the correct scripts.start in package.json), and check that the file is not excluded by .dockerignore.
6. "Database connection refused" or "timed out"
The app started and immediately tried to reach a database that was not there: an empty DATABASE_URL, a localhost address from development, an external database that pauses when idle (Neon, Supabase) and has not woken yet, or an Atlas cluster whose network access list does not allow the container. Fix: check the variable in the container settings, attach the add-on if you meant to, allow connections from anywhere on Atlas, and make the app retry the first connection a few times with a short timeout rather than exiting at once.
7. "The container exited with code 0"
Nothing crashed. The process ran to the end and quit, which for a web service means it never started a server: a script that prints and exits, a Node file that forgot app.listen, a build command used as the start command. Fix: make the start command the thing that listens. If the container is meant to be a one-shot job, this platform expects a server; run jobs from inside a service or from an external scheduler.
8. "Killed" with exit code 137
Not an error in your code; the container used more than its 512 MB and the kernel stopped it, often with no stack trace. Fix: find what grew (a JVM heap sized for a laptop, four Gunicorn workers, an in-memory cache without a bound, an image resize on a 20 MB upload), reduce it, or move to Always-On Medium (2 GB). The language guides on this blog each have a memory section for their runtime.
9. "Segmentation fault"
A native module crashed, usually because it was built for a different platform than the one it runs on: an npm install on macOS committed with node_modules, or a wheel for the wrong architecture. Fix: never commit node_modules, install dependencies inside the build, and let the build produce the native pieces for ARM64.
10. "Did not become healthy in time"
The process is running and no other category matched. Usually the app takes longer to start than the window allows (a JVM on a quarter CPU, a framework running migrations against a slow database), or a health path returns 500 while the app is still warming up. Fix: answer the health check as early as possible with a route that does not depend on heavy initialisation, cut start-up work, and check the previous nine categories in the runtime log; sometimes the real reason is there and the health timeout is only the symptom.
A habit that avoids most of this
docker build -t app .
docker run --rm -p 8080:8080 -e PORT=8080 --env-file .env.local app
curl -i http://localhost:8080/health
docker logs <container> # in another terminal, if it exitedEight of the ten categories above reproduce on your machine with those four commands, at no cost to the free tier's deploy allowance (5 per rolling 12 hours, failed builds included). The two that do not, architecture and memory, are the ones to think about when the local run passes and the deploy still fails.
Where to look: every container's page has the build log and the runtime log; a failed deployment shows its reason at the top. Free tier limits are on the free tier docs page; environment variables on the environment variables docs page.