LuanMartins.com
← log

My own S3 with Garage on Coolify

Self-hosting · Infrastructure11 min

I was reorganizing my personal site, one of the projects I'm picking back up to make it feel like mine and to keep a log of everything I do. As part of that, I ran it through an AI to review the content and suggest improvements.

One of the points it raised stung. My site's pitch says I host and run my personal projects on a VPS, and the link to my resume points at a Google Drive. That made no sense: I have a server up and running, and no S3 of my own to hold a PDF.

So I decided to build my own S3 with Garage. I looked at several options first, and this one won for having good reviews, decent documentation, and a ready-to-run app in Coolify's catalog. Installing it was easy. Getting it to work came with a few discoveries I wasn't expecting.

Why Garage

The obvious pick would have been MinIO, the name that was synonymous with "self-hosted S3" for years. That reflex is out of date. The community edition was dismantled piece by piece: the admin console went away, the binaries and Docker images stopped being published, and in the end the repository was archived. The template still sits in plenty of app catalogs, which makes it easy to install software today that no longer gets security patches.

Garage is an open source project by Deuxfleurs, a French non-profit collective. It's written in Rust, ships as a single binary with no external dependencies, and, most important for me, has a native ARM64 build, which is what my Oracle VPS runs on. It became the most common recommendation in self-hosted communities after MinIO's demise.

It's genuinely light, runs comfortably on the 2 cores I have on the VPS, speaks S3 with SigV4 signing, and stores data with compression and deduplication. I found two limitations: no object versioning, and no official web console. Neither matters for my case, which is serving portfolio assets, pipeline uploads, and backups.

The Garage logo

The Coolify template

In fairness, installing Garage on Coolify really is simple. It's in the one-click catalog, and the template generates the variables plus a garage.toml as a file mount under the Persistent Storage tab. Deploy, container comes up.

Except it didn't work on the first try. The container came up, went green on the dashboard, and no request got a useful answer back. It took a few trips to the log to find the real cause, and along the way I almost fixed something that wasn't broken.

The wrong suspect

The first thing that looked wrong to me was the garage.toml. The template's three secret lines came like this:

rpc_secret_file = "env:GARAGE_RPC_SECRET"

[admin]
admin_token_file = "env:GARAGE_ADMIN_TOKEN"
metrics_token_file = "env:GARAGE_METRICS_TOKEN"

I went to the Garage documentation and that env: scheme doesn't exist there. What exists is rpc_secret taking the value directly, rpc_secret_file taking a path to a file on disk, or the GARAGE_RPC_SECRET environment variable carrying the value. Everything pointed to a broken template.

Digging further, the file was fine. It's generated by Coolify itself, and the actual secrets come in through the environment variables the template already passes in the compose file, GARAGE_RPC_SECRET, GARAGE_ADMIN_TOKEN and GARAGE_METRICS_TOKEN, which are exactly Garage's native mechanism. That's also why the template sets GARAGE_ALLOW_WORLD_READABLE_SECRETS=true. Those odd lines were never the problem, and I didn't touch them.

The real error was somewhere else.

Layout not ready

The container was up, the servers were listening on the right ports, and every request to the web endpoint came back with this:

GET 500 Internal Server Error web-garage.yourdomain.com/
API error: Internal error: Layout not ready

Garage is not like MinIO here. It refuses every write until you define the cluster layout, which is the assignment of a zone and a capacity to each node. Even when the whole cluster is a single node. Coming from MinIO, you start the container and expect it to just work; here, a running container is only half the way there.

On older versions, the layout is done by hand:

garage layout assign -z dc1 -c 40G <NODE_ID>
garage layout apply --version 1

The template shipped with the v2.1.0 image. Starting at v2.3.0 there's a --single-node flag that creates the layout on its own at startup. Since the upgrade path has no breaking changes in between, I bumped the version and used the flag. That ended up being the only change I made to the template:

services:
  garage:
    image: 'dxflrs/garage:v2.3.0'
    command:
      - /garage
      - server
      - '--single-node'

After the redeploy, the log showed the layout being created:

INFO garage::server: Created initial layout for single-node configuration:
Partitions are replicated 1 times on at least 1 distinct zones.
  0b4ebababb34909a  [default]  256 (256 new)  192.7 GiB

And the 500 turned into a 404, which is the correct error for "this host doesn't match any bucket". Debugging with progress measured in status codes.

Disk capacity is just a weight

The --single-node flag assigned 192.7 GiB to the node, the whole disk. My first instinct was to lower that number to save room for the other services running on the same VPS.

Wrong instinct. The node's capacity is not a limit, it's a distribution weight: the layout algorithm uses it to decide how many partitions each node gets. The thing that enforces a size limit is the filesystem. With a single node, the number is decorative: all 256 partitions land on it anyway, whether it says 192 GiB or 40. Lowering it would have protected nothing.

What actually protects you comes down to two things. Per-bucket quotas, the only real enforcement:

garage bucket set-quotas --max-size 40GB <bucket>

And disk monitoring, which Coolify already has, with a configurable threshold that fires a notification. The risk is not theoretical: the volumes live in /var/lib/docker/volumes, on the same filesystem as Coolify and any database I spin up later. A full disk takes everything down with it.

The bucket name is the domain

Since the goal was to serve my resume publicly through a URL, the next step was creating the bucket and enabling the website on it. But I needed a public bucket.

And it was at creation time that the strangest discovery of the whole process showed up: the bucket couldn't be named just anything. Garage's web endpoint, the one on port 3902, decides which bucket to serve by looking at the request's Host, in one of two ways:

  1. If the host ends with the root_domain configured under [s3_web], the prefix is the bucket name: mybucket.web.yourdomain.com serves the bucket mybucket.
  2. Otherwise, it looks for a bucket whose alias is the entire hostname.

The first path is virtual-hosted style, and it needs a wildcard certificate, which on Coolify means setting up a DNS challenge in Traefik. It wouldn't even work the way the template ships, by the way: the root_domain is generated as .web.garage.localhost, which will never match a real domain. The second path works with a plain Let's Encrypt certificate. I went with the second, and the consequence is that the bucket has to be named exactly like the domain. Hence bucket create web-garage.yourdomain.com.

Which means every public bucket needs its own domain. For two or three, fine; beyond that it's worth investing in the wildcard. Private buckets have no such restriction and can be called whatever you want.

For the same reason, on the S3 API I use path style: addressing_style = path in the aws CLI, forcePathStyle: true in the JavaScript SDK. Without it, the client builds bucket.s3-garage.yourdomain.com and the certificate doesn't cover it.

The padlock was missing. Https worked, but hitting the site over http served the page in plain text without redirecting. Coolify decides this based on the scheme you type into the domain field: without the https:// in front, it doesn't generate the redirect middleware in Traefik. I edited the domain to https://web-garage.yourdomain.com and hit Redeploy. It has to be a deploy: the Traefik labels are only rewritten there, a restart isn't enough.

The image has no shell

When it was time to create the buckets, I opened Coolify's web terminal and got:

Terminal Not Available
No shell (bash/sh) is available in this container.

The Garage image is distroless: the static binary and nothing else. No shell, no coreutils. Great for size and attack surface, bad for tools that assume there's a sh on the other end.

But I didn't need a shell, I needed to run the binary. Over SSH on the host:

docker exec <container> /garage bucket create web-garage.yourdomain.com

No -it, no shell invoked. It works. And it's the same reason the template's healthcheck uses the list form, with CMD calling /garage stats -a directly, instead of CMD-SHELL: there's no shell in there to parse a command line.

The security model

Before the commands, it's worth understanding the structure, because it's simpler than AWS's, and knowing that saves you time hunting for things that don't exist.

Three ports, each with its own credential:

PortWhatCredential
3900S3 API, objectssigned key (SigV4)
3902web endpointanonymous, only buckets with the website enabled
3903Admin APIbearer token

And the permission hierarchy has four levels, except the last one doesn't exist:

  1. RPC secret and admin token: cluster level, full control.
  2. Access key (GK... + secret): an identity, born with access to nothing.
  3. Per-bucket permission: the key gets read, write or owner on each bucket, one by one.
  4. Object: nothing. There is no control at this level.

Garage doesn't implement AWS-style ACLs or bucket policies. It's a table of key, bucket and permission, and that's it. There's no such thing as a public object inside a private bucket, which means a bucket with the website enabled is entirely public. Don't mix sensitive content into it.

Anonymous access through the S3 API simply doesn't exist, and you can watch that in the log as soon as the domain goes public, with the internet's scanners taking the hit:

GET / → 403 Forbidden: Garage does not support anonymous access yet

The buckets I created

The rule I followed: a bucket is a trust boundary, a prefix is organization. A new bucket is only worth creating when the answer changes for three questions: who reads, who writes, what's the quota.

I ended up with three. web-garage.yourdomain.com, public, for the portfolio assets. A private one for pipeline uploads. And a private one just for backups. That last one being separate is not organization, it's containment: if the application's key ever leaks, the backups are out of its reach.

The setup, once. The <container> in the commands is the container's real name, which you can find with a docker ps on the host, looking for garage in the list:

# creates the bucket, empty and private
docker exec <container> /garage bucket create web-garage.yourdomain.com

# enables anonymous reads through the web endpoint (writes still require a key)
docker exec <container> /garage bucket website --allow web-garage.yourdomain.com

# check it: look for "Website access: true"
docker exec <container> /garage bucket info web-garage.yourdomain.com

# creates an identity (born with access to nothing)
docker exec <container> /garage key create laptop

# connects identity and bucket
docker exec <container> /garage bucket allow --read --write web-garage.yourdomain.com --key laptop

The key create answers with the key's Key ID and secret.

And notice I didn't give the key --owner. That's on purpose: without it, a leaked key can't turn off the bucket's website or remove the quota.

Connecting the client

From here on, no more SSH. On the laptop, the aws CLI handles it with a profile:

aws configure --profile garage
# Key ID, Secret, region = garage, output format blank

aws configure set endpoint_url https://s3-garage.yourdomain.com --profile garage
aws configure set s3.addressing_style path --profile garage

The Key ID and secret it asks for are the ones from the laptop key. I hadn't copied them at creation time, and it wasn't a problem: a /garage bucket info web-garage.yourdomain.com shows the bucket's information, including the keys connected to it, and that's where I picked up what aws configure wanted.

The endpoint_url in the config file needs aws CLI 2.13 or newer. On older versions, it's --endpoint-url on every command.

The test:

aws s3 ls --profile garage

And the quick diagnosis, for when it doesn't work:

  • SignatureDoesNotMatch is a wrong secret.
  • A timeout is a wrong endpoint.
  • AccessDenied is a missing bucket allow.

The resume, at last

This whole thing started because of a PDF on a Google Drive. So the closing move is taking it out of there:

aws s3 cp resume.pdf s3://web-garage.yourdomain.com/assets/resume.pdf \
  --content-type application/pdf \
  --content-disposition 'attachment; filename="Luan_Martins_Resume.pdf"' \
  --profile garage

The --content-disposition is what makes the link download the file instead of opening it in the browser. It has to go in at upload time, because it becomes the object's metadata. If the file is already up there, you can rewrite it by copying the object over itself with --metadata-directive REPLACE, without downloading anything.

The obvious solution, HTML's download attribute, doesn't help here: it only works for same-origin links, and the bucket lives on another subdomain.

The resume link on my site now points at my own server. The inconsistency the AI called out at the start is gone.

The final configuration

For whoever just wants the result. The garage.toml stays the way the Coolify template generates it, with no real secret inside, because they come in through the environment variables:

metadata_dir = "/var/lib/garage/meta"
data_dir = "/var/lib/garage/data"
db_engine = "lmdb"

replication_factor = 1
consistency_mode = "consistent"

compression_level = 1
block_size = "1M"

rpc_bind_addr = "[::]:3901"
rpc_secret_file = "env:GARAGE_RPC_SECRET"
bootstrap_peers = []

[s3_api]
s3_region = "garage"
api_bind_addr = "[::]:3900"
root_domain = ".s3.garage.localhost"

[s3_web]
bind_addr = "[::]:3902"
root_domain = ".web.garage.localhost"

[admin]
api_bind_addr = "[::]:3903"
admin_token_file = "env:GARAGE_ADMIN_TOKEN"
metrics_token_file = "env:GARAGE_METRICS_TOKEN"

The compose file, with the one change I made, the image and the command:

services:
  garage:
    image: 'dxflrs/garage:v2.3.0'
    command:
      - /garage
      - server
      - '--single-node'
    environment:
      - GARAGE_S3_API_URL=$GARAGE_S3_API_URL
      - GARAGE_WEB_URL=$GARAGE_WEB_URL
      - GARAGE_ADMIN_URL=$GARAGE_ADMIN_URL
      - 'GARAGE_RPC_SECRET=${SERVICE_HEX_64_RPCSECRET}'
      - GARAGE_ADMIN_TOKEN=$SERVICE_PASSWORD_GARAGE
      - GARAGE_METRICS_TOKEN=$SERVICE_PASSWORD_GARAGEMETRICS
      - GARAGE_ALLOW_WORLD_READABLE_SECRETS=true
      - 'RUST_LOG=${RUST_LOG:-garage=info}'
    volumes:
      - 'garage-meta:/var/lib/garage/meta'
      - 'garage-data:/var/lib/garage/data'
      -
        type: bind
        source: ./garage.toml
        target: /etc/garage.toml
    healthcheck:
      test:
        - CMD
        - /garage
        - stats
        - '-a'
      interval: 10s
      timeout: 5s
      retries: 5

And the domains in Coolify, both with https:// in front: s3-garage.yourdomain.com on port 3900, web-garage.yourdomain.com on port 3902, and port 3903 with no public domain at all.

What I'd do differently

Read the whole log before hunting for a culprit. I almost "fixed" a garage.toml that was correct. The container was up, the S3 API server listening lines were there from the start, and so was the Layout not ready. The answer was in the log the entire time; I'm the one who chased an assumption before reading to the end.

Separate noise from signal. Once the domain went public, the scanners started hammering non-stop, and the healthcheck adds three lines every 10 seconds. The most confusing part: each healthcheck connection shows up with a different ID, looking like new nodes joining the cluster. It's not: it's the CLI generating an ephemeral identity on every invocation.

Set up the wildcard from the start if the plan is several public buckets. Bucket names can't be renamed. Migrating later means creating a new bucket and syncing everything over.

Was it worth it

This all started with an AI pointing out a silly inconsistency on my site. It ended with my own object storage, with quotas, keys scoped per bucket, and my resume served from my server, on my domain.

For the time it cost, yes. Garage itself gave me less trouble than I expected: what held me up was an outdated template and a concept I didn't know, not bugs. And what's now standing serves far more than one PDF: the images on this log, my pipeline uploads and my backups all have somewhere to go without depending on a third party.

Three things are still pending, and at least one of them becomes a post: the wildcard certificate, to unglue bucket names from hostnames; garage-webui, the community-built web console, to retire SSH for good; and the buckets, quotas and website config described in Terraform, since there's a provider for it.