S3 Compatible Storage: What It Is and How to Choose
September 2026 · 26 min read · Surya
S3 compatible storage is object storage whose provider implements the Amazon S3 REST API well enough that your existing S3 client, SDK or CLI talks to it without code changes. That is the whole definition, and almost every argument about it comes from people quietly using a looser one, where compatible means the provider has an S3-shaped button in its dashboard and a blog post that says the word S3 a lot.
We build a control plane for s3 compatible storage, so we have a commercial interest in this definition being sharp. It also means we spend our weeks reading other people's API responses, which is where the sharpness comes from. When your aws s3 sync command exits zero against a provider that then drops a third of your multipart uploads on the floor, the badge on the pricing page meant nothing and the wire behaviour meant everything.
This is a long piece. It covers the definition, the category error people make when they type "s3 compatible block storage" into a search box, the eleven behaviours we now audit before we let a provider into a customer's migration path, and the parts of this market where we tell you to go use something else. We would rather you leave with the right tool than the wrong one from us.
S3 compatible storage means one thing: the provider answers the S3 REST API without you rewriting your client
Compatibility is a contract about wire behaviour, not a badge. The contract says: when my client sends an HTTP request shaped like the S3 REST API, your endpoint returns a response my client can parse, with status codes, headers and XML bodies that mean what the S3 specification says they mean. Nothing about the dashboard, the marketing page, or the fact that your provider's SDK wrapper exists changes that.
The fuzzy version is everywhere. People say "S3 compatible" to mean any of three things, and the blur between them is where projects die.
Tier one: protocol compatibility
The provider terminates HTTP requests on an endpoint and speaks enough S3 to satisfy a browser upload. This is the weakest tier. A provider can pass tier one with PUT Object, GET Object, DELETE Object and a hand-wave at everything else. It will pass a curl smoke test. It will fail the moment your application asks for ListObjectsV2 with a continuation token, or sends a presigned URL with a custom expiry, or tries a conditional write.
Tier one is fine for a static asset bucket behind a CDN. It is not fine for anything with a backup job, a sync job, or a retention policy attached.
Tier two: library compatibility
The provider passes the test suite of one particular SDK, usually the AWS SDK for the language its first customer used. This is where a lot of "we are S3 compatible" claims actually live. The provider tested with boto3 on Python 3.9, the calls worked, and the claim went on the website. Then someone arrives with the Go SDK, or the Rust client, or the AWS CLI v2 with a different signing path, and hits a header the provider never implemented.
Library compatibility is real and useful. It is not the same as protocol compatibility, and providers rarely tell you which one they mean. The way to find out is to run the client you are actually going to use, not the one in their docs.
Tier three: contract compatibility
The provider implements the S3 REST API to the point where the behaviours your workload depends on are stable, documented, and tested across clients. This is the tier that matters and the tier nobody advertises, because advertising it means publishing the list of things you have not implemented yet.
We connect 50+ S3-compatible providers at this point, plus custom endpoints, and the pattern is consistent: the providers that pass tier three are the ones that publish an explicit compatibility matrix with the gaps named. The ones that pass tier one tend to have a page that says "fully S3 compatible" and nothing else.
If you take one thing from this section: when a provider says "S3 compatible", ask which tier. If they do not know what you are asking, that is your answer. Our own provider list at storafleet.com/providers is the set of endpoints we have actually connected; it is not a claim that every one of them passes tier three for every workload.
Object and block are not two flavours of the same thing, and "s3 compatible block storage" is a category error worth explaining once
S3 is an object API. It has no block semantics at all, and no amount of compatibility changes that. This matters because "s3 compatible block storage" is a real search phrase with real volume, and the people typing it are usually about to make an expensive mistake.
Block storage gives you a device. You get a fixed-size volume, you attach it to a host, the host sees sectors, and a filesystem sits on top of those sectors. Reads and writes are addressed by offset. The device does not know what a file is, does not know what a bucket is, and does not care. AWS EBS is block storage. A local NVMe drive is block storage. iSCSI and NVMe-oF are the protocols that carry it.
Object storage gives you a flat namespace of keys and values. You PUT an object. You GET an object. You LIST keys with a prefix. There is no offset, no sector, no mount point, and no in-place edit of a byte range. Overwriting an object is a whole new object with the same key. AWS S3 is object storage, and so is every provider in this market.
These are different access patterns for different jobs. A database wants block. A backup archive wants object. You can run a database on object storage if you are very clever and very patient, and you will still be paying for that cleverness in latency for the rest of the system's life.
The only real bridge is a gateway
When someone genuinely needs S3-shaped access on top of a block backend, the thing in the middle is a gateway, and the gateway is where all the problems live. Products like s3gw, MinIO's gateway modes, and various vendor appliances take S3 requests and translate them into block operations underneath. The translation is lossy in both directions. Object listing becomes metadata the gateway has to maintain separately. Multipart upload becomes a staging area. Consistency guarantees get complicated fast.
If you are searching for "s3 compatible block storage", you are probably in one of three situations, and the answer is different for each.
- You want a filesystem that looks like a drive but stores to a bucket. That is not block storage, it is a FUSE mount, and the tools are rclone mount or Mountain Duck. We do not do this. More on that later.
- You want block storage and someone told you S3 is cheaper. It is cheaper per gigabyte for cold data and more expensive in engineering time the moment you need a filesystem. Price the engineering, not just the bytes.
- You want a gateway so an existing block-based appliance speaks S3. That is a real product category. It is not what any of the major object storage providers sell, and if a provider's page implies otherwise, read the fine print.
The category error is worth correcting once because the cost of not correcting it is a rewrite. Object storage is a destination. Block storage is a device. The two words do not stack.
The compatibility audit: eleven behaviours that decide whether a provider is actually compatible for your workload
Before you point a workload at a provider, run these eleven checks. Each one is a behaviour your code may depend on, and each one breaks something specific when it is missing. We run this list by hand when a customer asks us to move data between two endpoints we have not paired before.
1. Multipart upload thresholds
What it is: the minimum part size, maximum part count, and maximum object size the provider allows for multipart uploads. S3's own numbers are 5 MiB minimum part, 10,000 parts maximum.
What breaks: a provider with a different minimum part size will reject parts your client considers legal, and the error usually arrives after the upload has been running for an hour. A provider with a lower part count maximum caps your maximum object size, which you find out when the 400 GB file fails at 90 percent.
2. ListObjectsV2 pagination
What it is: whether the provider honours continuation-token, max-keys, start-after and delimiter the way S3 does.
What breaks: your inventory, your sync, your backup job, and anything that walks a bucket. This is the one that caught us. A provider that ignores max-keys and returns everything in one response will work in testing and fall over on a bucket with a hundred million keys.
3. ETag semantics
What it is: what the provider puts in the ETag header, and whether it is an MD5 for single-part uploads and a part-hash for multipart.
What breaks: any verification step that compares ETags across providers. Some providers return an opaque internal identifier instead of a hash. It is stable and it is useless for cross-provider comparison.
4. Presigned URL expiry limits
What it is: the maximum lifetime a provider will accept on a presigned URL. S3 caps at seven days for SigV4.
What breaks: any workflow that generates a long-lived link. A provider with a shorter cap will reject the URL at generation time, or worse, accept it and fail on use. If you generate links for external partners, check this before you commit to a number in your product docs.
5. Versioning
What it is: whether the provider supports bucket versioning, delete markers, and the ability to list and restore specific versions.
What breaks: your recovery story. Without versioning, an accidental overwrite is permanent. With versioning implemented partially, an accidental overwrite is permanent in a way that looks protected until you try to restore.
6. Object lock and retention
What it is: S3 Object Lock in governance and compliance modes, with retention periods and legal holds.
What breaks: any regulatory retention requirement. If your compliance team has signed off on WORM storage and your provider's object lock is a checkbox that does not enforce immutability, you have a problem that surfaces during an audit, which is the worst time.
7. Conditional writes
What it is: If-None-Match and If-Match on writes, the mechanism behind atomic create-if-absent and optimistic concurrency.
What breaks: any application doing distributed locking or coordinating writers without a separate lock service. Without conditional writes, two writers race and the last one wins silently. Many S3-compatible providers do not implement this. Check before you build on it.
8. Checksum headers
What it is: support for x-amz-checksum-sha256, x-amz-checksum-crc32 and the rest of the additional checksum family.
What breaks: your verification step, and your ability to prove data integrity to anyone who asks. A provider that accepts the header and ignores it is worse than one that rejects it, because you will believe you have verification you do not have.
9. Bucket naming rules
What it is: the character set, length limits, and reserved prefixes a provider accepts for bucket names. S3's rules are lowercase, 3 to 63 characters, no consecutive periods, not formatted as an IP address.
What breaks: migration between providers when the source bucket name is legal on one side and not the other. If your application has the bucket name hardcoded in config, a name that fails validation on the destination is a real blocker at cutover time.
10. Region reporting
What it is: what the provider returns for GetBucketLocation and how it handles region-specific endpoints.
What breaks: SDKs that infer endpoint URLs from region, and any multi-region setup where your client needs to know which endpoint to talk to. A provider that returns an empty region or a region string your SDK does not recognise will send requests to the wrong host.
11. Error code fidelity
What it is: whether the provider returns the S3 error codes (NoSuchKey, AccessDenied, SlowDown, RequestTimeout) with the right HTTP status codes, or collapses everything into a generic 500.
What breaks: your retry logic. A client that retries on SlowDown and gives up on NoSuchKey depends on the distinction. A provider that returns 500 for a missing key will send your application into a retry loop against an object that does not exist.
Eleven checks. Each one is a paragraph in a runbook and a possible rewrite if you skip it.
Here is the actual migration sequence we run now, step by step, after getting it wrong twice
This is the order. Skip a step and you will meet the failure that step exists to prevent. It is not elegant and it is not fast. It is the sequence that has survived contact with real migrations.
1. Inventory, and reconcile it
List every object on the source, page by page, with the continuation token followed to exhaustion. Do not stop on a short page. Do not assume ordering. Write the manifest to a local file with key, size, and last-modified for each object.
aws s3api list-objects-v2 \
--bucket source-bucket \
--endpoint-url https://source.example.com \
--page-size 1000 \
--query 'Contents[].[Key,Size,LastModified]' \
--output text >> manifest.tsv
Then reconcile: total object count, total byte count, and a comparison against the provider's own reported usage number. If those three do not agree within a small margin, you have a listing bug or a provider reporting bug, and you need to know which before you copy anything.
2. Scope the credentials
Mint a dedicated key pair for the migration. Read-only on the source, write-only on the destination, scoped to the specific buckets. Do not use the account root key. Do not reuse a key that has delete permissions on the source, because the day you fat-finger a bucket name in a delete command is the day you want that key to be useless.
If your destination is AWS, connect by IAM role assumption rather than a long-lived key. That is what we do where it is available, and it removes a whole class of leak.
3. Dry run
Run the copy with the destination set to a scratch bucket, or with a --dryrun flag if your tool supports it, and confirm the object count that would be copied matches the manifest. This is cheap and it catches the class of bug that cost us a day.
4. Checksum strategy, decided before you start
Decide now whether you are verifying by size, by ETag, or by a stored checksum, and write down which one you chose and what it does not cover. This is the step people skip and regret.
The ETag on a single-part upload is usually the MD5 of the object. The ETag on a multipart upload is not the MD5 of the object, it is a hash of the part hashes, and it carries a -N suffix where N is the part count. That means ETag comparison across providers is only valid when both sides used the same multipart chunking, which they frequently did not. Size comparison is always valid and always weaker. A stored x-amz-checksum-sha256 header, where both providers support it, is the strongest option.
Pick one. Know its limits. Do not discover them mid-cutover.
5. Cutover
Stop writes to the source. Wait for in-flight operations to drain. Run one final incremental copy pass to pick up anything written since the last sync. Flip the application's endpoint configuration. Start writes to the destination.
The drain wait is the part people underestimate. If your application has background jobs holding connections, "stop writes" is not instantaneous, and the objects written during the gap are the ones that go missing.
6. Verification
Run the checksum pass against the full manifest. Object count, byte count, and your chosen checksum. Read a sample of objects from the destination and compare their contents, not just their metadata, especially for the largest objects and the ones with the most unusual keys.
7. Rollback plan, written down before cutover
Know how you get back. If cutover fails at hour two, what is the command that points the application at the source again, and how long does it take. This is a paragraph in a runbook, not a feeling.
8. Keep the source warm
Do not delete the source on cutover day. Keep it for a week, or a month, depending on how much you trust your verification. The egress cost of keeping it is trivial next to the cost of discovering a gap after you have deleted the only other copy.
Multi-cloud S3-compatible storage is a control-plane problem, not a storage problem
The bytes live somewhere. The decision about where they live, and who moves them, is a separate system, and conflating the two is why multi-cloud projects get stuck.
Separate the data plane from the control plane and the architecture gets clear. The data plane is the objects sitting in buckets at Amazon S3, Cloudflare R2, Backblaze B2, Wasabi, MinIO, DigitalOcean Spaces, Oracle Cloud, IBM Cloud Object Storage, Scaleway, Linode, Vultr, Storj, IDrive e2, Hetzner, or wherever else. The control plane is the set of policies, schedules, and credentials that decide what gets copied where, when, and under whose authority.
Most multi-cloud storage conversations are actually about the data plane, which is the part nobody can unify. You cannot make one vendor's bytes live in another vendor's datacenter without copying them. That is physics, not a product gap.
The control plane is where the leverage is. Policies like "this bucket replicates to a second provider every night", "this namespace is searchable across all three clouds", "this bucket's lifecycle rules expire objects after 90 days", "these credentials are scoped to this provider only". Those are decisions, they are portable, and they can live in one place even when the bytes do not.
What we do, and what we do not
Storafleet is a control plane. It never stores customer file contents. We connect to your buckets, we move bytes between them, and we never keep a copy. That is not a marketing position, it is the architecture: we hold credentials and metadata, not objects.
We connect the S3-compatible family and any custom S3 endpoint. That covers the enterprise clouds and the smaller providers above. We do not connect Google Drive, Dropbox, OneDrive, SFTP, WebDAV or Box. If the job is "move a folder from Google Drive into a bucket", we are the wrong tool and we will tell you so, because building that connector is not on our roadmap and pretending otherwise wastes your evaluation time.
On the long-tail question about integration with AWS, Azure and GCP: we integrate with the S3-compatible endpoints. AWS is native S3. Azure Blob has an S3-compatible layer, and GCP Cloud Storage has an S3-compatible XML API, both with caveats, and we connect through those where they work. We are not a first-class Azure or GCP client. We are an S3 client, and if the Azure or GCP surface speaks enough S3 for your workload, we can drive it.
The honest summary: the data plane stays where you put it. The control plane is the thing you can centralise, and it is the thing worth centralising, because managing credentials and policies across five provider consoles is the actual pain of multi-cloud and it is a pain that a single console solves.
What it actually costs to run this at enterprise scale, and where the per-gigabyte pricing model quietly punishes you
The per-gigabyte number on the pricing page is the smallest line in your bill, and for a lot of workloads it is not even the second smallest. Enterprise object storage cost anatomy has six lines, and the pricing model you choose determines which of them grows.
Storage
The bytes at rest. This is the number everyone compares, and it is real. For a few hundred terabytes, the difference between providers at a fraction of a cent per gigabyte per month is a line item worth negotiating. Round illustrative numbers: a provider at a couple of cents per GB per month versus one at a fraction of a cent is a five-figure annual difference at petabyte scale.
This is also the line where the per-gigabyte model behaves exactly as advertised. You store more, you pay more, linearly, and everyone understands it.
Requests
The line that punishes you, and the one nobody models. Object storage bills per API request, with different rates for PUT, GET, and listing. A workload that reads small objects frequently can spend more on requests than on storage.
Here is the shape of it. A bucket holding a terabyte in ten million small objects, read once a day by an application that lists before it fetches, generates two requests per object per day. That is twenty million requests a day. At typical request pricing, that is a bill measured in hundreds of dollars a month for a terabyte of storage that costs a few dollars a month to hold. The bytes were never the problem. The access pattern was.
This is why the per-gigabyte number is a trap for some workloads and fine for others. If your objects are large and read rarely, storage dominates and per-gigabyte pricing is honest. If your objects are small and read often, requests dominate and the per-gigabyte number is a rounding error next to your actual invoice.
Egress
The line that decides whether you can leave. Some providers charge for bytes leaving their network. Some do not. This is the single biggest structural difference between providers in this market, and it is the reason a migration can cost more in egress than the destination costs in a year.
Egress pricing is also the reason the "just move if it gets expensive" advice is incomplete. Moving has a price, and the provider with the cheapest storage is often the one with the highest egress. Model the exit before you commit to the entrance.
Retrieval
The line that exists only on archive tiers. If you use an infrequent-access or archival storage class, reads carry a retrieval fee on top of egress. For backup data that is read only during a restore, this is fine. For anything read more than occasionally, the retrieval fee can exceed the storage savings that made the archive tier attractive.
Replication
The line that appears when your multi-cloud policy starts doing what you asked. Every scheduled sync between two providers is a read on the source and a write on the destination. If the source charges egress, you pay it on every run. A nightly replication of a large bucket is a nightly egress bill, and nobody puts that number in the architecture diagram.
Control plane
The line that is usually a flat fee, and the one that scales worst if it is per-gigabyte. This is our line. Storafleet is a flat subscription and we never charge per gigabyte to move data. Migrated bytes are unmetered on every tier, including Free. Tiers differ on concurrent migration jobs and on automation like scheduled sync.
We chose flat deliberately, because the alternative is that our bill grows when your job runs more, which gives us a reason to want your jobs to run more, which is a bad incentive to build into a product. A flat fee means the only way we make more money is by being worth more to you, which is the incentive we want.
Where flat pricing does not help
Flat subscription pricing is not universally better. If you run one migration a year and then nothing, a flat subscription is a worse deal than a per-gigabyte tool that charges you once. If your data volume is tiny, the per-gigabyte cost of anything is trivial and the subscription is the whole bill. If your workload is a single bucket in a single cloud, you do not have a control plane problem and you should not pay for a control plane.
The workloads where flat wins are the ones that move data continuously: scheduled replication, ongoing migration, multi-provider sync. The workloads where it loses are the ones that move data once and stop. Be honest about which one you have.
Where rclone genuinely beats us, and where we tell people to go use it instead
rclone is a better tool than us for a large number of jobs, and this section is not a setup for a "but". We are going to name the specific cases where it wins, and we are not going to walk any of them back.
rclone is free, permanently, with no account and no vendor. You download a binary and you have a tool. There is no signup, no subscription, no seat count, no renewal conversation. For an individual, a small team, or an organisation that has decided not to add another vendor relationship, that is the end of the discussion.
It supports more backends than we do. We connect the S3-compatible family and custom S3 endpoints. rclone supports over seventy backends, including the consumer clouds we deliberately do not touch: Google Drive, Dropbox, OneDrive, Box, and the rest. If your job involves moving data between a consumer cloud and a bucket, rclone is the answer and we are not.
It mounts. rclone mount presents a remote bucket as a local filesystem, which means Finder and Explorer and every tool that expects a path can read from it. We have no FUSE surface. We cannot do this, and if you need it, the conversation about us is over.
It is scriptable in the way a Unix tool is scriptable. It composes with cron, with systemd timers, with CI pipelines, with Makefiles, with shell scripts and with whatever else you have wired together. It exits with a status code. It writes to stderr. It takes flags. A web console never composes that way, and we are a web console.
It runs inside your own security boundary. The binary is on your machine, the config is on your machine, the credentials are on your machine, and no third party holds a copy. Our credentials are encrypted at rest with AES-256-GCM and per-namespace HKDF keys, and AWS can be connected keylessly by IAM role assumption, and none of that is the same as "the keys never left my laptop". The latter is a stronger answer than ours and it would be dishonest to pretend otherwise.
It has been maintained for over a decade by a community that is still active. On the question "will this still exist in five years", a decade-old open-source project has a better answer than a company that started in 2026. We are a small team. We have no published case studies and no named customers to point at, and we will not invent them.
So: if you have no budget, use rclone. If you need a mounted filesystem, use rclone mount or Mountain Duck. If you need Drive, Dropbox, OneDrive or SFTP, use rclone or MultCloud. If your credentials must never leave hardware you control, use rclone. If your workflow is entirely code and always will be, use rclone or another CLI, because a console does not compose and never will.
We say this because the alternative is worse. A customer who buys us for a job rclone does better is a customer who churns in three months and tells everyone we oversold. A customer who uses rclone for that job and comes to us for the multi-provider scheduled sync later is a customer for years. We would rather be honest about the first case and earn the second.
Who should not use Storafleet for this, and what to use instead
Here is the blunt list. If you are in one of these categories, we are not the right tool and the alternative is named.
- Budget is zero. Use rclone. The conversation is over and rclone wins. It is free forever, it has no account requirement, and it does everything a zero-budget migration needs.
- You need storage mounted as a filesystem. Use
rclone mount, Mountain Duck, or ExpanDrive. We have no FUSE surface and we are not building one. If remote buckets need to appear as a drive letter, we cannot help and we will not pretend the console is a substitute. - You need Google Drive, Dropbox, OneDrive, SFTP, WebDAV or Box. Use rclone or MultCloud. We connect the S3-compatible family and nothing else. CloudFuze covers the consumer clouds for enterprise workflows if you need a managed service there.
- Credentials must never leave hardware you control. Use rclone, or if it is AWS only, connect through our IAM role assumption so no long-lived key is stored. For any non-AWS provider, rclone on your own hardware is the stronger answer and we will not argue otherwise.
- Everything lives in one cloud and always will. Use that provider's own console. It is free, it is always current with that provider's newest features, and it is authoritative in a way a third-party control plane is not. We add value when there are two or more providers in the picture. With one, we are an extra layer with an extra bill.
- The workflow is entirely code and always will be. Use a CLI. Storafleet is a console. It does not compose with cron, systemd, or a Makefile the way a binary does. If your migration is a line in a CI pipeline, a binary is the right shape and a web UI is not.
- You need a single person managing a few buckets with no subscription. Use Cyberduck, S3 Browser, or Mountain Duck. One-time licence or free, files stay on your machine, and Mountain Duck mounts. Better fit than us for that scale.
- You need a vendor with published case studies and named enterprise customers. We do not have them and we will not manufacture them. Choose a provider with a longer track record if that is a requirement of your procurement process.
That is the list. We lose those sales on purpose, because the alternative is selling to people who should have used something else and then explaining to them why the thing we said we do is not the thing they needed.
What is left, if you are still reading, is the case we are built for: multiple S3-compatible providers, ongoing movement between them, credentials and policies that need one place to live, and a workload that runs continuously rather than once. That is a control plane problem, we are a control plane, and we charge a flat subscription for it because we would rather compete on being worth the fee than on metering your bytes.
Run the eleven-behaviour audit first. It is free, it takes an afternoon, and it will tell you more about your shortlist than any pricing page.