Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Prompting Claude Opus 5.5 (claude.com)
    37comments
  2. Owed a billion dollars in Nvidia stock (colo.to)
    275comments
  3. Thinking fast and slow in AI: The role of metacognition (2021) (arxiv.org)
    23comments
  4. Ember-1 (fireworks.ai)
    211comments
  5. When did Google get so weird? (sancho.bearblog.dev)
    704comments
  6. Malleable software: Restoring user agency in a world of locked-down apps (2025) (inkandswitch.com)
    34comments
  7. Made by Mechanical Means (felixrieseberg.com)
    1comments
  8. Nissan's third generation e-POWER powertrain (nissan-global.com)
    54comments
  9. Footguns with Postgres "at time zone 'UTC'" (bookofrevenue.com)
    2comments
  10. Maybe don't let Muse run your Facebook Marketplace account (threads.com)
    15comments
  11. Functional Mechanical Sympathy [video] (youtube.com)
    1comments
  12. Alan Kay's answer to “Did the ENIAC have a BIOS”? (quora.com)
    40comments
  13. Self-Hosting on the Dark Web (alvarezrosa.com)
    67comments
  14. Lunar Terminator Paradox (secretsauce.net)
    46comments
  15. Was silent reading unusual during Augustine's time? (historyofinformation.com)
    29comments
  16. The state of SIMD in Rust in 2026 (shnatsel.github.io)
    33comments
  17. Guitar amp and effects pedal built on the Waveshare ESP32-S3-Touch-AMOLED-2.06 (github.com/dashersw)
    32comments
  18. Don't couple your Go code to GitHub (iain.rocks)
    108comments
  19. Deterministic Concurrency [video] (youtube.com)
    3comments
  20. There is more to code review than (automatable) detection (adaptivecapacitylabs.com)
    82comments
  21. Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi (loficities.com)
    112comments
  22. In an $80 motel room, a discovery to shed light on the origins of life (nytimes.com)
    92comments
  23. What I did at Recurse Center (thill.me)
    32comments
  24. Imp is a full port of DSPy to the BEAM (github.com/deepfates)
    7comments
  25. Reading’s Bayeux Tapestry (diamondgeezer.blogspot.com)
    16comments
  26. Replacing the old battery on rechargeable bike lights (jvns.ca)
    91comments
  27. Previously unheard recordings of John Coltrane, captured by Frank Tiberi (jazzwise.com)
    33comments
  28. Oral history of John Chowning, inventor of FM synthesis [video] (youtube.com)
    21comments
  29. Writing Efficient C++ Code (2013) (asawicki.info)
    118comments
  30. Packing Binary Is Fun (hereticpleb.vercel.app)
    4comments

S3 Is the Future, S3 Is the Past

66 pointsby 2d agobtrblocks.com
78 comments
10h agoHN ↗

S3’s dominance is due to its many advantages: effectively infinite capacity, high durability, and low per-gigabyte capacity cost.

It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it

S3 is terrible deal on any front. it's just easy

10h agoHN ↗

Everything in AWS is expensive. But at least S3 adds a lot of value.

6h agoHN ↗

Ever tried Ceph? It's a bit painful to set up and operate, but it seems to work pretty well.

(Don't bother with Ceph Object Gateway unless you really really need it - since you're programming your own application, access the storage pool directly with librados)

10h agoHN ↗

As someone with TBs of SSDs and HDDs laying around I am someone who agrees with your point, but this is not a good way to compare costs.

For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.

Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.

8h agoHN ↗

If you have only 2GB of data it's trivial, it hardly matters how you store it. You might as well hold a full copy in RAM on each server and synchronize it with your developer laptop every half hour in case they all go offline at once, for all it matters.

7h agoHN ↗

To reliably hold a bit over 2GB of data in RAM on AWS, you'll probably need something like a t4g.medium, about $0.0336/hr on-demand. 730 hours in a month, so $24.528, but ideally you'll have like three or so instances so like $74/mo. That's before thinking about stuff like IP address costs, the EBS volumes underpinning those boxes, etc. And then managing all the syncing and what not.

Of course, that's list prices, you can get savings plans and RIs and discounts.

The cost of RAM per hour in AWS isn't cheap.

6h agoHN ↗

Once again proving that AWS is expensive any way you slice it. Even in today's RAM crisis, an extra 2GB is what, $50? and you probably already have 2GB free since they don't make RAM modules that small.

10h agoHN ↗

It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it

This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.

That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.

The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.

10h agoHN ↗

Even if you factor that it, S3 is pretty expensive for most small to medium size data.

10h agoHN ↗

Try doing the math and ask why your time costs. You need to buy a lot of storage to pay for the engineering work and most places would prefer to spend that time and attention on the product rather than shaving a few percent off of the storage bill (especially since the savings will be negative for quite some time).

10h agoHN ↗

Wait, I'm confused, are you arguing for or against S3? Did you mean factoring the time needed to obtain the multiple PhD levels of information in tracking the cost, usage, and security of those interoperable services plugged into S3?

9h agoHN ↗

Many of us just don't buy your argument. "Engineering work" for maintaining our own storage is not rocket science.

You just plug it in and basically it runs. Thats pretty much it.

Having to explain to an HN audience how installing an SSD is trivial is weird.

9h agoHN ↗

I have done the math as part of my job and often found that yes, it was literally worth the money to do our own storage in cases. Running a highly available, durable storage system is not particularly difficult. There was an era when it was a common skill for a sysadmin. In the current era there are off the shelf solutions for it, both open and proprietary.

The biggest reason people pay for S3 is because of data transfer cost to other cloud services they are using.

3h agoHN ↗

a few percent

Really? No, it's not a few percent.

9h agoHN ↗

Recent Iranian attacks on AWS locations in Bahrain resulted in loss of data that simply could not be recovered. https://www.infoq.com/news/2026/09/aws-middle-east-data-loss...

Of course, if you had paid extra for the backups and redundancy, your data would survive.

So to address your original point, unless you are paying extra for redundancy, it is cheaper if you use a 4TB SSD at retail price as the original commenter discussed.

8h agoHN ↗

In this case AWS did not save you. AWS is supposed to be resilient against AZ failures within a region. Iran destroyed all AZs at once, so all the data was lost. If you had a hard drive in your office, either directly serving your project or as an off-site backup, you still have your data. In the end, was it worth paying several times as much for AWS to provide durability for you, only to have them lose the data anyway?

8h agoHN ↗

You seem to have misunderstood. What you said is exactly my argument. I am with you :).

The parent commentator is under the illusion that AWS automatically means security and scale and reliability.

I was merely pointing out to him that Iranian attacks must serve as a wake up call for him.

8h agoHN ↗

Do you need that? If you want your own replicated storage cluster at home, you can use Ceph. I've seen Ceph deployed to utilize the spare hard drive ports in server clusters that otherwise mostly do compute (transcoding), alongside memcached to utilize the RAM, etc.

I'd bet the majority of projects, but perhaps not the majority of traffic, can run adequately on a single server and tolerate a few hour maintenance window one weekend every month. Don't waste money on overkill.

9h agoHN ↗

If you play the game correctly it can be a good deal. The ~1 USD per TB per month of glacier deep archive is hard to compete with if you (probably) don’t need to read the data back.

8h agoHN ↗

What use is data that is never read unless it is for redundancy, disaster recovery, legal, or audit purposes, etc.

Yes you can make the argument for that specific use case.

For everyday use case, using s3 as your primary storage is costly and is not at all ideal.

8h agoHN ↗

I agree, but that's cherry-picking what is perhaps the only cost-effective service in all of AWS.

7h agoHN ↗

S3 cost growth is so gradual most businesses will just absorb the cost. We have some physical backups, but everyone is more than happy to not be shuffling around physical hard drives, and pay Amazon to deal with that and securely store it.

5h agoHN ↗

That's AWS in general for you. It has its place, but it's vastly overused in our industry. A lot of businesses would find their TCO would go down if they ditched cloud services, but that's not trendy enough for execs to do it.

10h agoHN ↗

Nice article. I agree that it is a bit of a shame that everything is forced to be so S3-centric (and I say that as someone whose work helped motivate a lot of people to do that), but right now its an unfortunate reality of running software in the cloud because cloud networking and SSDs are so expensive that you really are required to use S3 if you want a system that can handle "big data" scale workloads cost effectively

10h agoHN ↗

The unfortunate part of that graph is that I think it's not updated for today's prices given the memory shortage/crunch we're experiencing that's driven up the prices for all kinds of memory.

But also while it should be faster than it is, how would an S3 designed around SSDs look differently to an API user? I would think the API is basically the same.

10h agoHN ↗

It would probably look a lot like existing S3 concepts that have explicit hot and cold tiers. For example Intelligent Tiering, or Glacier.

9h agoHN ↗

That's S3 Express One Zone. Directory buckets are SSD-backed and regular buckets are HDD-backed. Some of the API differences off the top of my head:

- Directory entries are no longer returned in sorted order in ListObjectsV2

- There's an AppendObject API

- There's a RenameObject API

8h agoHN ↗

It's hard to imagine how these API differences can be explained by the different underlying block device. I don't see any good reason you couldn't support these operations on a HDD.

I suspect it's more to do with the fact that with One Zone is a clean rewrite of large parts of the application stack that makes up S3.

S3 is made up of hundreds of microservices[0], there probably isn't anyone at Amazon that actually understands the whole system. Refactoring it to support these features probably requires coordination between a lot of different teams. They might have petabytes of metadata, making a change to how metadata is persisted probably requires a massive risky data migration.

[0]: "All in, S3 today is composed of hundreds of microservices" - https://www.allthingsdistributed.com/2023/07/building-and-op...

23m agoHN ↗

For what it's worth I think you are correct, and that these differences have nothing to do with being SSD-based.

10h agoHN ↗

We ran EBS for years. Finally the costs became ridiculous and we moved our entire dataset to S3. We’re saving 20k/month. There’s just no beating the price.

8h agoHN ↗

Well of course, by doing absolutely anything on AWS it's about 5-10 times the price it would be to do yourself, or 100 times if it involves egress data.

10h agoHN ↗

One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade:

  2006-03-14  $0.150/GB-month
  2010-11-01  $0.140/GB-month
  2012-02-01  $0.125/GB-month
  2012-12-01  $0.095/GB-month
  2014-02-01  $0.085/GB-month
  2014-04-01  $0.030/GB-month
  2016-12-01  $0.023/GB-month

Today it's still $0.023/GB-month.

9h agoHN ↗

That's true for the base S3 product, but there are a lot more storage tiers than there used to be with cheaper pricing. All the way to $0.99/TB for Deep Archive.

9h agoHN ↗

Well there was also a period of low inflation in them during that time.

God why am I defending Amazon?

9h agoHN ↗

Not that deep archive isn't a valid option for the appropriate work load, but it has different effective costs.

9h agoHN ↗

Egress pricing of $0.09/GB has been around for over a decade too.

For reference transit cost has dropped over the years and is now about at $0.000247/GB or free if there is a peering agreement with the network the data is sent to.

8h agoHN ↗

It's the lock-in cost. They don't want you to move all your data out, they want it to remain trapped. Europe forced them to allow a one-time free exit, but you have to negotiate it with their support, so they're hoping nobody uses it.

7h agoHN ↗

Egress pricing pays for the AWS network infrastructure that is incredibly reliable (hardware failures happen regularly yet almost no one notices) and allows it to operate at scale sufficient to absorb even the largest DDoS attacks.

This isn’t just AWS BTW; all tier 1 cloud providers recoup their costs this way.

7h agoHN ↗

Cloudflare R2 has zero egress fees.

7h agoHN ↗

CloudFlare is not a tier 1 cloud provider.

7h agoHN ↗

CloudFlare is not a tier 1 cloud provider.

Yes, you're legally correct, but the point is AWS are clearly overcharging for egress.

As you say:

Egress pricing pays for the AWS network infrastructure that is incredibly reliable (hardware failures happen regularly yet almost no one notices) and allows it to operate at scale sufficient to absorb even the largest DDoS attacks.

Cloudflare is also pretty good at absorbing DDoS attacks, yet charge no egress costs.

Maybe the real question is, are the tier-1 providers colluding in overcharging for egress?

6h agoHN ↗

I don’t think we can say with certainty without knowing more details about their respective network investments. CloudFlare is architected quite differently than the others, and has a much less comprehensive service portfolio. It’s not an apples to apples comparison unless we’re strictly comparing S3 to R2. Choosing to charge less may also be a loss leader or a differentiation play.

6h agoHN ↗

If they're charging what customers are willing to pay, are they overcharging? I think an axiom of capitalism is that that's the amount you should charge.

Many people avoid T1 clouds because they're so expensive, but they seem to have found an extremely profitable market segmentation consisting of the remainder.

3h agoHN ↗

If they were offering it in a competitive market without the special bundling factors, basically nobody would choose their egress. So yes they're overcharging, because they have control of your data.

6h agoHN ↗

This is just Stockholm Syndrome. All T1 (except Cogent) and most T2 ISPs (that's most ISPs) are also very reliable, how much do they charge for transit?

6h agoHN ↗

You’re confusing an ISP with a tier 1 cloud provider. They are different animals altogether.

And kindly refrain from accusing others of Stockholm Syndrome here. It’s incredibly rude.

6h agoHN ↗

I am not confusing T1 ISPs with T1 clouds. When I said T1 ISPs, I meant T1 ISPs. Because those are really reliable ISPs (that you should probably not use as your ISP, for reasons irrelevant to the comparison).

Finding any excuse to justify a 100x profit margin is a form of Stockholm syndrome.

9h agoHN ↗

I guess inflation is the price reduction we get. 23 cents in 2016 is 32 cents in 2026

8h agoHN ↗

And that's by official inflation numbers. If you go by how much prices of food and rent increased i think you get a number more like 50 or 60 cents

8h agoHN ↗

Just be happy they keep the GB as large as they used to...

5h agoHN ↗

Is S3 priced in actual power-of-two GB? Or the shrinkflation power of ten GiB?

59m agoHN ↗

From the S3 pricing details[1]:

Amazon S3 storage usage is calculated in binary gigabytes (GB), where 1 GB is 2^30 bytes. This unit of measurement is also known as a gibibyte (GiB), defined by the International Electrotechnical Commission (IEC). Similarly, 1 TB is 2^40 bytes, i.e. 1024 GBs.

I guess they could just change that overnight in the future. Still, a lot better than charging for Storage Units or similarly arbitrary unit.

[1]: https://aws.amazon.com/s3/pricing/#S3_Pricing_Details

7h agoHN ↗

This is absolutely the case for nearly all of AWS products. Some new services have filled lower price gaps, but I previously looked at EC2 and other service prices in the past and late 2016 is where it all seemed to stop getting price drops. The M5+ upgrades for example all came with price increases along with the performance gains.

1h agoHN ↗

Every new EC2 generation is typically slightly more expensive, but the gains in the compute performance are 15-25% or even higher depending on your workload. The same dollar buys you a lot more compute power than a decade ago.

9h agoHN ↗

Right now it is not even clear how to interface with SSDs even on a single host, there has been all sorts of attempts to move away from the traditional plain block device model. NVMe has extensions for KV, ZNS, and FDP, all which offer different characteristics. And then there are of course open channel SSDs and some others too. I kinda expect the future foundational IO interface to be more S3-like than block-device or unixy filesystem-like.

9h agoHN ↗

I really wish reviewers would harp on FDP (Flexible Data Placement) and perhaps KV support. FDP supposedly somewhat ate ZNS as a spec, allegedly, but there might be gaps, reasons to keep ZNS.

There's only a small little mention, if we are lucky, on the couple drives that have it (expensive enterprise flagships). It should be a regular sticking point, whether it's there or not. Without pressure it's not going to get regularly available, it feels like.

FDP is so simple. Declare a number for what pool of data you want to write into. Data of the same pool gets written to the same storage such that you can wipe it latter together. It has huge wins though against write amplification! Massive wins. For so close to free.

Some day I want to own a FDP drive. And then I can finally start using the tokio/io-uring support that I contributed! https://github.com/tokio-rs/io-uring/issues/380

The NVMe-KV is more radical. Still worth putting some pressure on, but your drive as KV, as object store, feels harder. Side note, really enjoyed this ceph nvme-kv offload post thing, my favorite tech write up in a while! https://ceph.io/en/news/blog/2026/for-whom-the-door-bell-tol...

6h agoHN ↗

FDP relies on putting logic on the drive side, then they can upcharge for drives with this feature and still give you limited control. If you go the other direction instead, you have MTD devices which gives the host system kernel full control over data placement, page erasing and error correction, and for this reason they need specialized filesystems. These devices usually aren't attached over PCIe as they use controllerless raw flash interfaces instead, but in principle they could be.

5h agoHN ↗

I'm very much a fan of open channel flash. My understanding is that the Open Channel flash people when they tried a decade ago basically got told no by drive makers, that the drive makers weren't interested in becoming commodity vendors, and were intent on keeping product lines somewhat as they were. I wish I had some links to share to back this up, but that's basically my recall. ZNS and then FDP was sort of a compromise to give people some of the wins they were looking for, while basically not disrupting the product.

FDP is a very minimal addition of control. But yes it is another box to ticket, is another place to upcharge. Yet still, we're only just seeing mainstream products emerge. Kioxia's CM10 for example. https://www.techpowerup.com/351218/kioxia-introduces-first-p...

I'm hoping that the need is great enough to break the industry control. I think for a while the market felt relatively well enough served such that it was unclear whether clear wins would actually result in customers. With AI need for speed, I can definitely imagine incredibly crazy CXL controllers or what not, that allow low latency acces to many many open channel flash systems. Wouldn't that be a thing.

Again though, FDP is such a ridiculously tiny add, and it helps SSDs so much for so many use cases. I really hope it becomes an expectation, not a feature, sooner rather than latter.

Edit: happy to see a new group has shown up asking for open channel flash, Open Flash Platform. https://openflashplatform.org/

8h agoHN ↗

Why would it be S3-like? S3 is a very general abstraction on storage. If there's room below the current abstractions, it's below, not above - Linux has drivers for raw flash devices (mtd devices) which gives Linux full control of the program/erase cycle and responsibility for wear-leveling.

4h agoHN ↗

I would assume “S3-like” in the sense that objects are immutable, large writes, separate metadata storage, etc. Patterns that fit modern SSDs better than the abstractions of block devices.

9h agoHN ↗

I find Cloudflare R2 to be much cheaper as there are zero egress fees irrespective of data size. One is only billed on the count of requests with 10 million free download requests per month (check pricing for precise info).

If you do a clever bit of caching work in your app like we have with our apps Slyp and SlypBusiness, you could even make it insanely cheap.

We have a file manager that is ultimately responsible for images, files, videos and whatnot. It loads the images/files from the downloaded local cache when requested by a feature in the app.

A standard feature uploads and downloads using signed urls obtained from the backend service. The manager increments the file's version in the backend on each successful upload. For download, the manager compares the version against the locally cached version. If there is an update, the new file is downloaded in the background and overwrites the local cache and notifies all features using the file inside the app.

8h agoHN ↗

Is there a catch? So you can host a 1TB file and if 9,999,999 people download it you pay $0?

7h agoHN ↗

There's no catch. They don't care about the bandwidth. If 20 million people make a single request each then you'll pay $3.60 to cover the requests.

And you're going to have to pay $15 of storage costs for your terabyte.

4h agoHN ↗

The tl;dr for anyone skimming the comments: "CloudFlare doesn't like you" in this case means a gambling site that was getting CloudFlare shared IP addresses legally blocked by ISPs in countries where the gambling site wasn't licensed/was illegal. And "abusing low prices" means CloudFlare wanted them to pay for a bring-your-own-IP plan to give them dedicated IPs so their blocks wouldn't affect other CloudFlare customers.

3h agoHN ↗

Cloudflare was really unclear about their motivation was and it was effectively a matter of kicking them off for not liking them. CF was annoying them about an overpriced enterprise plan, before suddenly throwing out accusations that if true would have meant cutting the company off even if they did buy enterprise. And they were demanding a ton of money, much more than bringing an IP is worth.

Maybe maybe bringing their own IP would have solved the problem, but Cloudflare was obstructing everything and then did a cutoff without proper warning. It was really bad on Cloudflare's part.

2h agoHN ↗

This the same cloudflare that gets blocked in the entire country of Spain during every football match and says nothing about it?

They really hated that particular customer.

6h agoHN ↗

I forgot about R2. And I just got a problem perfect for it. Thanks for the reminder. HN helpful yet again!

7h agoHN ↗

I finally stopped using (S)FTP when i realized that all the ftp clients now have native support for S3. It turns out if you have a business partner who wants to use sftp, 99.9% of the time of they upgrade their client, the upgrade has S3 support, and then you can just use modern tooling on your side. And when they finally automate, they can also use modern tooling.

6h agoHN ↗

Had the same experience with trying to use the built-in SFTP support in Azure Storage accounts.

It turned out that all of our peers supported blob storage better than SFTP, which has some “show stopper” problems like forced outages caused by mandatory host key rotations.

3h agoHN ↗

DDEX choreography, which is used to distribute musical recordings and other assets within the music industry to sites like Spotify followed this same evolution.

The standards say SFTP. Most everyone ignores that part and have been using S3 buckets by bilateral consensus instead for years.

1h agoHN ↗

Has anyone figured out how glacier works yet? I wish they’d just tell us. I’d happily sign an NDA, I just want to know how it works!

I’d also like a “shitter” storage service. I.e no redundancy, no multi-AZ. Just really cheap storage for stuff I don’t really care about losing, so that I can store a bunch of stuff in bulk that can easily be recovered.

4m agoHN ↗

What I've read, and it seems reasonable to me, is that glacier is mostly just stored on the same HDDs that everything else is stored on.

What you need to understand is that a modern hard drives are severely limited in the amount of IOPS they can perform. Disk size has grown, but read and write speeds have not kept up with it.

In 1990 you could read a whole hard drive in 2 minutes, in 2026 as sizes have grown it takes the entire day.

As a system designer you want the maximum utilisation of your hardware, so you want to be close to saturating the IO that your drives are capable of.

What that means in practice is that only a tiny portion of the drive can be allocated to files that are read often.

That leaves the rest of the drive for files which are not used very often.