As someone with TBs of SSDs and HDDs laying around I am someone who agrees with your point, but this is not a good way to compare costs.
For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.
Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
Try doing the math and ask why your time costs. You need to buy a lot of storage to pay for the engineering work and most places would prefer to spend that time and attention on the product rather than shaving a few percent off of the storage bill (especially since the savings will be negative for quite some time).
Wait, I'm confused, are you arguing for or against S3? Did you mean factoring the time needed to obtain the multiple PhD levels of information in tracking the cost, usage, and security of those interoperable services plugged into S3?
Nice article. I agree that it is a bit of a shame that everything is forced to be so S3-centric (and I say that as someone whose work helped motivate a lot of people to do that), but right now its an unfortunate reality of running software in the cloud because cloud networking and SSDs are so expensive that you really are required to use S3 if you want a system that can handle "big data" scale workloads cost effectively
The unfortunate part of that graph is that I think it's not updated for today's prices given the memory shortage/crunch we're experiencing that's driven up the prices for all kinds of memory.
But also while it should be faster than it is, how would an S3 designed around SSDs look differently to an API user? I would think the API is basically the same.
We ran EBS for years. Finally the costs became ridiculous and we moved our entire dataset to S3. We’re saving 20k/month. There’s just no beating the price.
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
S3 is terrible deal on any front. it's just easy
Everything in AWS is expensive. But at least S3 adds a lot of value.
As someone with TBs of SSDs and HDDs laying around I am someone who agrees with your point, but this is not a good way to compare costs.
For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.
Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
Even if you factor that it, S3 is pretty expensive for most small to medium size data.
Try doing the math and ask why your time costs. You need to buy a lot of storage to pay for the engineering work and most places would prefer to spend that time and attention on the product rather than shaving a few percent off of the storage bill (especially since the savings will be negative for quite some time).
Wait, I'm confused, are you arguing for or against S3? Did you mean factoring the time needed to obtain the multiple PhD levels of information in tracking the cost, usage, and security of those interoperable services plugged into S3?
Nice article. I agree that it is a bit of a shame that everything is forced to be so S3-centric (and I say that as someone whose work helped motivate a lot of people to do that), but right now its an unfortunate reality of running software in the cloud because cloud networking and SSDs are so expensive that you really are required to use S3 if you want a system that can handle "big data" scale workloads cost effectively
The unfortunate part of that graph is that I think it's not updated for today's prices given the memory shortage/crunch we're experiencing that's driven up the prices for all kinds of memory.
But also while it should be faster than it is, how would an S3 designed around SSDs look differently to an API user? I would think the API is basically the same.
It would probably look a lot like existing S3 concepts that have explicit hot and cold tiers. For example Intelligent Tiering, or Glacier.
We ran EBS for years. Finally the costs became ridiculous and we moved our entire dataset to S3. We’re saving 20k/month. There’s just no beating the price.
One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade:
Today it's still $0.023/GB-month.