Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. You can run Git on object storage if you re-make packfiles(tigrisdata.com ↗)
    discuss
  2. Fourteen AI models against a customer edge. Divert blocked first(divert.cloud ↗)
    1comments
  3. Fault tolerance in low-bandwidth model parallelism(tplr.ai ↗)
    discuss
  4. Driscoll's Gave China Its Blueberries–Then China Swiped the Secret(wsj.com ↗)
    discuss
  5. Breakout List(breakoutlist.com ↗)
    2comments
  6. Show HN: Pixel Agents – A pixel-art mission control for your Claude Code agents(mateovalle.github.io ↗)
    discuss
  7. A Trump-backed crypto bill just suffered a bruising defeat in the Senate(npr.org ↗)
    discuss
  8. Oil Prices Could Hit Highest Levels in Months After Saudi Pipeline Attacks(nytimes.com ↗)
    discuss
  9. The first (public) System One Model; Jev gives AI the properties of code(typesafe.ai ↗)
    discuss
  10. Iran strikes on Amazon data centers caused permanent loss of customer data(arstechnica.com ↗)
    2comments
  11. Amazon bumps minimum starting pay to $20/hour, adds new discount for employees(thehill.com ↗)
    discuss
  12. Google confirms Pixel phones exploited in 'targeted' modem based attack(9to5google.com ↗)
    discuss
  13. AI Compilers Are Not Just Compilers for AI(aicompilers.github.io ↗)
    discuss
  14. History of the BBC TV Idents(bbc.com ↗)
    discuss
  15. A mouse whose brain cortex is made up of human cells(technologyreview.com ↗)
    discuss
  16. Ask HN: AI Panic, you buying it?
    3comments
  17. The FBI Doubles Down on Easing 'Bestiality' Hiring Standards(wired.com ↗)
    1comments
  18. Show HN: Smarter Shell History for Zsh(github.com/overflowy ↗)
    discuss
  19. OSRS Wiki and RuneLite are under strain from low-effort AI development(runescape.wiki ↗)
    1comments
  20. CO3: Toward the Optimal FFI(mversic.github.io ↗)
    discuss
  21. Training Text-to-Image Models 3.6× Faster(linum.ai ↗)
    discuss
  22. Gnome 51 Released with Improved Frame Scheduling, Many App Improvements(phoronix.com ↗)
    1comments
  23. We killed our in-app AI chat and made the agent a first-class user(sensefold.app ↗)
    discuss
  24. Omarchy on Mac(github.com/omacom ↗)
    discuss
  25. Whiteboard Defense(twitter.com/mitchellh ↗)
    discuss
  26. Show HN: Ivx/AI Chat, AI client that doesn't track you (with a better GUI imo)(ivx.run ↗)
    discuss
  27. As Someone Who Has Written This Blog Post,(zachholman.com ↗)
    discuss
  28. Numerical Tours(numerical-tours.com ↗)
    discuss
  29. IPC Panic, a 4-call kernel panic in Apple XNU, up to macOS and iOS 26 included(seriot.ch ↗)
    discuss
  30. Inside the MIT Report on How AI Is Eroding Campus Culture(learningcurve.fm ↗)
    1comments

Salesforce Global Outage

226 pointsby 6h agostatus.salesforce.com
127 comments
6h agoHN ↗

That status page is the most salesforce thing ever.

Scroll down. >_<

6h agoHN ↗

At least they're consistent about UX. My only complaint is that it needs more tabs.

5h agoHN ↗

I see tabs with a spinner loading infinitely, which is indeed very salesforce like

5h agoHN ↗

Wow, it's almost as long as the Every UUID V4 or Every Floating Point Number pages.

5h agoHN ↗

Yup. No mention of outage. Even drilling down gets nothing more than "Service Disruption".

5h agoHN ↗

It lists all instances, you can drill down into each one of them, see which services are affected, and each affected service pops up the incident timeline?

Isn't it actually amazing, and not "the most salesforce thing ever"?

6h agoHN ↗

This is certainly a unique status page.

6h agoHN ↗

At least now we can figure out what Salesforce does.

5h agoHN ↗

"And in tonight's news, the worldwide CRM solution Salesforce had a global outage affecting one hundred percent of its customer base. We interviewed users of the service to find out the scope of the impact. Everyone agreed that they were impacted, but strangely, nobody could describe _in what way_ they were affected."

6h agoHN ↗

Perfect timing with Dreamforce this week.

5h agoHN ↗

Remind me please, what are folks currently paying per seat for this glorified CRUD app?

5h agoHN ↗

It's this thing called golf course driven development

5h agoHN ↗

If I had been just out of Uni I would think this is edgy

Now I'm just glad I'm not responsible for this fire

4h agoHN ↗

They let their agentic AI take over system maintenance at Dreamforce yesterday like they were pushing in the talks

/s (partly)

5h agoHN ↗

Could it be people doing Claude/GPT automations and they just can't handle it?

3h agoHN ↗

Could be. With a big conference going on and the updates mentioning resource exhaustion it could be a bunch of people doing demos, AI driven or not. Basically slashdotting themselves.

5h agoHN ↗

And nothing of value was lost. God, I hate everything about Salesforce. Sometimes I have to integrate against their services, and it is always a pain, not to mention what the core project actually is: optimization of marketing and spam.

3h agoHN ↗

not really defending salesforce but OAuth+REST is a pain? Pretty plain vanilla in terms of integration requirements.

2h agoHN ↗

Sure, if you’re struggling with weird custom endpoint behavior, or custom triggers and validation, that’s on your admins and devs.

But if you’re having trouble with the standard rest or bulk apis, that’s 100% on you.

1h agoHN ↗

What is painful about a plain REST API?

5h agoHN ↗

You better bet someone started their agents with a prompt "Make a salesforce clone but with 100% uptime"

5h agoHN ↗

"Make a salesforce clone but with 100% uptime"

  ⎿  You've hit your session limit · resets 2:53am (48°52.6′S, 123°23.6′W Etc/GMT+8)
  /upgrade to increase your usage limit.
4h agoHN ↗

More like 70% human-configured DNS, 25% human-configured routing configuration, 5% interesting software bug.

3h agoHN ↗

Entirely correct.-

(Nowadays any of those need to fit in an "agent dropped all tables. Apologized" moment.-)

4h agoHN ↗

Feel like it has to be for all of this to go down at the same time.

5h agoHN ↗

Never before in the history of global compute outages was so little lost by so many down servers, whose purpose was known to so few.

5h agoHN ↗

So I guess today everyone gets actual work done

5h agoHN ↗

Wow, the intern must have tripped over a very big power cable this time

5h agoHN ↗

Sam speaks at Salesforce.

Salesforce goes down.

No causation here...move on.

1h agoHN ↗

Huang, Sam and Dario speaking at Dreamforce.

5h agoHN ↗

My gut instinct is that this is about when all of their on prem servers were EOL and their /public cloud solution was required. This must have had something to do with that

5h agoHN ↗

My gut instinct is that this is about Dreamforce with the rickshaws and whatnot

5h agoHN ↗

Ah yes. Exactly what a status page should look like: an endless list of random ID’s that don’t mean anything and no information whatsoever

At least salesforce is consistent with their design language

5h agoHN ↗

random ID’s that don’t mean anything and no information whatsoever

If you use salesforce you know what all of that stuff means. Just click on one, it’s not rocket surgery.

5h agoHN ↗

Was really just poking fun at them - AWS’s status page isn’t much better

5h agoHN ↗

Looks like a region list to me, maybe just with a lot of regions

3h agoHN ↗

Random Ids? If you mean the “USA324” ones, those are pods. If you’re a customer you know which one(s) you care about.

53m agoHN ↗

This might actually be one of the most useful status pages ive seen. Its not just random green tick marks representing the entire service that only change yellow when someone gives and admits that 5 hours of bad service is an outage.

5h agoHN ↗

You can click on any of the instances and then the service that is down to read the updates. It’s not 100% clear but some sort of issue with a “legacy login service”. The latest updates say a fix is rolling out.

5h agoHN ↗

Haven't they got some kind of new fancy ai interface they can use to fix it?

5h agoHN ↗

Have you tried turning it off and then on again?

We're no longer pursuing restarts as a path to remediation.

Oh you have

4h agoHN ↗

Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it.

We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.

At least it didn't fix the problem so they can actually start finding the real cause.

We're no longer pursuing restarts as a path to remediation.

Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?

4h agoHN ↗

"Yeah, I'm with Rob. Just let's reboot and see what happens"

2h agoHN ↗

"If that doesn't work, clear the cache and reboot again."

3h agoHN ↗

In my experience it’s a safe way to do something useful while everyone is getting their bearings. It immediately partitions the situation space between being persisted or systemic vs local or caused by long-running processes. Plus everyone’s going to ask if you’ve tried that already, so you might as well get it out of the way if it makes any amount of sense

2h agoHN ↗

Followed quickly by "Redeploying with more log lines", the next logical step.

3h agoHN ↗

I don't know, that reads exactly like an AI troubleshooter working through a plan without the implicit contextual understanding an experienced human might bring to either the actions or the communications.

"Oops, we forgot to tell it that this is the hyperscaled Salesforce production environment and that its choices need to project competence and consider brand embarrassment. WILLFIX"

2h agoHN ↗

Restart should be a very last emergency step, as if it works, a restart often might wipe out evidence of why.

So hopefully it's not done often.

18m agoHN ↗

I wouldn't say so, rather you need to balance recovery time and evidence preservation. A good incident manager will give the service owning team a chance or two to debug, but not let them fall into the trap of needing to understand the problem fully before attempt a clumsy potential fix. And of course will take into account the total business impact of the ongoing disruption and the known and unknown risks of the proposed clumsy fix (it could make things worse).

2h agoHN ↗

Been many an MIR that I've seen where after initial assessment, the next log entry was "service restart attempted"

31m agoHN ↗

In my experience this is very common on linux hosts. Not that you would do a full system reboot as a generic first attempt (this was much more common when I worked with Windows) but restarting a wedged or misbehaving service/daemon is a pretty common thing.

2h agoHN ↗

Yup. Could be that everyone was rushing to get all their products and demos ready leading up to it.

52m agoHN ↗

Most places at this scale have code freezes in place well before conferences. The most likely issues are some launch couldn't handle the scale or periodic deployments have been saving them from some sort of long-standing leak bug, and pausing going into Dreamforce meant some service hasn't been restarted in a week. Historically, Salesforce sharded by customer, so that goes against both of these, unless it's in a routing layer.

3h agoHN ↗

This is what happens when more than half the company is away attending the Salesforce cult-indoctrination stuff while spending all their bandwidth making customers/partners feel good.... The stuff that matters to keep the lights on gets overlooked.

3h agoHN ↗

Arguably, making customers and partners feel good is the more important part of the business

3h agoHN ↗

They wont feel good if the product they pay for doesnt work

2h agoHN ↗

From what I can tell, the business model is "it doesn't work, but you can pay folks exorbitant fees to 'customize' it for you"...

Also see: Oracle

3h agoHN ↗

arguably, this is what sales cares about and the half-measures taken to tackle what must be the Mt. Everest of tech debt at Salesforce is what leads to large, systemically degraded customer trust in products that keep shipping bugs

2h agoHN ↗

I don’t think engineering and SRE of the organizing company are ever invited to those events. They’re mainly for marketing and sales (which includes solution architects).

2h agoHN ↗

this event is only for customers. its not a company event.

1h agoHN ↗

Say what you want about the product and leadership... they do throw a good conference.

4h agoHN ↗

Cause: Legacy Salesforce login service got into a resource-exhaustion cascade.

Fix: Rolling some unspecified fix they proved in testing out over the fleet seemingly very slowly (After their earlier attempts to roll something out faster failed).

Details at https://status.salesforce.com/incidents/20004433

2h agoHN ↗

I wonder if “legacy login” is the shared login gateway.

It’s optional but everyone uses it. And it was flaky for an hour or so, like two months ago.

3h agoHN ↗

At this point OpenAI really ought to let us know when they're testing again.

3h agoHN ↗

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

3h agoHN ↗

I honestly don’t get the snark. The status page has:

Seemingly meaningful IDs

Search

Region filter

Email update signup

Predictable URLs for instance status so they can be deep linked in runbooks

What appears to be the actual live instance status.

What appears to be the actual live service status in each instance.

An update log with frequent detailed updates.

1h agoHN ↗

I despise Salesforce, but when I landed on this page I was like, huh. wow. honesty. Looks at GitHub

So yeah you're exactly right, the snark is not deserved if you ask me, and I'm 82% snark.

3h agoHN ↗

This is the case for every single B2B saas product. This is like the "bar is rolling on the floor" level of competence required. Please have higher standards for paid products.

1h agoHN ↗

Terrible, I'd argue the vast majority of modern big tech offerings are extremely poor quality where the need for surveillance in the form of constant monitoring/advertising metrics deliberately makes these types of services more costly to maintain and repair over time.

Sure there are like 3 or 5 decent services out there (like S3) but the vast majority are over engineered to be user hostile while extracting out whatever resources they can from their customers.

2h agoHN ↗

Hacker News is much easier to read when you realize that 95% of people have never worked on a "high" (maybe we could say >1B requests per day as a starting point) scale distributed service and think it's trivial to run one with more than 2 nines. You see comments all the time here mentioning that their own desktop at home is achieving more than that which belies deep misunderstanding of how systems are measured. Or that unofficial github status page repeatedly posted here that counts all github services together into one number.

1h agoHN ↗

which belies deep misunderstanding

I think you are missing the point. When I state my Exchange server is more reliable than Exchange Online, I don't think I'm a better engineer. I recognize Microsoft has harder problems to solve than I do. I think building overengineered, oversized SaaS environments is introducing extreme risk. It's an inherent flaw of the current approach.

Smaller is, in fact, better, because it's easier to operate reliably.

1h agoHN ↗

Indeed, the scale Anon1096 refers to wrt distributed systems is anti pattern. It is designed to vacuum up revenue and create enterprise value with scale, not to create resiliency for customers (although resiliency might be a byproduct of a well architected and operated distributed system at scale).

"Simplicity is the ultimate sophistication." -- Da Vinci

32m agoHN ↗

Hidden in this discussion around self-hosting reliability are other options as well.

Depending on your time and appetite for tinkering with all of this, it's not hard to imagine a home setup that fails over to a cheap Hetzner or DO VM. A manual failover at the DNS level isn't overly complex, and could be scripted.

Keeping a database in sync between home and the instance might be simple or more complex depending on needs, but would it really be that hard to have Claude help you setup a replicating Postgres server? If your database (or data files) are 1 gigabyte and don't update that often... maybe just rsync it every night or something

There's a thread you and others are pulling on here, and we need to pull it. Hosting doesn't have to be the domain of the big vendors anymore.

1h agoHN ↗

Is it? When your internet is out for five days because your ISP takes a few days to get to you, do you acknowledge that you're now at 98.5% availability for the year, far worse than any SaaS email service?

I think people forget that those large environments are there for a reason. To make sure the service stays up in the face of problems outside your own control.

1h agoHN ↗

In my entire adult lifetime (mid 40s), my ISP has never been out for five days. Compare to Github, Microsoft, Salesforce, and AWS outages that are always occurring in some fashion. Reddit is down constantly in various ways and still continues to operate as a business, public no less, so I disagree about the need to chase five nines and broadly speaking, large distributed systems that are potentially unnecessary for the use case and target outcome.

https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

1h agoHN ↗

Consider yourself lucky that you’ve never been the victim of a fiber cut. But what about if the power to your house goes out? Or what if your server blows the power supply?

My entire point is that you have no redundancy in your system and you also aren’t big enough to have any pull with the vendors who can fix these types of outages so you’re basically at the mercy of your providers with no recourse.

That’s why these systems are built the way they are.

And generally four nines is considered the gold standard these days. I can tell you for sure that both Netflix and Ebay would lose money anytime they drop below four nines because I have at some point been responsible for both. You’re correct that Reddit has a lot more leeway and outage time before they start losing money but not that much leeway.

1h agoHN ↗

I've contributed to building out data centers, as well as managed colos for others, primarily in downtown Chicago at Level3 and at 350 E Cermak. I am familiar with architecture required for reliability and diversity, from power and fiber in all the way up the stack to the Kubernetes cluster and software defined networking. If you participate in the capital markets, your data traverses systems I've participated in designing and implementing. There is a time and place for complexity (in this context, large/global distributed systems), but too often, complexity exists where it need not (imho).

"What are you optimizing for?" is always an important question.

1h agoHN ↗

I've dealt with a fiber cut, it wasn't nearly that bad. Fiber cuts impacting my SaaS providers were worse because there was nothing I could do about it.

1h agoHN ↗

Curious, what AWS outage has affected you for days?

(I hope you'll agree that the middle east outage is a true outlier)

1h agoHN ↗

While not a home-run server, the NTP system is a distributed service that receives 100 billion to trillions of requests per day, and it's running pretty smoothly - it's never gone down completely since it started in 1985. It's also very simple. The reason it has so many 9's uptime is because it is simple. Given a low amount of complexity, it's not unreasonable to think that an individual could run a >1B requests per day service.

Salesforce is not simple. It's wildly, overly complex. It's amazing it has any 9's at all and not 8's or 7's. Salesforce offers three 9's, which allows for 43 minutes downtime per month. The current outage is at 8 hours (and counting) so Salesforce is now at 98.9% uptime for the month - there's an "8" in there now. Not good, but considering the complexity of Salesforce, it's still kind of amazing.

2h agoHN ↗

Hmm, could the use of genAI have anything to do with this failure and the inability to quickly fix it?

1h agoHN ↗

It's not impossible, but Salesforce has had big outages before LLMs. For a disruption that began at 1am pacific, the response time isn't that bad. 3 hours total to give up on restarts, 4 hours total to validate a quick fix and begin rollout, and the rest of the time since has been waiting for the rollout + addressing subsets of instances that had some issues with restarting+the quick fix. It's nearly 9am pacific now, so Dreamforce is saved~ (It's Dreamforce week this week. Most devs are either focused on that or on soft-vacation / working on lower priority non-feature-work items, it's surprising anything would be updated to production this week that could do this.) The architecture and approval process of everything there has long been setup so that things can't be changed quickly.

2h agoHN ↗

What I'm curious about is why it is a single-PaaS; I'd have expected Salesforce to have the customers quite isolated so the chance of bringing down multiple customers at once was much smaller.

1h agoHN ↗

a few things are global (login service) to some extent, just as AWS places several such things in us-east-1

3h agoHN ↗

Kind of ironic. Salesforce is basically one of the major spiritual grandfathers of Slop. It is not uncommon in production systems to find that objects like Contact and Account have hundreds of custom fields. Sometimes, you find out that several of them have the same meaning and semantics, but were used at different times. Digging out you discover that some Marketing guy that used to work at the company did some task in a certain way that was lost when he was gone, and then a few months later his substitute had the same need and went ahead and created the same field with a slightly different name.

Doing data engineering work with Salesforce data is an exercise on archeology, psychology and organizational politics.

Slop is basically the ontological and teleological philosophy behind Salesforce very existence. Despite the official discourse that the "No Software" meant no infrastructure, no toil with updates and configuration, the subtext as intended for executives was very clear: "No need for you to be blocked by those pricks from engineering and their stupid, bureaucratic and gatekeeping rules".

"No software" was a call-to-arms to a certain subset of managers that were radicalized by Nicholas Carr's 2023 HBR article "IT Doesn't matter". It doesn't matter that Carr was a journalist and a writer with a masters in English that has never ever run even a small bodega, or has never managed an IT department. Anti-intellectualism and the abundance of capital brought in by the petrodollar that allowed the US government to run deficits year by year while exporting the ensuing inflationary effects to rest of world, would ensure that this message would ressonate and then even be amplified during the years of ZIRP and the Baillouts. Play fast and loose, first come, first served, a rising tide rises all boats and all that jazz. Wall Street favors bold, and the heck with the long term! This quarter will only live once!

Frankly, this is just poetic justice: Kill by slop, be killed by slop.

2h agoHN ↗

Something tells me Troy the Salesforce Admin/BD Analyst did not cause the SAAS infrastructure to go down.

And I think you're confusing crud with slop.

2h agoHN ↗

I've never worked somewhere that had a Salesforce integration which wasn't an eternal disaster. Have you?

Why is every company's Salesforce team absolute bottom of the barrel developers with super high churn, no responsibility, and little competency?

Something about the product and its positioning attracts catastrophe. That's what GP is talking about.

2h agoHN ↗

Typo: Carr wrote that in 2003, not 2023 for those of you who missed the foolishness and insane wreckage that article caused.

The lost business value and competitiveness caused by outsourcing IT overseas to unmotivated parties under Carr's premise is hard to put your finger on but I have seen the aftermath and it's pretty massive.

2h agoHN ↗

I can't understand how such a huge company can have such a lousy UX.

2h agoHN ↗

Peak Tech Salesforce was 2010 +/- 2 years - i.e. after Visualforce and before Aura era. It used to be a developer oriented platform and it became shiny/flashy garbage eventually. But all these shiny things allowed them to get a large market cap with very brilliant sales people, it's hard to deny.

1h agoHN ↗

VF and Aura overlapped. Aura was just a bad start and janky. We sometimes just did React instead, for a while.

LWC is worlds better. And the local tooling with the cli and VSCode extensions is miles better than the old Eclipse/Sublime FMT days.

16m agoHN ↗

Aura existed simply because Salesforce thought to be smarter than open source, well. Can't deny it though: Salesforce engineers were great on the backend, but frontend dev has never been their thing. Back then there was Angular 1 which was miles ahead. React was released shortly after Aura itself, so to say. LWC is what Aura shall have been 10+ years ago. And that ties back to what OP wrote: awful UX extremely slow bloated with JS.

Now, don't talk me about VSCode Extensions. This is the perfect example of an awful dev experience. apex-jorje-lsp.jar with a JVM to parse Apex taking GB of memories, extensions taking dozens of seconds to load (when they load) ... In fact, the only decent LSP is aer, a simple decently working Go binary rather than the monster Salesforce shipped. The one good tooling Salesforce built in the last 15 years is, to some extent, the SF CLI - which came after the `force` CLI from the same guys who built `aer`, anyway. And nowadays, people can use that with their preferred editor from Zed to Vim with shortcuts from built upon the SF CLI.

So no, Salesforce didn't do great with tooling, they just did the bare minimum waiting on the (small) community to give them the right ideas.

2h agoHN ↗

I'm sure the cause of this outage will not be connected to vibe coding in any way

2h agoHN ↗

Salesforce has turned into the monolithic messy bloatware it set out to replace.

Please VCs stop with the AI FOMO and find a few good startups to just go destroy Salesforce and give folks a simple inexpensive replacement.

2h agoHN ↗

As someone who has never used Salesforce nor hubspot. What are the core features that has people using these services? Is it just the integration between all the different areas where customer data lives?

2h agoHN ↗

It’s like a lot of these big SaaS platforms that lock people in to big contracts. They promise the world with all sorts of features and integrations but most customers don’t use that and get locked into buying a very expensive solution to a very simple problem (in this case tracking sales pipelines).

Where customers do use the features and integrations it’s often a giant mess that needs a whole separate ecosystem of consultants and “partners” to get the thing working and maintaining it.

45m agoHN ↗

They're an incredible marketing machine.

But mostly it's a mix of the integration network effects you mention, cost of reimplementation if you want to leave, and the good old "nobody gets fired for buying IBM" dynamic.

1h agoHN ↗

There are 10,000 “simple inexpensive replacements”. Always have been. If you need a glorified three object contact list… you shouldn’t buy Salesforce.

Those just, obviously, can’t do almost any of the forty million serious things that Salesforce does, and real businesses do need.

54m agoHN ↗

This is an honest question: which serious things?

2h agoHN ↗

Does it mean I won’t get any AI generated slop spam for a few hours?

1h agoHN ↗

I was just working on a zigpoll integration and thought it was me...

1h agoHN ↗

RPC errors in my Dunkin app and my wife's Poshmark app. Is it an infrastructure problem?