Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Astra for Law(openai.com ↗)
    303comments
  2. Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint(prismml.com ↗)
    53comments
  3. Bend – A language that blocks AI mistakes via proof, on CPU and GPU(bend-lang.com ↗)
    129comments
  4. Hister: A private search engine for the pages you visit and the files you keep(github.com/asciimoo ↗)
    130comments
  5. Wax motor(wikipedia.org ↗)
    42comments
  6. Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA(global.fujitsu ↗)
    190comments
  7. More than 100k people in Japan are now aged 100 or older(bbc.com ↗)
    56comments
  8. Flet 1.0 – Build cross-platform apps in Python(flet.dev ↗)
    15comments
  9. Diplodocus, Long Thought Exclusively American, Turns Up in Spain(sci.news ↗)
    15comments
  10. CrowdSec Source Code Leak(crowdsec.net ↗)
    35comments
  11. Landing the Space Shuttle – A Flying Machine and the Thrill of a Lifetime(eaa.org ↗)
    2comments
  12. Rate limits on GitLab.com are changing(about.gitlab.com ↗)
    105comments
  13. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data(arxiv.org ↗)
    29comments
  14. How Uber Protects Against Retry Storms(uber.com ↗)
    15comments
  15. How GLM built its own inference infrastructure(z.ai ↗)
    260comments
  16. Why I didn’t sign the Fields medallists’ letter(gowers.wordpress.com ↗)
    275comments
  17. TSMC revealing details about next gen A14 node(mapyourshow.com ↗)
    31comments
  18. How do we prevent mathemathics from devolving into the Medieval Era of secrecy?(mathoverflow.net ↗)
    46comments
  19. The American Religion of Self-Storage Facilities(newyorker.com ↗)
    320comments
  20. Plugin4Shell – Zero Click RCE Vulnerability found in top four coding agents(air.security ↗)
    2comments
  21. Zettascale (YC S24) Is Hiring ASIC/FPGA Engineers to Build Chips for ASI(zscc.ai ↗)
    discuss
  22. CCC invites all model citizens to 40C3(ccc.de ↗)
    178comments
  23. The most important product decision is what you don't build(liamnugent.me ↗)
    14comments
  24. Running Ubuntu on the Lenovo IdeaPad Duet(vhaudiquet.fr ↗)
    23comments
  25. Show HN: Snapdrop: Instantly share files between devices. No setup, no signup(snapdrop.me ↗)
    12comments
  26. Sex, AI, and the Apocalypse(iankduncan.com ↗)
    110comments
  27. Launch HN: Skillsync (YC W26) – AI chat sessions made portable across agents
    46comments
  28. Show HN: Share your AI Setup, Learn from others(mysetup.ai ↗)
    89comments
  29. André Weil and the Hodge Conjecture(jiahao116.github.io ↗)
    10comments
  30. My temporary PHP fix from 2014 has nearly 20M installs. Today I'm deprecating it(jakeasmith.com ↗)
    96comments

How Uber Protects Against Retry Storms

28 pointsby 3h agouber.com
15 comments
2h agoHN ↗

I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?

1h agoHN ↗

This is trading a good developer experience for a bad user experience. There are situations where it makes sense to force manual retry, but there's no reason to apply one universal rule to all possible situations. Lack of considering nuance for your situation is just intellectual laziness.

43m agoHN ↗

My point was that just throwing exponential backoffs at the retry problem is not a magic solution

I don't follow how being cautious about avoiding multiplicative layers of backoffs is trading a good developer experience for a bad user experience. The described situation is an awful user experience. Simply adding a retry and calling it a day sounds like the easy developer experience at the expense of the user experience

39m agoHN ↗

The described situation is an awful user experience

Sure, the worst case scenario is. 99.9999% of the time, a transient error will actually just work on the first or second auto-retry and save your users the effort of paying attention and manually retrying things. This is especially prudent for background tasks where the failure may not be noticed right away; coming back to something fire-and-forget 30m later to see it never tried to finish is not a good user experience.

My point was that just throwing exponential backoffs at the retry problem is not a magic solution

Nobody said it was. In fact, I suggested the exact opposite - a proper solution takes dev effort. Adhering to an iron rule of "just make them manually retry" is throwing your hands up and not even trying to solve the problem because laziness is convenient.

17m agoHN ↗

Nobody said it was

I responded to a post that merely linked to the Wikipedia article for exponential backoff (in response to "I'd be interested to hear other strategies in [protecting against retry storms]")

The original article is precisely about the degenerate case and the difficult work of dealing with it

This approach works for transient or low-rate failures. However, during moderate or severe degradation, it becomes counterproductive. Aggressively retrying against an already struggling service increases load, accelerates failure, and amplifies retry traffic across upstream dependencies. What begins as a localized outage can quickly escalate into a stack-wide incident—ultimately degrading, or in the worst case, completely breaking, the end user experience.

1h agoHN ↗

Yes, exponential back off and jitter are the first things to work on, and good if you don’t have a better signal (like loss of network).

Also, a simple signal status server or queue system helps to keep global state such that everyone doesn’t retry all at once.

If you have a central error rate server you can skip your retry based on the error rate (100% error rate, don’t retry, etc).

1h agoHN ↗

So many variables, but the simple thing is to set things up like normal rate limiting (which you would want to do anyways). The one generating the errors passes back a retry time. You can add jitter here, tell low priority requests to wait longer, etc.

BTW: do keep track of priority. It’s like having a database that gets flooded with connections and won’t allow new ones in—but will for admin users (btw, it did not used to be that way in the early days of MySQL).

1h agoHN ↗

There's a good amount of literature about this (check the other comments), but you can vastly simplify this into two things you need to do:

1. Your service that retries should have some retry budget. This is a good place to be "smart", because you can reason entirely locally instead of turning it into a distributed systems problem. The best library I've seen for this was doing Exponential Moving Average of requests per second sent down that pipe (not counting retries) and only allowing 20% more requests per second as retries, total. Each individual request could be retried 3 times. This was critical as it bounds the additional load from retries.

2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load.

Everything else is nice-to-have, but those two alone should bound the total requests you get in a retry storm.

42m agoHN ↗

2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load.

Ooh, I like the idea of propagating "no retries" hints in the responses back upstream. Have you seen it implemented in the wild, or in public discussions about the practice?

30m agoHN ↗

I've only seen it in bigcorp cross-service typedefs, or in startup's code that re-implements the checks in every service.

1h agoHN ↗

I'm suspicious of load shedding not mentioned in the article. Combine that with exp backoff in the caller and you got yourself a pretty robust starting point

46m agoHN ↗

429 (and sometimes 503) errors returned by servers might well be a symptom of intentional load shedding. Perhaps it's just not explicitly called out as a server behavior that induces client retries.