- 206comments
- 100comments
- 25comments
- 120comments
- 37comments
- 174comments
- 22comments
- 3comments
- 34comments
- 4comments
- 103comments
- 26comments
- 169comments
- 3comments
- 292comments
- 6comments
- 254comments
- —discuss
- 27comments
- 240comments
- 6comments
- 29comments
- 18comments
- 136comments
- 11comments
- 8comments
- 175comments
- 43comments
- 74comments
- 82comments
I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?