- 55comments
- 9comments
- 187comments
- 6comments
- 92comments
- 102comments
- 453comments
- 2comments
- 2comments
- 218comments
- 36comments
- 65comments
- —discuss
- 15comments
- 254comments
- 2comments
- 4comments
- 80comments
- 36comments
- 29comments
- 336comments
- 1comments
- 72comments
- 169comments
- 307comments
- 29comments
- 428comments
- 18comments
- 97comments
- 49comments
Estimating the cost of unaccepted runs (those that were not successful?) is left as an exercise for the reader.
Apparently they updated it perhaps based on your comment?
Oh - I see that now. It's possible I missed it originally - I did read the article but I was skimming quickly. Mea culpa, if so.
You can have 100 runs for the price of the Claude Max plan?
I find this - or perhaps the title - a bit surprising.
I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.
Perhaps DS works better where targets have low-hanging fruits than can be found fast?
In your example, couldn't you parallelize DS's work more ? You could have 11 times as many agents for the same price.
When the history books are written and all is said and done, the hubris of this moment where all the American labs decided to punk their investors and join hand in hand in agreeing to let the Chinese win forever is going to be the main story.