One of the Anthropic salesmen told everyone to use loops all over the place, auto mode and use Fable 5 and Opus 5 with multiple sub agents.
Then it seems they forgot about their own infrastructure which goes down once every two weeks.
Looks like most of all the engineering knowledge that was needed at Google was lost after the whole org at Anthropic is now vibing their work.
Rather than screaming “coding is solved!”, “AGI with a month” and spooking everyone with the bogus mysticism of “Mythos” which now everyone has an equivalent strong cyber model, maybe keep the lights on first before plotting the next doomsday story.
Because coding is solved and there’s no need for software engineers any more.
Being charitable to Anthropic though: coding is sort of solved for a very limited definition of “solved”.
Software engineering of huge distributed systems at massive scale is very much not solved, and it seems Anthropic don’t have the human capital to do it particularly well
If we want to be even more charitable to Anthropic, they're going from nothing to hyperscaler in like 3 years. I'm as happy as anyone to dunk on the general bubble situation and people claiming their technology is too dangerous to share, then something like this happens, it's hilarious. But it's still actually hard to build a service with that much usage and every org has had some outages on their way to figure it out.
Anthropic's product is inherently unreliable even when operating at its best. Why would they bother with reliable delivery when their customers are the type clearly not interested in reliability?
I think it's because customers don't seem to mind enough to switch.
I really want to support Anthropic because of their commitments to effective altruism, but it's getting harder and harder to justify on technical merits. Sol is just faster, cheaper, and more reliable, and as good as Opus for my use cases
> I really want to support Anthropic because of their commitments to effective altruism
If you're really committed to the ideas of EA, shouldn't you dispassionately select whatever model works best for you, then make as much money as possible using that and donate it? :)
I've recently started using codex for most of my coding. Haven't looked back so far: It seems to reason deeper about the code it writes and has much less downtime.
> I really want to support Anthropic because of their commitments to effective altruism
They are doing the opposite while saying the other. Zero trust about that. They are using every malicious playbook tactic these days to increase users and be more profitable and lock-in the users. Even OpenAI feels better these days.
I mean, that’s exactly what “effective altruism” is about? Getting filthy rich and powerful no matter the methods and then the vague promise of doing something good?
Not really. They seems to chase IPO so that alone proves that they can’t hold any promise about doing something good. They would need to stay as private company to have any power to do some ”good” in the end.
Estimate of the number of people in that latter group? Nick Beckstead is probably the cleanest example of an extreme longtermist, but even he has taken the GWWC pledge. Similarly, Nick Bostrom argues in a pretty longtermist way and is neither vegan nor has taken the pledge, but I don’t see any particular evidence that he has a morality that an average person would call evil. Toby Ord and Will MacAskill are both visibly mundanely extra-super-ethical despite their longtermist views.
> Notably, people in the latter group never assume THEY will be the ones suffering for the future.
What makes you think so? What evidence do you have?
> Yes, some people think effective altruism means helping people here and now as efficiently as possible.
Most effective altruists seem pretty relaxed about the 'here' part, but you are right that the 'now' part receives different answers from different subgroups.
> Others think the sheer number of potential sentients in the future justify any amount of suffering tolerated in the here and now.
There's also the people who care a lot about shrimp. They are less crazy than it seems at first.
Only if you're incapable of throttling to assure quality of service. Are you saying they have so much demand that even their load balancers were overloaded?
Throttling helps to prevent a 30% failure rate (with a quick recovery time) turning into a 100% failure rate (with practically infinite recovery time), but some users are still going to see failures, and the news is still going to be "Anthropic is down again" because some users are seeing an outage.
I'm not really sure how they can avoid getting in the news with "Anthropic is down", TBH. Although, at least by throttling they might get "Anthropic is down for some requests" and not "Anthropic is down for all requests".
There is no point having this discussion. Some people have had their brains genuinely broken by Anthropic and will defend them to the last breath, despite having no clue what they are talking about.
It is a reliability problem - because if the API ends up rejecting 30% of the incoming queries, no one cares if it's 500 Internal Server Error or 529 Overloaded.
Having your infrastructure at the knife's edge of load to capacity also means you have no redundancy when something fails.
If this incident had happened during the day I could buy that argument. But it is happening in the dead of the night. Are there really that much people scheduling claude runs over-night, taking more capacity than day-time? (assume most claude users are west coast programmers)
My other guess is they are scheduling training runs and their capacity isolation is bad.
Clearly they just want to be down all the time. They have the ability to throttle down traffic. I have no clue why they decide to throttle it down just enough to have the worst of both worlds, throttling and incidents. Throttle more and have zero incidents, what an incredible revelation.
They didn’t want to be down so bad they rented capacity from Elon and uptime improved dramatically. Super flaky before that, they were even playing games with response quality to lighten their load nontransparently. Not sure what the issue is today. But on average, past 90 days, good availability and quality right?
"Coding is solved", how ironic.
Bad coding is solved.
Great reminder to be using multiple providers (i.e. Anthropic 1P, Bedrock, Vertex AI, Azure) with auto-fallback.
Bedrock has been fine throughout today; just pick two providers and you have much more stable Claude :)
Also, sometimes older models work fine, Opus 4.8/4.6/4.5 are worth trying.
Never put all AI taxes in one basket!
Imagine being left behind because you don't have an API, network connection, or credits to burn; shame.
How they go down globally? All regions use the same infra?
It was a configuration error ...
One of the Anthropic salesmen told everyone to use loops all over the place, auto mode and use Fable 5 and Opus 5 with multiple sub agents.
Then it seems they forgot about their own infrastructure which goes down once every two weeks.
Looks like most of all the engineering knowledge that was needed at Google was lost after the whole org at Anthropic is now vibing their work.
Rather than screaming “coding is solved!”, “AGI with a month” and spooking everyone with the bogus mysticism of “Mythos” which now everyone has an equivalent strong cyber model, maybe keep the lights on first before plotting the next doomsday story.
> Looks like most of all the engineering knowledge that was needed at Google was lost after the whole org at Anthropic is now vibing their work.
Sorry, what does Google have to do with Anthropic's outage?
They also serve inference from GCloud
Anthropic has constant failure since morning.
Yes yes it is key that we vibe code critical infrastructure. I can't wait to see a claude-ism in my next BIOS update.
They cant solve the infrastructure problem until Claude is back up....
Anyone can explain why do they fail so often and so seriously? Is this inherent to managing LLMs or is their Ops really just that bad?
Because coding is solved and there’s no need for software engineers any more.
Being charitable to Anthropic though: coding is sort of solved for a very limited definition of “solved”.
Software engineering of huge distributed systems at massive scale is very much not solved, and it seems Anthropic don’t have the human capital to do it particularly well
If we want to be even more charitable to Anthropic, they're going from nothing to hyperscaler in like 3 years. I'm as happy as anyone to dunk on the general bubble situation and people claiming their technology is too dangerous to share, then something like this happens, it's hilarious. But it's still actually hard to build a service with that much usage and every org has had some outages on their way to figure it out.
Anthropic's product is inherently unreliable even when operating at its best. Why would they bother with reliable delivery when their customers are the type clearly not interested in reliability?
I think it's because customers don't seem to mind enough to switch.
I really want to support Anthropic because of their commitments to effective altruism, but it's getting harder and harder to justify on technical merits. Sol is just faster, cheaper, and more reliable, and as good as Opus for my use cases
> I really want to support Anthropic because of their commitments to effective altruism
If you're really committed to the ideas of EA, shouldn't you dispassionately select whatever model works best for you, then make as much money as possible using that and donate it? :)
Just like the coding is solved by redefining what solved means, so is the current iteration of "altruism".
Do you actually believe in their altruistic goals?
The goals are altruistic. The torrent client was purely instrumental.
I've recently started using codex for most of my coding. Haven't looked back so far: It seems to reason deeper about the code it writes and has much less downtime.
> I really want to support Anthropic because of their commitments to effective altruism
They are doing the opposite while saying the other. Zero trust about that. They are using every malicious playbook tactic these days to increase users and be more profitable and lock-in the users. Even OpenAI feels better these days.
I mean, that’s exactly what “effective altruism” is about? Getting filthy rich and powerful no matter the methods and then the vague promise of doing something good?
Not really. They seems to chase IPO so that alone proves that they can’t hold any promise about doing something good. They would need to stay as private company to have any power to do some ”good” in the end.
> their commitments to effective altruism
I think you mistook their declamations for commitments.
> commitments to effective altruism
That's a downside if anything.
Different folks have different preferences.
Yes, some people think effective altruism means helping people here and now as efficiently as possible.
Others think the sheer number of potential sentients in the future justify any amount of suffering tolerated in the here and now.
Notably, people in the latter group never assume THEY will be the ones suffering for the future.
Estimate of the number of people in that latter group? Nick Beckstead is probably the cleanest example of an extreme longtermist, but even he has taken the GWWC pledge. Similarly, Nick Bostrom argues in a pretty longtermist way and is neither vegan nor has taken the pledge, but I don’t see any particular evidence that he has a morality that an average person would call evil. Toby Ord and Will MacAskill are both visibly mundanely extra-super-ethical despite their longtermist views.
> Notably, people in the latter group never assume THEY will be the ones suffering for the future.
What makes you think so? What evidence do you have?
> Yes, some people think effective altruism means helping people here and now as efficiently as possible.
Most effective altruists seem pretty relaxed about the 'here' part, but you are right that the 'now' part receives different answers from different subgroups.
> Others think the sheer number of potential sentients in the future justify any amount of suffering tolerated in the here and now.
There's also the people who care a lot about shrimp. They are less crazy than it seems at first.
Because they can’t keep up with the demand?
Struggling with demands is a performance problem, not a reliability problem.
overly high demand can often put a system in an unstable state
Only if you're incapable of throttling to assure quality of service. Are you saying they have so much demand that even their load balancers were overloaded?
Throttling helps to prevent a 30% failure rate (with a quick recovery time) turning into a 100% failure rate (with practically infinite recovery time), but some users are still going to see failures, and the news is still going to be "Anthropic is down again" because some users are seeing an outage.
I'm not really sure how they can avoid getting in the news with "Anthropic is down", TBH. Although, at least by throttling they might get "Anthropic is down for some requests" and not "Anthropic is down for all requests".
There is no point having this discussion. Some people have had their brains genuinely broken by Anthropic and will defend them to the last breath, despite having no clue what they are talking about.
I'll leave it to the retards to confidently claim load has no bearing on the stability of a distributed system.
Load is not demand.
It is a reliability problem - because if the API ends up rejecting 30% of the incoming queries, no one cares if it's 500 Internal Server Error or 529 Overloaded.
Having your infrastructure at the knife's edge of load to capacity also means you have no redundancy when something fails.
If this incident had happened during the day I could buy that argument. But it is happening in the dead of the night. Are there really that much people scheduling claude runs over-night, taking more capacity than day-time? (assume most claude users are west coast programmers)
My other guess is they are scheduling training runs and their capacity isolation is bad.
The world is more than just the Americas.
Clearly they just want to be down all the time. They have the ability to throttle down traffic. I have no clue why they decide to throttle it down just enough to have the worst of both worlds, throttling and incidents. Throttle more and have zero incidents, what an incredible revelation.
They didn’t want to be down so bad they rented capacity from Elon and uptime improved dramatically. Super flaky before that, they were even playing games with response quality to lighten their load nontransparently. Not sure what the issue is today. But on average, past 90 days, good availability and quality right?
Everything is vibe coded, would you expect anything different?
Quick, release a new claude product then depricate it after 90 days
Mind you, They use latest Anthropic Claude models for programming. /s
It's a reminder to never depend of something as flaky as AI on the Internet for your important business processes.
my senior suggested quoting prices based off an online LLM fed a rules table and what the customer did :thinking:
Lately Dario has to many bad news from customer side. Of course there is no bad or good news, just marketing, but still…