It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.
9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
It's a standard term and has been for ages. It distinguishes end to end time vs eg the amount of CPU time an individual process uses (which excludes time spent waiting for the system or time when the process was otherwise not scheduled on the CPU).
“Wall time” is a pretty common systems term (you see it when comparing total runtime to kernel time, for example). I wouldn’t have indexed on “wall clock time” or other variants as an LLMism.
It’s great, but you still need to know what you are doing, not just the goal/s. I built a platform years ago from scratch, and now I am remaking it with more features and more polished design, I know exactly what needs to be done to tiniest details. The first prompt was very well detailed about the architecture and how everything should work, after an hour work at Xhigh, it did create the blueprint artifacts that I asked for, then I spent 3 days reading every single thing and writing notes, turned out it made the system overly complicated without adding extra value, plus I can see how some of the architecture design will be potentially a security risk. So after few days I fed my notes, this time took 4hours and 1M tokens! Later I spent more few days reviewing and writing notes, it was closer to what I want but still made architecture errors, the third run took around an hour and finally made it how it supposed to be, although there are still more notes on non critical stuff. So I don’t think we are yet at the stage where sitting goals and some high level is enough to produce quality results.
Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
I've given it some big tasks and asked it to parallelize as much as possible etc.
It did burn through my weekly tokens in about a day (20x max), but the output was completely on point.
(I knew there was a "reset token usage - opus 5.5" button in my account.)
I've now come to a point where I even delegate my discovery for new features to it.
You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.
Opus 5.5 is great and cheaper if you compare to fable with close quality in coding (tested in refactoring java to nodejs), but I do not understand why the week before the release Opus 5 started hallucinating (long loop and waste of token for single tasks)
It’s amazing at debugging too. I had it running in Powershell controlling an lldb session in MSYS2. The way it can read addresses and so root cause analysis is amazing! It takes a lot of mind power to do those things.
I like to watch it work though because it honestly teaches me some tricks.
I think what makes it great is that they trained it to write harnesses for the code it writes, so it can test stuff even if the supplied code is not complete.
It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.
9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
Is "wall-clock" an actual term you used before Claude? I had never heard it before the model used it and I can't stand it.
It's a pretty old term, to distinguish from e.g. CPU time. This was in common usage even 30 years ago.
Example from 15 years ago: https://stackoverflow.com/questions/7335920/what-specificall...
It's a standard term and has been for ages. It distinguishes end to end time vs eg the amount of CPU time an individual process uses (which excludes time spent waiting for the system or time when the process was otherwise not scheduled on the CPU).
I've used wall clock for many years, normally when compared to CPU time when talking about parallelizing some process.
CPU time might go up while wall clock time goes down
“Wall time” is a pretty common systems term (you see it when comparing total runtime to kernel time, for example). I wouldn’t have indexed on “wall clock time” or other variants as an LLMism.
It’s great, but you still need to know what you are doing, not just the goal/s. I built a platform years ago from scratch, and now I am remaking it with more features and more polished design, I know exactly what needs to be done to tiniest details. The first prompt was very well detailed about the architecture and how everything should work, after an hour work at Xhigh, it did create the blueprint artifacts that I asked for, then I spent 3 days reading every single thing and writing notes, turned out it made the system overly complicated without adding extra value, plus I can see how some of the architecture design will be potentially a security risk. So after few days I fed my notes, this time took 4hours and 1M tokens! Later I spent more few days reviewing and writing notes, it was closer to what I want but still made architecture errors, the third run took around an hour and finally made it how it supposed to be, although there are still more notes on non critical stuff. So I don’t think we are yet at the stage where sitting goals and some high level is enough to produce quality results.
9 hours?!
Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
Meh... There's a reason Opus 4.6 is still an option.
Pros know these are lower cost models.
Superb model indeed.
I've given it some big tasks and asked it to parallelize as much as possible etc.
It did burn through my weekly tokens in about a day (20x max), but the output was completely on point. (I knew there was a "reset token usage - opus 5.5" button in my account.)
I've now come to a point where I even delegate my discovery for new features to it.
You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.
"Don’t ask it to show its reasoning in the reply" “Explain why you chose this approach in three sentences” says it all really
Opus 5.5 is great and cheaper if you compare to fable with close quality in coding (tested in refactoring java to nodejs), but I do not understand why the week before the release Opus 5 started hallucinating (long loop and waste of token for single tasks)
Agreed on Opus 5.5 being a great model. It's the first one that I trust for long running (>1 hour) tasks.
It’s amazing at debugging too. I had it running in Powershell controlling an lldb session in MSYS2. The way it can read addresses and so root cause analysis is amazing! It takes a lot of mind power to do those things.
I like to watch it work though because it honestly teaches me some tricks.
I'm convinced gpt 4.5 was the best model humanity has ever made.
Now we are getting downgraded models that do 100x COT because it's cheaper.
Phenomenal model, not sure what they did, but I have been able to do so much with my $20 plan!
I think what makes it great is that they trained it to write harnesses for the code it writes, so it can test stuff even if the supplied code is not complete.
If enough people keep saying it I'm sure they will nerf it