One thing about these numbers that's absolutely shocking to me is how low the energy use is:
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's driving miles. It's boiling 10 gallons of water.
With the talk of AI Data Center's impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
Your only talking about variable direct energy. Does it take into account the entire lifecycle, building the data centre, running the cooling, building the chips, the % the chips are not utilised.
I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
> Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
Hmm I might need to rephrase. Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget (big price difference between models)
It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control
Flash is pretty decent coder, but it should be paired with good planner and reviewer. I would pick astra low for planning and sol 6.1 medium for reviews.
One thing about these numbers that's absolutely shocking to me is how low the energy use is:
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's driving miles. It's boiling 10 gallons of water.
With the talk of AI Data Center's impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
Your only talking about variable direct energy. Does it take into account the entire lifecycle, building the data centre, running the cooling, building the chips, the % the chips are not utilised.
I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
> Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
Hmm I might need to rephrase. Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget (big price difference between models)
GLM5.3-flash has been fantastic for me to make minor fixes in ambigious ways. "Fix x feature, whats going wrong. " It does the job.
GLM-5.3-flash is my implementation model after GLM-5.3 writes the plan.
It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.
Was this post generated with LLM, did he properly mention anywhere why exactly did it fail with example or i have trouble reading.
How did you measure energy usage?
Edit: I found a linked article that mentions the inference provider who does the measurements.
Yes, all from Neuralwatt, GPU energy use only. Makes models’ “efficiency” much more visible than tokens.
It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control
Flash is pretty decent coder, but it should be paired with good planner and reviewer. I would pick astra low for planning and sol 6.1 medium for reviews.