I recently read a few articles about how the harness is a bigger factor to successful LLM usage and wish they discussed this here.
I use GLM-5.3, Qwen3.8, Claude (all of 'em), GPT Sol/Luna/Terra across direct API calls + local models where I can (128GB Macbook Pro)... The harness and whether the model or underlying system prompts know how to make the best use of iterative LLM calls makes such a big difference...
For example: one-off articles on a news topic (e.g., "Update me on the US-Canada relations") yields very similar results across all models... But "run a web search, write a draft perspective from three points of view, and structure data around it" will make everything but Claude + GPT struggle.
I recently read a few articles about how the harness is a bigger factor to successful LLM usage and wish they discussed this here.
I use GLM-5.3, Qwen3.8, Claude (all of 'em), GPT Sol/Luna/Terra across direct API calls + local models where I can (128GB Macbook Pro)... The harness and whether the model or underlying system prompts know how to make the best use of iterative LLM calls makes such a big difference...
For example: one-off articles on a news topic (e.g., "Update me on the US-Canada relations") yields very similar results across all models... But "run a web search, write a draft perspective from three points of view, and structure data around it" will make everything but Claude + GPT struggle.
The data and graphs are great, but it would have been a much higher quality report if the text and titles were written by a human.