Qwen 3.5 9B scores 2-3x higher than 4o (depending on the 4o version), on the benchmarks.
Whether it's actually better for the kind of things people actually use it for... the benchmarks don't really tell you that. (In my experience, no.)
I often have funny experiences where models do great on benchmarks and are awful, or do poorly and are great for my use cases.
And different people use them in different ways, which probably explains why some people think one models is great and others think it sucks.
In my experience even small local models are now surprisingly good at programming and using a computer (bash), i.e. completing agentic tasks, but fall apart quickly in conversation (especially knowledge and understanding).
Whether it's actually better for the kind of things people actually use it for... the benchmarks don't really tell you that. (In my experience, no.)
I often have funny experiences where models do great on benchmarks and are awful, or do poorly and are great for my use cases.
And different people use them in different ways, which probably explains why some people think one models is great and others think it sucks.
In my experience even small local models are now surprisingly good at programming and using a computer (bash), i.e. completing agentic tasks, but fall apart quickly in conversation (especially knowledge and understanding).