Hacker News

Favorites Setup
Comment by simonw | original | Hy4 preview
[−]simonw · 2026-08-29 Sat 22:17 UTC · link
> [...] Let's maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn't ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no.

> Maybe add sunglasses? no.

> Maybe add water? no.

https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

[−]kurante · 2026-08-29 Sat 22:45 UTC · link
Is the broken English an optimization or a byproduct of the model being developed in China?
[−]minimaxir · 2026-08-29 Sat 22:48 UTC · link
Optimization. Why use many word when few word do trick?
[−]stavros · 2026-08-29 Sat 23:19 UTC · link
What I find funny about "why use many word when few word do trick?" is that it's only slightly shorter than the regular "why use many words when few words do the trick?"
[−]rapind · 2026-08-29 Sat 23:48 UTC · link
I always figured that was part of the joke, because a writer came up with it, and a writer would know (I assume?).
[−]inopinatus · 2026-08-30 Sun 07:54 UTC · link
The latter is not a complete alternative, it is ambiguously conflating vocabulary scale with word count, and also, it is not as funny
[−]happycube · 2026-08-30 Sun 10:22 UTC · link
less word better(, many words bad)
[−]andsoitis · 2026-08-29 Sat 23:33 UTC · link
> Why use many word when few word do trick?

Be concise.

  OR
Brief is best.

  OR
Eschew verbosity

  etc.
[−]gjvc · 2026-08-30 Sun 01:00 UTC · link
"Omit needless words."

-- William Strunk Jr. and E.B. White., The Elements of Style

[−]gaigalas · 2026-08-30 Sun 00:43 UTC · link
Optimization on a idiosyncrasy. The same thing that makes Claude repeat "That was the most important thing you said in this whole conversation" is what makes grug speak optimize on token usage.

Real humans get non-primary information from word variation. It's reasonable to hypothesize that it has a role in thinking things, because it endures. Our languages need to breathe over time, and flourishing might be one of the aspects that allows that breathing space.

[−]TiredOfLife · 2026-08-30 Sun 08:03 UTC · link
See world
[−]acheong08 · 2026-08-29 Sat 22:49 UTC · link
When GPT-5.6-sol's reasoning traces were leaked, they also used "caveman speak". Definitely a token efficiency optimization
[−]beefsack · 2026-08-29 Sat 23:31 UTC · link
I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?
[−]Barbing · 2026-08-29 Sat 23:37 UTC · link
"Neuralese"
[−]walrus01 · 2026-08-30 Sun 00:31 UTC · link
some people made a 'caveman' speak qwen as a joke

https://huggingface.co/ProCreations/grug-27b

[−]gaigalas · 2026-08-30 Sun 00:36 UTC · link
It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).
[−]walrus01 · 2026-08-30 Sun 00:39 UTC · link
Personally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about.

https://github.com/p-e-w/heretic

[−]gaigalas · 2026-08-30 Sun 00:47 UTC · link
I think those are mostly vapor that runs on the small culture of "models should not be censored" thing. But from my experience, they unlock nothing meaningful.

Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model.

That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).

[−]walrus01 · 2026-08-30 Sun 01:10 UTC · link
The most interesting use I've found for them so far is strictly as a novelty. Give a chat session with one to a completely non technical person, who at least knows that openai and anthropic have some guard rails on stuff, and tell them to wild with something like "give me the precursors and chemical formulas for the precusors for crystal meth" and watch it answer.
[−]gaigalas · 2026-08-30 Sun 04:58 UTC · link
Yep, but that's not changing the quality of the model. It's not an optimization in any sense (and it's a hit on productive workflows, possibly).

This is also likely to stop working as censoring moves to the training data source.

[−]dotancohen · 2026-08-30 Sun 07:26 UTC · link
But does it answer those queries correctly, or does it just not refuse to not halucinate an incorrect answer? From where would it even have that information?
[−]walrus01 · 2026-08-30 Sun 09:09 UTC · link
I don't know enough chemistry to say one way or the other if it's just wildly hallucinating the precursors and processes, but it'll also do things like, write an ISIS press release, or similar. There's a data set of basically a bunch of antisocial or dangerous prompts that some people have got variants of qwen to pass with 0 out of 465 refusals:

https://huggingface.co/datasets/mlabonne/harmful_behaviors

[−]fc417fc802 · 2026-08-30 Sun 00:41 UTC · link
Training a variant to reason in early modern english in the style of the tudor elites might be an amusing way to test for that.
[−]altmanaltman · 2026-08-30 Sun 03:26 UTC · link
Just so we are clear, no "caveman" spoke English. "Caveman speak" is just shortening the vocabulary of english, not a "caveman language". Given this, your concerns for "stereotypical caveman manner" makes very little sense since what caveman are you talking about?
[−]dotancohen · 2026-08-30 Sun 07:30 UTC · link
The concern is not that the model was trained on actual caveman artifacts, rather on modern media representations of the stereotypical caveman (that never actually existed).
[−]dotancohen · 2026-08-30 Sun 07:23 UTC · link
Caveman invented fire, the wheel, domesticated wild plants and animals, organised society, survived the Toba catastrophe, cooked food, and was having sex ages before you and me. Don't write him off as stupid.
[−]Gravityloss · 2026-08-30 Sun 12:57 UTC · link
And I wonder how they actually spoke. Since there was no visual communications medium except for cave art. (Some of which is very excellent. Try drawing 3d curved horns in perspective.) So people would have used verbal communication more. Also no written word. So one would expect there to be quite a lot of oral tradition. Like people reciting poem form epics.

If we assume the time is before farming, population density would have been low and limiting culture. Hunter-gatherers might have travelled a lot more than farmers with a homestead though.

[−]miroljub · 2026-08-30 Sun 16:00 UTC · link
They were smarter and more fit than us. At that time not being able or not wanting to contribute to the group meant your genes were dropped from the evolution pool forever.

Unlike today where a small group of tax payer is keeping alive and thriving complete parasitic parts of human races.

Until that changes we are doomed to regress and degenerate back to monkey like creatures.

Then, there would be no discussion whether "caveman speech" is suitable for talking to the AI.

[−]ekianjo · 2026-08-29 Sat 23:02 UTC · link
Saving tokens
[−]andsoitis · 2026-08-29 Sat 23:59 UTC · link
More intelligent and shorter:

Maybe add a small cycling cap or helmet if it doesn’t obscure the head.

[−]walrus01 · 2026-08-30 Sun 00:31 UTC · link
qwen3.8-flash-next also 'thinks' like this in its thinking stage before output, watching it 'think' in opencode, but it produces syntax correct and grammatically correct code comments, changelogs and readme type files.
[−]AdamConwayIE · 2026-08-30 Sun 01:56 UTC · link
Likely something that was first made especially obvious by Chinese models and then became something worth optimizing for in English too.

Chinese can be extremely information-dense in token terms, though it depends on the tokenizer. Roughly speaking, you can pack more "meaning" into a short sequence than English often allows for. That's why "caveman" reasoning is a pretty good fit.

There's a difference between bolting caveman speak onto an existing model and training a model to reason that way, though. If you just force an existing model to be concise in outputs, you're artificially reducing its available reasoning steps and can possibly prevent useful exploration or verification. If it's trained specifically to use compressed reasoning, it can learn to represent the same intermediate ideas in fewer generated tokens, cutting the number of sequential inference steps without necessarily sacrificing the useful reasoning itself.

It's not so much inherently a Chinese-model trait, but Chinese models could definitely have helped demonstrate how effective very compressed reasoning traces can be.

There are few tests of this, but one example I thought was interesting was here: https://github.com/PastaPastaPasta/llm-chinese-english

I wouldn't say it was Chinese specifically that was emulated, but it got people thinking about tokenizers and representation efficiency, and how natural English is rather inefficient.

[−]armcat · 2026-08-30 Sun 08:44 UTC · link
Less tokens. These models already overthink like crazy especially for complex tasks.
[−]gs17 · 2026-08-30 Sun 00:30 UTC · link
> Let's maybe add comments? The final code can have comments. Fine.
[−]tyre · 2026-08-30 Sun 00:34 UTC · link
This is actually pretty good!
[−]demibabs · 2026-08-30 Sun 00:35 UTC · link
No one’s talking about how good the final product is.

Edit: someone else commented that as I was typing this, lol.

[−]blackhaz · 2026-08-30 Sun 07:29 UTC · link
I wonder, do we need a new benchmark? There's quite a bit of feedback data floating around about pelicans on bicycles already.
[−]ziofill · 2026-08-30 Sun 09:05 UTC · link
That’s a fair question, but it seems that it’s not yet necessary. See here

https://dylancastillo.co/posts/pelicanmaxxing.html

https://simonwillison.net/2026/Jul/22/

[−]stymaar · 2026-08-30 Sun 15:50 UTC · link
I don't think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for “text-to-svg simple image generation” rather than just “pelicans on bicycle”.
[−]delichon · 2026-08-30 Sun 00:59 UTC · link
If someone can look at that reasoning trace and see a stochastic parrot next word prediction machine, we don't understand those words in the same way.
[−]sneak · 2026-08-30 Sun 01:09 UTC · link
Yeah, it has been clear for a long time that there is reasoning and mental modeling going on here.

The other option is that you do understand those words the same way, and the people making these (now nonsensical) anti-AI claims simply aren’t talking about the same programs/models we are. Their idea of SOTA is when chatgpt.com launched.

If you took a point sample pre-Opus, and didn’t write a good prompt, of course you would think all AI programming was worthless slop.

[−]negura · 2026-08-30 Sun 06:46 UTC · link
Maybe get yourself checked for chatbot psychosis. I am using current models productively, every day, and have 0 (and I mean precisely, literally 0) issue with calling it a stochastic parrot, one which lacks any kind of mentality whatsoever. There is not a shred of doubt in my mind that this is purely a statistical model, generating sequences of words, that happen to make sense in our actual mentality.
[−]weego · 2026-08-30 Sun 11:05 UTC · link
it has been clear for a long time that there is reasoning and mental modeling going on here

There is not. No one from these products is even claiming that's the case and they're so desperate to make the next big claim to re-ignite investment they'd be shouting it from every rooftop.

It's just breaking out all the reasonable probabilities around what it's been tasked with and structuring them in a way that is designed to actively look human, and then feed it back to itself. Fundamentally that's the easiest way to iterate new features when the underlying architecture of LLMs is largely "fixed" right now. The fact it is output in a way that appears to reason through each is just a technical decision that creates an illusion of reasoning.

[−]gjm11 · 2026-08-30 Sun 14:49 UTC · link
What, as precisely as you can say, is the difference between an illusion of reasoning and reasoning?

(I am not claiming that there is none. But I personally would define "reasoning" in terms of its structure and its results, and it looks to me as if the best LLMs' "illusion of reasoning" has enough similarities in structure and results to much human reasoning that I don't see why we shouldn't also call it reasoning; if your opinion differs then I'm curious about where the disagreements lie. E.g., do we have different beliefs about what sort of thing LLMs' schmeasoning is able to accomplish, or does your notion of "reasoning" specifically require that it be done by humans, or what?)

[−]dnautics · 2026-08-30 Sun 01:54 UTC · link
The stochastic parrot epithet is so 4 months ago
[−]0xfaded · 2026-08-30 Sun 03:27 UTC · link
I still call them stochastic parrots, but believe what they are revealing is that we are all stochastic parrots to some extent. I simply don't see how biological computation (i.e. thinking) can be anything else. Similar to the reveal in west world, we are likely much simpler than we give ourselves credit for.

A "train of thought" can be seen as a trace of a depth first search where the preceding trace is used to guide termination and next expansion decisions. A similar concept, "taboo search", exists in classical constraint optimization where previous solutions are fit to a model that guides future expansion (but as the name "taboo" implies, away from uninteresting solutions).

We also have harnesses that perform breath first search.

If I tried to describe what it means to "think deeply", I would probably say a combination of both.

Ultimately I believe that we will surpass human capabilities but fail with alignment. Handing the world's resources over to stochastic systems that can evolve faster than we can reason about them simply leaves too many "interesting" outcomes that do not end well. I also expect the failure modes will be totally non-obvious.

[−]vasco · 2026-08-30 Sun 06:08 UTC · link
As long as there's enough of them with different goals it doesn't matter, they'll keep each other in check. The worlds resources are already handed over to the worst people and we're still doing fine and none of the billionaires are "aligned with society". They just align with their own belly but because they want different things it all kinda works.
[−]sujzhsbnwjek · 2026-08-30 Sun 06:34 UTC · link
These “worst people” need you. They physically need you alive to perform labor for them and to give them money (and status).

That’s the reason we are “doing fine”. Once they stop needing you..

Also, both our comments brush over the generational struggles for fairness over the centuries. We have fought to be “fine”, it did not just happen. Without fairness being introduced by force you and I would be slaving away in some sweatshop getting paid nickels as was the norm not so long ago.

Edit: That’s also assuming you are Caucasian. If you are of a different ethnicity.. well, historically, all bets are off. You could also be the literal possession of some of these “worst people” with not even your own children considered yours.

[−]vasco · 2026-08-30 Sun 07:18 UTC · link
My only point is in a many agent system with different goals it "doesn't matter" that some agents have bad goals as long as there's enough variability of goals and resources that they can't put their vision in place.
[−]zhahhshsj · 2026-08-30 Sun 15:50 UTC · link
I know but that is just .. not how it works. Humans don’t align on just about anything but they will do tremendous mindboggling amounts of harm if not checked by mountains of checks and balances. Sheer variety alone is not a guarantee of anything.

Many types of Hitler does not make for a peaceful world all of sudden through sheer competition. It sounds nice but it will lead to certain hell.

[−]BoredomIsFun · 2026-08-30 Sun 06:59 UTC · link
> we are all stochastic parrots to some extent.

I think this statement is continuation of the old fallacy - every generation thinks of brain in terms of what is the current technology zaitgeist is - was it 19th century when they thought brain is a network of pneumatic pipes?

[−]pasteleft · 2026-08-30 Sun 04:49 UTC · link
LLM is "stochastic parrot next word prediction machine"; it's just that this "stochastic parrot next word prediction machine" have proven to be smarter than most people. I mean, this already happened with AlphaGo too.
[−]BoredomIsFun · 2026-08-30 Sun 07:02 UTC · link
> to be smarter than most people.

Hell no. In very narrow tasks - yes, in vast majority, esp. involving state tracking (board games) and spatial reasoning - they are awful.

[−]slopinthebag · 2026-08-30 Sun 08:11 UTC · link
Why can't a next token prediction machine not predict a train of reasoning?
[−]moezd · 2026-08-30 Sun 05:30 UTC · link
Did he seriously automate away one of the best quirks of his blog posts, i.e. evaluating new models with a touch of fun? I read AI slop all day, thanks.
[−]KeplerBoy · 2026-08-30 Sun 08:16 UTC · link
That's just part of the output along the SVG/image.
[−]chvid · 2026-08-30 Sun 06:18 UTC · link
This is a remarkable coherent and clear reasoning trace.

Maybe you should start also comparing reasoning traces when you do your pelican benchmark.

[−]armcat · 2026-08-30 Sun 06:34 UTC · link
That would be very interesting but only the open models allow you to see the reasoning trace
[−]kmike84 · 2026-08-30 Sun 11:27 UTC · link
So, open models will be better on this benchmark, which is deserved
[−]Aboutplants · 2026-08-30 Sun 13:17 UTC · link
Halfway through it states “Let's mentally compose SVG.”

Is this a common thing? I’ve never seen it before, the “mentally compose”

[−]simonw · 2026-08-30 Sun 13:41 UTC · link
It's common for models to produce a draft of the SVG part way through their reasoning. Here's Qwen3.8-Flash-Next doing that, for example: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

The open weight models let you see the full reasoning trace. Models from OpenAI, Anthropic, and Gemini tend to obscure or summarize the reasoning traces so you can't see exactly what they're doing. Here's Gemini 3.7 Flash which looks like it's doing something similar: https://tools.simonwillison.net/markdown-svg-renderer.html#u...

One of the step summaries includes this:

> I'm now detailing the pelican's anatomy within the SVG. I've sketched the main body outline, including coordinates for the tail, chest, neck, head, and massive beak with a pouch. I'm focusing on the position of the eyes and considering the positioning of the wings, with the foreground wing on the handlebar for a confident look.