Hacker News

Favorites Setup
Comment by kurante | original | Hy4 preview
[−]kurante · 2026-08-29 Sat 22:45 UTC · link
Is the broken English an optimization or a byproduct of the model being developed in China?
[−]minimaxir · 2026-08-29 Sat 22:48 UTC · link
Optimization. Why use many word when few word do trick?
[−]stavros · 2026-08-29 Sat 23:19 UTC · link
What I find funny about "why use many word when few word do trick?" is that it's only slightly shorter than the regular "why use many words when few words do the trick?"
[−]rapind · 2026-08-29 Sat 23:48 UTC · link
I always figured that was part of the joke, because a writer came up with it, and a writer would know (I assume?).
[−]inopinatus · 2026-08-30 Sun 07:54 UTC · link
The latter is not a complete alternative, it is ambiguously conflating vocabulary scale with word count, and also, it is not as funny
[−]happycube · 2026-08-30 Sun 10:22 UTC · link
less word better(, many words bad)
[−]andsoitis · 2026-08-29 Sat 23:33 UTC · link
> Why use many word when few word do trick?

Be concise.

  OR
Brief is best.

  OR
Eschew verbosity

  etc.
[−]gjvc · 2026-08-30 Sun 01:00 UTC · link
"Omit needless words."

-- William Strunk Jr. and E.B. White., The Elements of Style

[−]gaigalas · 2026-08-30 Sun 00:43 UTC · link
Optimization on a idiosyncrasy. The same thing that makes Claude repeat "That was the most important thing you said in this whole conversation" is what makes grug speak optimize on token usage.

Real humans get non-primary information from word variation. It's reasonable to hypothesize that it has a role in thinking things, because it endures. Our languages need to breathe over time, and flourishing might be one of the aspects that allows that breathing space.

[−]TiredOfLife · 2026-08-30 Sun 08:03 UTC · link
See world
[−]acheong08 · 2026-08-29 Sat 22:49 UTC · link
When GPT-5.6-sol's reasoning traces were leaked, they also used "caveman speak". Definitely a token efficiency optimization
[−]beefsack · 2026-08-29 Sat 23:31 UTC · link
I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?
[−]Barbing · 2026-08-29 Sat 23:37 UTC · link
"Neuralese"
[−]walrus01 · 2026-08-30 Sun 00:31 UTC · link
some people made a 'caveman' speak qwen as a joke

https://huggingface.co/ProCreations/grug-27b

[−]gaigalas · 2026-08-30 Sun 00:36 UTC · link
It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).
[−]walrus01 · 2026-08-30 Sun 00:39 UTC · link
Personally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about.

https://github.com/p-e-w/heretic

[−]gaigalas · 2026-08-30 Sun 00:47 UTC · link
I think those are mostly vapor that runs on the small culture of "models should not be censored" thing. But from my experience, they unlock nothing meaningful.

Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model.

That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).

[−]walrus01 · 2026-08-30 Sun 01:10 UTC · link
The most interesting use I've found for them so far is strictly as a novelty. Give a chat session with one to a completely non technical person, who at least knows that openai and anthropic have some guard rails on stuff, and tell them to wild with something like "give me the precursors and chemical formulas for the precusors for crystal meth" and watch it answer.
[−]gaigalas · 2026-08-30 Sun 04:58 UTC · link
Yep, but that's not changing the quality of the model. It's not an optimization in any sense (and it's a hit on productive workflows, possibly).

This is also likely to stop working as censoring moves to the training data source.

[−]dotancohen · 2026-08-30 Sun 07:26 UTC · link
But does it answer those queries correctly, or does it just not refuse to not halucinate an incorrect answer? From where would it even have that information?
[−]walrus01 · 2026-08-30 Sun 09:09 UTC · link
I don't know enough chemistry to say one way or the other if it's just wildly hallucinating the precursors and processes, but it'll also do things like, write an ISIS press release, or similar. There's a data set of basically a bunch of antisocial or dangerous prompts that some people have got variants of qwen to pass with 0 out of 465 refusals:

https://huggingface.co/datasets/mlabonne/harmful_behaviors

[−]fc417fc802 · 2026-08-30 Sun 00:41 UTC · link
Training a variant to reason in early modern english in the style of the tudor elites might be an amusing way to test for that.
[−]altmanaltman · 2026-08-30 Sun 03:26 UTC · link
Just so we are clear, no "caveman" spoke English. "Caveman speak" is just shortening the vocabulary of english, not a "caveman language". Given this, your concerns for "stereotypical caveman manner" makes very little sense since what caveman are you talking about?
[−]dotancohen · 2026-08-30 Sun 07:30 UTC · link
The concern is not that the model was trained on actual caveman artifacts, rather on modern media representations of the stereotypical caveman (that never actually existed).
[−]dotancohen · 2026-08-30 Sun 07:23 UTC · link
Caveman invented fire, the wheel, domesticated wild plants and animals, organised society, survived the Toba catastrophe, cooked food, and was having sex ages before you and me. Don't write him off as stupid.
[−]Gravityloss · 2026-08-30 Sun 12:57 UTC · link
And I wonder how they actually spoke. Since there was no visual communications medium except for cave art. (Some of which is very excellent. Try drawing 3d curved horns in perspective.) So people would have used verbal communication more. Also no written word. So one would expect there to be quite a lot of oral tradition. Like people reciting poem form epics.

If we assume the time is before farming, population density would have been low and limiting culture. Hunter-gatherers might have travelled a lot more than farmers with a homestead though.

[−]miroljub · 2026-08-30 Sun 16:00 UTC · link
They were smarter and more fit than us. At that time not being able or not wanting to contribute to the group meant your genes were dropped from the evolution pool forever.

Unlike today where a small group of tax payer is keeping alive and thriving complete parasitic parts of human races.

Until that changes we are doomed to regress and degenerate back to monkey like creatures.

Then, there would be no discussion whether "caveman speech" is suitable for talking to the AI.

[−]ekianjo · 2026-08-29 Sat 23:02 UTC · link
Saving tokens
[−]andsoitis · 2026-08-29 Sat 23:59 UTC · link
More intelligent and shorter:

Maybe add a small cycling cap or helmet if it doesn’t obscure the head.

[−]walrus01 · 2026-08-30 Sun 00:31 UTC · link
qwen3.8-flash-next also 'thinks' like this in its thinking stage before output, watching it 'think' in opencode, but it produces syntax correct and grammatically correct code comments, changelogs and readme type files.
[−]AdamConwayIE · 2026-08-30 Sun 01:56 UTC · link
Likely something that was first made especially obvious by Chinese models and then became something worth optimizing for in English too.

Chinese can be extremely information-dense in token terms, though it depends on the tokenizer. Roughly speaking, you can pack more "meaning" into a short sequence than English often allows for. That's why "caveman" reasoning is a pretty good fit.

There's a difference between bolting caveman speak onto an existing model and training a model to reason that way, though. If you just force an existing model to be concise in outputs, you're artificially reducing its available reasoning steps and can possibly prevent useful exploration or verification. If it's trained specifically to use compressed reasoning, it can learn to represent the same intermediate ideas in fewer generated tokens, cutting the number of sequential inference steps without necessarily sacrificing the useful reasoning itself.

It's not so much inherently a Chinese-model trait, but Chinese models could definitely have helped demonstrate how effective very compressed reasoning traces can be.

There are few tests of this, but one example I thought was interesting was here: https://github.com/PastaPastaPasta/llm-chinese-english

I wouldn't say it was Chinese specifically that was emulated, but it got people thinking about tokenizers and representation efficiency, and how natural English is rather inefficient.

[−]armcat · 2026-08-30 Sun 08:44 UTC · link
Less tokens. These models already overthink like crazy especially for complex tasks.