> [...] Let's maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn't ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no.
I don't think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for “text-to-svg simple image generation” rather than just “pelicans on bicycle”.
That is the point though, if labs are maximizing svg image generation capabilities it is a very good thing. That's a general skill that is useful. So assuming they aren't specifically maximizing pelican bicycle svgs (and it doesn't look like they are) then incentives are aligned that the "benchmark" is measuring a general desirable capability.
What I find funny about "why use many word when few word do trick?" is that it's only slightly shorter than the regular "why use many words when few words do the trick?"
Optimization on a idiosyncrasy. The same thing that makes Claude repeat "That was the most important thing you said in this whole conversation" is what makes grug speak optimize on token usage.
Real humans get non-primary information from word variation. It's reasonable to hypothesize that it has a role in thinking things, because it endures. Our languages need to breathe over time, and flourishing might be one of the aspects that allows that breathing space.
I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?
It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).
Personally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about.
I think those are mostly vapor that runs on the small culture of "models should not be censored" thing. But from my experience, they unlock nothing meaningful.
Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model.
That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).
The most interesting use I've found for them so far is strictly as a novelty. Give a chat session with one to a completely non technical person, who at least knows that openai and anthropic have some guard rails on stuff, and tell them to wild with something like "give me the precursors and chemical formulas for the precusors for crystal meth" and watch it answer.
But does it answer those queries correctly, or does it just not refuse to not halucinate an incorrect answer? From where would it even have that information?
I don't know enough chemistry to say one way or the other if it's just wildly hallucinating the precursors and processes, but it'll also do things like, write an ISIS press release, or similar. There's a data set of basically a bunch of antisocial or dangerous prompts that some people have got variants of qwen to pass with 0 out of 465 refusals:
Just so we are clear, no "caveman" spoke English. "Caveman speak" is just shortening the vocabulary of english, not a "caveman language". Given this, your concerns for "stereotypical caveman manner" makes very little sense since what caveman are you talking about?
The concern is not that the model was trained on actual caveman artifacts, rather on modern media representations of the stereotypical caveman (that never actually existed).
Caveman invented fire, the wheel, domesticated wild plants and animals, organised society, survived the Toba catastrophe, cooked food, and was having sex ages before you and me. Don't write him off as stupid.
And I wonder how they actually spoke. Since there was no visual communications medium except for cave art. (Some of which is very excellent. Try drawing 3d curved horns in perspective.) So people would have used verbal communication more. Also no written word. So one would expect there to be quite a lot of oral tradition. Like people reciting poem form epics.
If we assume the time is before farming, population density would have been low and limiting culture. Hunter-gatherers might have travelled a lot more than farmers with a homestead though.
The Pleiades cluster is called the seven sisters in Greek. That's curious, because the human eye under the best conditions can discern only six stars in there. Even more curious, the aboriginal Australians also called this cluster the seven sisters.
Ancient Greeks' and aboriginal Australians' last common ancestors split about 60,000 years ago. And astronomers tell us that 60,000 years ago, there were seven discernable stars in that cluster.
One could this conclude not only is speech likely 60,000 years old, but also that the tale of the seven sisters might be a tale from so long ago.
Interestingly, back in my ill-spent youth, a few of my fellow astronomers and I were out in a very, very dark-sky location in the late 70s/early 80s, and we were able to consistently count and draw between 9 and 11 stars. Although we would tease those who could see 11 stars as using averted imagination. :-) Today, if I can see six stars, it's an okay night in an okay sky.
fwiw, if you can get out to dark skies where you can see fifth- or sixth-magnitude stars with the naked eye, I highly recommend getting out there when it's a low-moisture atmosphere and the Milky Way through Cassiopeia and Perseus is vertical, as it's a rather dramatic sight of this stream of stars heading down to the northern horizon.
The summertime Milky Way overhead down to Sagittarius tends to get all the love, but the wintertime Milky Way is also visually rich and worth spending time on.
They were smarter and more fit than us. At that time not being able or not wanting to contribute to the group meant your genes were dropped from the evolution pool forever.
Unlike today where a small group of tax payer is keeping alive and thriving complete parasitic parts of human races.
Until that changes we are doomed to regress and degenerate back to monkey like creatures.
Then, there would be no discussion whether "caveman speech" is suitable for talking to the AI.
qwen3.8-flash-next also 'thinks' like this in its thinking stage before output, watching it 'think' in opencode, but it produces syntax correct and grammatically correct code comments, changelogs and readme type files.
Likely something that was first made especially obvious by Chinese models and then became something worth optimizing for in English too.
Chinese can be extremely information-dense in token terms, though it depends on the tokenizer. Roughly speaking, you can pack more "meaning" into a short sequence than English often allows for. That's why "caveman" reasoning is a pretty good fit.
There's a difference between bolting caveman speak onto an existing model and training a model to reason that way, though. If you just force an existing model to be concise in outputs, you're artificially reducing its available reasoning steps and can possibly prevent useful exploration or verification. If it's trained specifically to use compressed reasoning, it can learn to represent the same intermediate ideas in fewer generated tokens, cutting the number of sequential inference steps without necessarily sacrificing the useful reasoning itself.
It's not so much inherently a Chinese-model trait, but Chinese models could definitely have helped demonstrate how effective very compressed reasoning traces can be.
I wouldn't say it was Chinese specifically that was emulated, but it got people thinking about tokenizers and representation efficiency, and how natural English is rather inefficient.
The open weight models let you see the full reasoning trace. Models from OpenAI, Anthropic, and Gemini tend to obscure or summarize the reasoning traces so you can't see exactly what they're doing. Here's Gemini 3.7 Flash which looks like it's doing something similar: https://tools.simonwillison.net/markdown-svg-renderer.html#u...
One of the step summaries includes this:
> I'm now detailing the pelican's anatomy within the SVG. I've sketched the main body outline, including coordinates for the tail, chest, neck, head, and massive beak with a pouch. I'm focusing on the position of the eyes and considering the positioning of the wings, with the foreground wing on the handlebar for a confident look.
Did he seriously automate away one of the best quirks of his blog posts, i.e. evaluating new models with a touch of fun? I read AI slop all day, thanks.
Yeah, it has been clear for a long time that there is reasoning and mental modeling going on here.
The other option is that you do understand those words the same way, and the people making these (now nonsensical) anti-AI claims simply aren’t talking about the same programs/models we are. Their idea of SOTA is when chatgpt.com launched.
If you took a point sample pre-Opus, and didn’t write a good prompt, of course you would think all AI programming was worthless slop.
Maybe get yourself checked for chatbot psychosis. I am using current models productively, every day, and have 0 (and I mean precisely, literally 0) issue with calling it a stochastic parrot, one which lacks any kind of mentality whatsoever. There is not a shred of doubt in my mind that this is purely a statistical model, generating sequences of words, that happen to make sense in our actual mentality.
it has been clear for a long time that there is reasoning and mental modeling going on here
There is not. No one from these products is even claiming that's the case and they're so desperate to make the next big claim to re-ignite investment they'd be shouting it from every rooftop.
It's just breaking out all the reasonable probabilities around what it's been tasked with and structuring them in a way that is designed to actively look human, and then feed it back to itself. Fundamentally that's the easiest way to iterate new features when the underlying architecture of LLMs is largely "fixed" right now. The fact it is output in a way that appears to reason through each is just a technical decision that creates an illusion of reasoning.
What, as precisely as you can say, is the difference between an illusion of reasoning and reasoning?
(I am not claiming that there is none. But I personally would define "reasoning" in terms of its structure and its results, and it looks to me as if the best LLMs' "illusion of reasoning" has enough similarities in structure and results to much human reasoning that I don't see why we shouldn't also call it reasoning; if your opinion differs then I'm curious about where the disagreements lie. E.g., do we have different beliefs about what sort of thing LLMs' schmeasoning is able to accomplish, or does your notion of "reasoning" specifically require that it be done by humans, or what?)
I still call them stochastic parrots, but believe what they are revealing is that we are all stochastic parrots to some extent. I simply don't see how biological computation (i.e. thinking) can be anything else. Similar to the reveal in west world, we are likely much simpler than we give ourselves credit for.
A "train of thought" can be seen as a trace of a depth first search where the preceding trace is used to guide termination and next expansion decisions. A similar concept, "taboo search", exists in classical constraint optimization where previous solutions are fit to a model that guides future expansion (but as the name "taboo" implies, away from uninteresting solutions).
We also have harnesses that perform breath first search.
If I tried to describe what it means to "think deeply", I would probably say a combination of both.
Ultimately I believe that we will surpass human capabilities but fail with alignment. Handing the world's resources over to stochastic systems that can evolve faster than we can reason about them simply leaves too many "interesting" outcomes that do not end well. I also expect the failure modes will be totally non-obvious.
As long as there's enough of them with different goals it doesn't matter, they'll keep each other in check. The worlds resources are already handed over to the worst people and we're still doing fine and none of the billionaires are "aligned with society". They just align with their own belly but because they want different things it all kinda works.
These “worst people” need you. They physically need you alive to perform labor for them and to give them money (and status).
That’s the reason we are “doing fine”. Once they stop needing you..
Also, both our comments brush over the generational struggles for fairness over the centuries. We have fought to be “fine”, it did not just happen. Without fairness being introduced by force you and I would be slaving away in some sweatshop getting paid nickels as was the norm not so long ago.
Edit: That’s also assuming you are Caucasian. If you are of a different ethnicity.. well, historically, all bets are off. You could also be the literal possession of some of these “worst people” with not even your own children considered yours.
My only point is in a many agent system with different goals it "doesn't matter" that some agents have bad goals as long as there's enough variability of goals and resources that they can't put their vision in place.
I know but that is just .. not how it works. Humans don’t align on just about anything but they will do tremendous mindboggling amounts of harm if not checked by mountains of checks and balances. Sheer variety alone is not a guarantee of anything.
Many types of Hitler does not make for a peaceful world all of sudden through sheer competition. It sounds nice but it will lead to certain hell.
I think this statement is continuation of the old fallacy - every generation thinks of brain in terms of what is the current technology zaitgeist is - was it 19th century when they thought brain is a network of pneumatic pipes?
LLM is "stochastic parrot next word prediction machine"; it's just that this "stochastic parrot next word prediction machine" have proven to be smarter than most people. I mean, this already happened with AlphaGo too.
> Maybe add sunglasses? no.
> Maybe add water? no.
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...