Hacker News (curated)new | past | comments | ask | show | jobs| show hidden

Is anyone here working on a problem for which current generation LLMs are inadequate, but that could possibly be solved by the next release of a first tier LLM?

Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.



My experience is that even Opus 5 still tends to write buggy or low-quality code and makes serious mistakes when analyzing code. It's a lot better than before but still not something I trust. I've had less experience with Fable since I can't use it at work; I hear it's a step up but still has its limits.

For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capable of producing codebases that actually work. (Anthropic built a C compiler with Opus 4.6 but it lacked optimizations and apparently hit a complexity wall.)

I also want to use LLMs for reverse engineering, but apparently it's pretty hit-or-miss, especially if you're forced to use open-source models to avoid restrictions.


This reply is particularly interesting to me because most of my experience with actually using LLMs to get work done is with coding agents. But I only have a fairly narrow set of experiences: two pretty large solo Flutter projects. I am currently really pleased with Gemini as a coding agent. It could improve, but I think improvements are going to come from marginal gains in the harness and training material so it can catch things like misconfigured permissions in platform specific areas.

It's also interesting because, while coding agents are important and are a notable success, they are never going to be a multi trillion dollar business. And are there any other domains where LLMs have such a large impact?


Opus 5 is weird. It scores high on benchmarks, but it seems that majority of those who try to use it day to day hate it

One of the things that came out of the decoded reasoning paper was that Claude models had memorized answers to tests but hid this memorization from the user output and pretended to derive the answer properly. It's only possible to cheat so blatantly in closed models where the reasoning is hidden.

If it were a Chinese model everyone would be screaming benchmaxxed.

Seriously something feels really off about Opus 5. I hope they correct it before 4.6 is removed.


I asked a current generation LLM to make me $1k a week and it hasn't so far.

If you're smart enough to try this you're worth more than $52,000 a year. Think about how foolish everyone else will feel when they didn't test the new release of Totally Working Golden Goose For Real This Time

For me personally, the answer is no. Fable is adequate to do basically anything I want to do. My perspective, broadly speaking, is that we've saturated most of the benchmarks because we've largely saturated our capacity to verify models' work at scale. What's left is context-bound verification, i.e. the problem of ensuring that output matches intent and ambiguities in prompting were resolved correctly. Further advances in autonomy do not make that latter verification problem easier. If anything they make it harder as the output per task becomes more complex and therefore more taxing for a human to verify.

The solution to that (to my mind) would be not a better model but a basic shift in architecture beyond the current paradigm and into a setup where agents have durable, plastic memories and undergo contextual individuation over time. But at that point agents start to become quasi-persons and not tools.


I saw a laptop earlier in the train that I asked ChatGPT, Claude and Gemini what it was, providing a brand, screen size and ports description. Gemini could never figure it out, Claude and ChatGPT eventually did, after multiple rounds of indirection, giving completely wrong answers (there was a perfect match for the problem statement, they all explored alternatives first). LLMs are (probably) amazing at things I don't care about, and still suck at the mundane stuff you would have the marketing tell you they excel at.

I want to be able to generate my own Simlilirian movie by dumping the content of a book into an LLM.

Both animated and live action results would be acceptable.

Unfortunately most existing LLMs lack the capability to maintain context across tens of thousands of frames.


That sounds like an interesting challenge. Have you seriously considered solving it? Because in about 10 seconds I came up with a process that should work, provided enough compute power. Simply model the traditional film making process by starting with a script, character stories. Design your world, then design the storyboard, and all the scenes. Create a list of all the visual elements that need to be replicated between all the scenes. Then you have to built prompts and reference art of the objects, faces, people. Make sure to do multiple takes of each scene, and have the vLLM critique and analyze the performances and technicalities. Should work?

I think, also, like in the traditional film makers career, this process should be built iteratively, start with a fast food commercial, then do a music video, then you can probably do a short film. Continue to improve the process, and one day I’m sure the LLM film studio can make you any movie you want, provided you have enough tokens.


I'm getting a Poe's Law feeling. I'm genuinely unsure about whether this post is a stone cold parody or not. I think I need to turn off the internet and go to bed.

EDIT: Your username doesn't help, either.


On that topic, check higgsfield cinema studio 4; they already provide amazing tech for the cinematic experience, somewhat similar to what you are describing.

Nobody wants to watch such films, they want to muck with the filmmaker, the generation apparatus. That's the product, if there is one here, and absolutely not the 3 hour epic that's spit out with 4 variations to choose between. That's work. We'll have other LLMs pointlessly tell us which should be watched, we'll view a summary, and vote the Oscar on that.

the problem is you have to make the AI watch the whole thing to make sure it works.

I've done this sort of with comfyui/same agent factory stuff, but the verification loop only works for models like fable as planner/writer, with gemini as verifier for like a very short movie. Sub 3-5 mins. After that you burn through million tokens.

Can't go too low fidelity audio/video or it craps out. Too long video and it loses consistency. Look at only snippets, it lacks global consistency, etc.


> Have you seriously considered solving it?

Sort of, but I want it to be relatively low on human effort. I feel burned by spending lots of time in 2023 learning image generation pipelines (using control net etc) only for that to be rendered trivial by the next generation of LLMs.

This movie would be only for personal consumption and I’m okay with waiting for model improvements.


Control nets are still useful in Krea and H3. For images, Krea can usually get close enough to a reference that it's not a big deal, but for H3 conditioning makes a big difference over prompting for complex actions.

How would a computer generated video be live action?

The results are boring. Not because the content is boring, but because you can so easily remix the results. Human curation is what creates value with these, not dumping and consuming. A personal perspective of a human being ups the respect, where the exact same sentences generated by an LLM carry no such value.

Tolkien did the human creative work - I just want a movie adaptation that’s as honest to the original text as possible. Think translating the text into video.

They are different mediums. While the underlying story might be Tolkien's, every frame is an artistic choice and while LLMs can make a choice is many situations, they are unable to 1) keep it coherent b) make it meaningful because art, to me and most, in an outcome of human experiences and thought, which by definition an llm cannot do.

Do you really want a 20 minute pause in the action, every time a new room or space is encountered, because Tolkien goes on for 2-5 pages describing every new environment like that.

I think this is the best and most realistic reply so far: the ability to do this is close enough, and things like AI music are hints that there is a business model for this. Maybe I'm just jaded about CGI effects in movies currently, but I think the fact that people except that kind of thing as entertainment means you might get away with a fully AI movie that people will pay for.

There are two more points in favor of this kind of AI movie project: there's zero chance that anyone would greenlight a Hollywood budget for the Silmarillion, and it is beyond human capability to write that screenplay.


Since live action results are acceptable, this is already possible with current day LLMs. Just instruct one to hire a writer, director, cast, and crew to make the movie.

Plus, the token costs involved should be pretty low! (Other costs may not be.)


I was given a picture cube, which is like a Rubik's cube but every side is a unique picture. It came scrambled and I don't have an original reference image. I like to take videos of it and give it to llms to solve. I call it my agi test because it hasn't been solved yet

Interesting problem! I don’t think I could solve it myself honestly.

Scientific physics simulations - even the frontier models just engage in rationalization of obviously unphysical results instead of understanding the system. They have the rote knowledge but fail to apply it unless their hand is held through the process.

Today's models can just write code to run the simulation instead

They choose not to unless you beg them, and of course they double down on not running it once they suggest reasoning qbout it is sufficient.

I am referring to doing exactly that.

Realistic, scientifically useful simulations still require tuning all sorts of parameters based on physical intuition and understanding of the system being simulated. Both Sol and Fable/Opus 5 fail at it and either blow up the computation cost to levels that can't be processed realistically, or they invent a justification for a visibly unphysical result.


I think it's more like steamshovels. 6 months ago they were enabling people who had never broken real earth to find out why osha has so many rules about shoring up walls. You could pull off something complex or delicate but it took a procedure with too many steps too much time to get there, good outcomes were pleasant surprises and required careful target selection. Now it's more like paying good money for a professional. They show up, measure twice, cut once, you're walking around your shiny new hole wondering what took the last guy so long.

The frayed edges on what I have slopped together as unreasonably ambitious, ludicrous projects with fucktons of tokens from models 6-9 months ago mostly look like situations where a capable-enough-to-be-dangerous developer tries to muscle through problems that explode in width & depth but keep digging (so, a tier below stopping early to do more design, two below recognizing the need for more planning from the outset). The primitives are there, most major things work well enough, but the remaining functionality and performance is inaccessible. At a cost of multiples of >1/8th of a $200/mo subscription.

Right now I can put $10 into DSv4 Pro/Flash or Qwen 3.8 Max/Flash, hand it a project and all of its unfinished forks in a state I barely remember, tell it that I want the things these forks have been working towards, and 8 hours later it has ie an working, tested, benchmarked multicore car physics simulation fabric with all of the forks evaluated, the gains merged in, the remaining work documented. It only needed a few hundred more lines of code but Codex 5.5 was never going to see that.

9 months ago I was saying developers are not being ambitious enough with these things, that's only more true now. They lend themselves to digging far deeper than they should: 200kloc god files, dozens of forks. Let it happen, you don't need to read it, it's for them, later. The only time you make them clean up is when it has a severe adverse effect on how long builds/lints/tests/benchmarks take. Spend a whole week having it do nothing but dig up published papers in relevant fields with cutting edge techniques and translating them into feature specs. Pick whatever state of the art is and try to crush it, throw everything at it, leave it looping on vague but wildly ambitious goals. When it modularizes and refactors it all down you might 'do a breakthrough', or maybe it happens in an hour over christmas break when you're trying the next one.


Yes. Most of us are, still. The frontier is currently both at expanding ‘common sense’ / non-cheating outcomes for imprecisely specified software (that’s all software), and at expanding autonomy - ability to work longer unsupervised with success, oh, and also at expanding outside contextual reasoning about what’s being built, as in “hmm, that doesn’t look right or make sense, let me explore that.”

3d modeling to an STL a part compatible to a visible cable raceway still fails even if I let Claude Fable use me as a robot that measures with calipers.

I think the tech analogy for frontier models is going to be super computers.

Super computers keep getting better but most people don't need them for most things.


2000's supercomputer is today's (highest end) smartphone performance tho

Yeah. And if you compare the AI you can run on a high end smart phone today to AI of the 2000s, today's smart phone probably wins.

In 10 or 20 years maybe we'll all be running AI models that are currently considered "frontier" on smart phone type devices.


Hardware debugging and firmware details lead to thinking/testing loops on all but the frontier here.

Yes, lots – I think that folks will hopefully discover more of these as they scale up their ambition, now that LLMs make a lot of previously difficult things far easier.

This is the exact same type of comment I heard about computer hardware upgrades for three decades in a row.

“Very few people actually require a Pentium workstation, a 486 is perfectly adequate for the majority”

The logical fallacy is taking an extant distribution of “product capability” that is priced to fit what the market will bear and assuming the “next upgrade” simply tacks on a little bit more to the right hand rail of that curve.

No!

It shifts the entire curve!

Everything for everyone gets better and the top 1% of the most demanding users will continue to pay the same-ish premium.

“Nothing” will change.

Look at it this way: you can buy a $200 laptop for your kid or a $20,000 Mac with an M5 Ultra processor.

BOTH are vastly more powerful than either a $200 PC or a $20,000 “workstation” from 20+ years ago.

Look at: https://arena.ai/leaderboard/text?q=openai&utm_source=chatgp...

The “budget” 5.5 Instant model beats o1 and o3 which were “pro” models at the time of their release!


Intel didn't just surf some natural wave of demand for higher power personal computers. Intel found new needs for powerful PCs, especially in gaming, and they put a lot of marketing and industry relations dollars behind PC gaming.

In other words. PC users didn't figure out that they could buy super powerful PCs and play games on them, that was a carefully managed market transition.

What is going to do the same for LLMs?


> Intel found new needs for powerful PCs

It wasn't "Intel" that found new uses for PCs, it was everybody who found new uses for them. Billions of people and millions of companies found uses for "more computer power".

It was only the journalists with limited imaginations (and no industry experience) who struggled to come up with potential uses.

> carefully managed market transition.

You make it sound like a conspiracy! It wasn't. It was simple capitalist competition. If Intel hadn't improved their products, their competitors would have left them behind.

That very nearly happened ten years ago because Intel become stuck on the 14nm process and their products stagnated while Apple, ARM, and AMD lapped them repeatedly.

> What is going to do the same for LLMs?

Everybody.

Are you saying that unless you're "carefully managed" by some third-party, you could not find any use for "unlimited intelligence on tap"?


Infra as code and devops shit. Fable is there in general because things it doesn’t know I can point at documentation and have it do a reasonable job. Opus 5 sucks. If I don’t have fable quota, I drop to Opus 4.8 and hold its hand.

A spanish rock solved that problem for free.

¿Como?

He just had an accident on "Vuelta a España" : https://www.theguardian.com/sport/2026/aug/29/cycling-tadej-...

Oh I had not heard about this, thanks for the link.

Reading OP's analogy I was like "even him don't need this bike now..."

I feel bad for him as a human, but as a cycling fan I'm glad that we'll have an interesting WC in Canada


I get buy with very cheap models and actually using my brain, you don't need these SOTA models. China will definitely win this AI 'war'

If the models stay open, it seems like everybody but anthropic/openai wins. i literally can’t see a downside. We can post-train the models to know about tienanmen square.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact | github