I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted.
At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality.
We see a very janky pelican and declare the problem solved.
jmugantoday at 5:56 PM
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
bredrentoday at 4:55 PM
I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
jcimstoday at 8:29 PM
I’d like to see a human one shot a pelican on a bicycle in raw svg.
try-workingtoday at 7:59 PM
this is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media.
a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
HarHarVeryFunnytoday at 5:26 PM
It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
qwertoxtoday at 4:41 PM
I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.
I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
informal007today at 8:32 PM
One difference for human to understand the video is that we only care the changes on a picture compare to LLM
swe_dimatoday at 8:26 PM
In my experience SVGs are still too hard for LLMs.
I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it.
Results looked like drawing of a 5 year old.
toolslivetoday at 6:13 PM
Reading the title, I was thinking "Karpathy? I don't know this chess player."
(The Pelikan is a well known chess opening, and famous chess players often have book titles like "X's Y" where X is the player, and Y is the opening)
trentortoday at 5:15 PM
I always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?
Waterluviantoday at 7:46 PM
Speaking of benchmarks has anyone given AIs Where’s Waldo pages and asked it to find Waldo?
I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.
hooloovoo_zootoday at 8:06 PM
I suspect LotR is a singularly unrepresentative choice here considering how much info exists about it.
baron816today at 4:57 PM
IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
siliconc0wtoday at 7:37 PM
There is a tipping point between procedurally generating everything in SVG to maybe giving them tool access to something like 3dsmax (or having them build and then use a tool to do the thing vs doing the thing).
fzeindltoday at 5:02 PM
Regarding the argument about LLMs having difficulties auditing their work:
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
xyzsparetimexyztoday at 5:33 PM
How much are the hobbit houses described in the book? The ones here look exactly like the movie
matsemanntoday at 5:19 PM
I'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.
informal007today at 8:29 PM
it shows the possibility that SVG replace PNG/JPG even video.
barrenkotoday at 6:24 PM
This has started to feel a bit like the beginning of railroads and then the steampunk fiction of "let's just build railroads to everywhere". We don't need it and there's no use for it.
As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies.
This software crap is done, leave it to the nerds.
eichintoday at 6:54 PM
Is anyone else getting "mongodb is webscale" vibes? (Except 16 years ago that was a lot smoother, because it used some sort of "render this conversation" engine...)
sinaatalaytoday at 7:56 PM
On consumer devices, AI communicates with us through speakers and screens. Screens are the richer medium, so most consumer AI innovation will happen there.
Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.
cocoa19today at 5:25 PM
We must not be using the same opus 5, because if I tried to generate this it would refuse based on copyright grounds.
dekhntoday at 5:59 PM
I'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).
fwlrtoday at 5:59 PM
I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.)
It always brings to my mind some words from Rich Hickey:
I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!”
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.
mold_aidtoday at 7:05 PM
The tilde thing remains uniquely obnoxious in a field that seems want to mangle language for fun, so that's innovative I guess
skybriantoday at 5:05 PM
Still images seem like a better quick test because we can see them at a glance. Maybe ask it to make a comic?
hkalbasitoday at 5:44 PM
This makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?
wiradikusumatoday at 5:27 PM
Do you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)?
I wonder if Flash is still popular... LLM can use that instead...?
dofmtoday at 6:03 PM
Anthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money.
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.
serftoday at 5:02 PM
you don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
croestoday at 6:00 PM
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
OtherShrezzingtoday at 6:12 PM
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.
After watching this video I am absolutely positive that the issue is not a lack of stamina in humans. It is that humans have the capacity to realize that this is a bad idea long before they complete it.
throwaway89864today at 5:43 PM
It may make sense to switch this to USD/Omniverse.
xg15today at 5:41 PM
> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it.
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
I'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic.
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
bbstatstoday at 5:51 PM
This is awful
andy99today at 6:25 PM
Benchmarks like the pelican thing are about correlation with “how good the model is”. Better models produce better pelicans.
It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance.
Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway.
TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.
blitzartoday at 5:09 PM
I think the pelican test is better.
stackedinsertertoday at 6:26 PM
It would be better to ask model to render segmented 3d, with placeholders, like magenta is water, blue is sky, green is grass, purple is Frodo's face, etc, then pass the result through img2img model to properly "render" it.
c0rruptbytestoday at 4:38 PM
perfect benchmark to burn more tokens - convenient
epolanskitoday at 4:45 PM
I wish there was a timeline where I never ever had to see the pelican SVG test ever again.
forrestthewoodstoday at 5:19 PM
As a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment.
There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.
In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.
My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.
Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
"Check the current situation and make a new Iran Lego (tm) truth bomb video."
miltonlosttoday at 5:35 PM
Tech bros continue wasting money to make the absolute worst art
hansmayertoday at 8:18 PM
[dead]
trlhaqtoday at 5:00 PM
[flagged]
theproblemisyoutoday at 5:31 PM
[flagged]
yourewrongsorrytoday at 5:15 PM
[flagged]
andrewstuarttoday at 7:03 PM
This is equally bad as a pelican test.
LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.
So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).
You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.
Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.
No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.
Any test that anyone comes up with must clearly state the context and how the outcome is measured.
hn22fazjsvtoday at 6:08 PM
Screenshotting for later
matchagauchotoday at 4:51 PM
It's difficult to think in exponentials.
But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.