Vote on which of Hacker News' challenges for AI have been met

59 points - today at 5:32 PM

Source

Comments

bmenrigh today at 6:23 PM
At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
Retr0id today at 6:46 PM
Heh, there's one of mine: https://stoppels.ch/goalposts/?c=39727943

"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."

The vote is currently 64% yes, 18% no.

Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...

Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).

deleted today at 10:50 PM
ben_w today at 6:05 PM
Very pleased one of my predictions was totally wrong: https://news.ycombinator.com/item?id=23252711

Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.

ErrantX today at 6:37 PM
What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.

And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.

But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.

That alone tells you a lot IMO

6thbit today at 7:19 PM
Not sure why this thread got flagged ?

Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.

mrweasel today at 6:44 PM
The Turing test is interesting, because I believe that the current LLMs are perfectly capable of parsing the it in many situations. On the other hand we also have people are sound like they aren't real.

Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.

delichon today at 6:23 PM
If for each mistaken prediction there was some mild accountability, like someone shows up and slaps you with a trout, it would improve the site. But it should be added to the terms of service first.
travisgriggs today at 6:28 PM
How was this assembled? From a meta point of view, how much AI was used to curate and highlite the goals; how much was used to assemble the site itself? Or deploy it?
eternal_braid today at 7:11 PM
A chess scoresheet sometimes contains mistakes but chess players can figure out in many cases what was meant by thinking of what moves make sense and considering the level of play so far. Popular AIs tools fail at that.
deleted today at 6:17 PM
simianwords today at 6:19 PM
I made a bet with a guy on HN that the market value of OpenAI + Anthropic would get to at least 2.5T by 2027. I think I'm on track to winning.

https://news.ycombinator.com/item?id=48517353

I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropic

https://news.ycombinator.com/item?id=48500827

I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.

JBits today at 7:13 PM
Quite a few of the challenges revolve around asking for LLMs to complete tasks reliably and aren't about whether an instance of an LLM completing the task exists. Quite a few of the goalposts are consequently completely changed without the surrounding context, are not the same as what the HN commenter requested and hence seem disingenuous to me.
tamimio today at 7:58 PM
Well I said that before AI will soon make the pcb and electronics just like code, it seems some hw engineers didn’t like it, months later there are few products about the same idea :)
simianwords today at 6:38 PM
[flagged]