linux-nerds.org

Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.

Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).

ChatGPT 'got absolutely wrecked' by Atari 2600 in beginner's chess match — OpenAI's newest model bamboozled by 1970s logic

Technology

193 Beiträge 132 Kommentatoren 0 Aufrufe

O otp@sh.itjust.works

That is more a failure of the person who made that decision than a failing of ChatBots, lol
W This user is from outside of this forum
W This user is from outside of this forum
wewbull@feddit.uk

schrieb zuletzt editiert von

#116

Agreed, which is why it's important to have articles out in the wild that show the shortcomings of AI. If all people read is all the positive crap coming out of companies like OpenAI then they will make stupid decisions.
1 Antwort Letzte Antwort

2
L lifecoach5000@lemmy.world

This post did not contain any content.
A This user is from outside of this forum
A This user is from outside of this forum
arc99@lemmy.world

schrieb zuletzt editiert von

#117

Hardly surprising. Llms aren't -thinking- they're just shitting out the next token for any given input of tokens.
1 Antwort Letzte Antwort

18
A alecsadler@sh.itjust.works

ChatGPT has been, hands down, the worst AI coding assistant I've ever used.

It regularly suggests code that doesn't compile or isn't even for the language.

It generally suggests AC of code that is just a copy of the lines I just wrote.

Sometimes it likes to suggest setting the same property like 5 times.

It is absolute garbage and I do not recommend it to anyone.
A This user is from outside of this forum
A This user is from outside of this forum
arc99@lemmy.world

schrieb zuletzt editiert von

#118

All AIs are the same. They're just scraping content from GitHub, stackoverflow etc with a bunch of guardrails slapped on to spew out sentences that conform to their training data but there is no intelligence. They're super handy for basic code snippets but anyone using them anything remotely complex or nuanced will regret it.
A N 2 Antworten Letzte Antwort

5
S seven_phone@lemmy.world

You say you produce good oranges but my machine for testing apples gave your oranges a very low score.
W This user is from outside of this forum
W This user is from outside of this forum
wizardbeard@lemmy.dbzer0.com

schrieb zuletzt editiert von

#119

No, more like "Your marketing team, sales team, the news media at large, and random hype men all insist your orange machine works amazing on any fruit if you know how to use it right. It didn't work my strawberries when I gave it all the help I could, and was outperformed by my 40 year old strawberry machine. Please stop selling the idea it works on all fruit."

This study is specifically a counter to the constant hype that these LLMs will revolutionize absolutely everything, and the constant word choices used in discussion of LLMs that imply they have reasoning capabilities.
1 Antwort Letzte Antwort

3
N nutsack@lemmy.dbzer0.com

my favorite thing is to constantly be implementing libraries that don't exist
A This user is from outside of this forum
A This user is from outside of this forum
arc99@lemmy.world

schrieb zuletzt editiert von

#120

It's even worse when AI soaks up some project whose APIs are constantly changing. Try using AI to code against jetty for example and you'll be weeping.
1 Antwort Letzte Antwort

1
L lifecoach5000@lemmy.world

This post did not contain any content.
H This user is from outside of this forum
H This user is from outside of this forum
halosheep@lemm.ee

schrieb zuletzt editiert von

#121

I swear every single article critical of current LLMs is like, "The square got BLASTED by the triangle shape when it completely FAILED to go through the triangle shaped hole."
D I L 3 Antworten Letzte Antwort

46
X xavier666@lemm.ee

Have you tried feeding the toddler gallons of baby-food? Maybe then it can play chess
B This user is from outside of this forum
B This user is from outside of this forum
baggachipz@sh.itjust.works

schrieb zuletzt editiert von

#122

They’ve been feeding the toddler everybody else’s baby food and claiming they have the right to.
X 1 Antwort Letzte Antwort

2
I isaamoonkhgdt_6143@lemmy.zip

They used ChatGPT 4o, instead of using o1 or o3.

Obviously it was going to fail.
W This user is from outside of this forum
W This user is from outside of this forum
wizardbeard@lemmy.dbzer0.com

schrieb zuletzt editiert von wizardbeard@lemmy.dbzer0.com

#123

Other studies (not all chess based or against this old chess AI) show similar lackluster results when using reasoning models.

Edit: When comparing reasoning models to existing algorithmic solutions.
1 Antwort Letzte Antwort

0
H halosheep@lemm.ee

I swear every single article critical of current LLMs is like, "The square got BLASTED by the triangle shape when it completely FAILED to go through the triangle shaped hole."
D This user is from outside of this forum
D This user is from outside of this forum
drspod@lemmy.ml

schrieb zuletzt editiert von

#124

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
P I M 3 Antworten Letzte Antwort

37
B baggachipz@sh.itjust.works

They’ve been feeding the toddler everybody else’s baby food and claiming they have the right to.
X This user is from outside of this forum
X This user is from outside of this forum
xavier666@lemm.ee

schrieb zuletzt editiert von

#125

"If we have to ask every time before stealing a little baby food, our morbidly obese toddler cannot survive"
1 Antwort Letzte Antwort

3
A alecsadler@sh.itjust.works

ChatGPT has been, hands down, the worst AI coding assistant I've ever used.

It regularly suggests code that doesn't compile or isn't even for the language.

It generally suggests AC of code that is just a copy of the lines I just wrote.

Sometimes it likes to suggest setting the same property like 5 times.

It is absolute garbage and I do not recommend it to anyone.
I This user is from outside of this forum
I This user is from outside of this forum
ilikeboobies@lemmy.ca

schrieb zuletzt editiert von

#126

I’ve had success with splitting a function into 2 and planning out an overview, though that’s more like talking to myself

I wouldn’t use it to generate stuff though
1 Antwort Letzte Antwort

0
D drspod@lemmy.ml

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
P This user is from outside of this forum
P This user is from outside of this forum
pushbutton@lemmy.world

schrieb zuletzt editiert von

#127

You get 2 triangles in a single square mate...

CHECKMATE!
A 1 Antwort Letzte Antwort

6
M monkdervierte@lemmy.zip

LLM are not built for logic.
P This user is from outside of this forum
P This user is from outside of this forum
pushbutton@lemmy.world

schrieb zuletzt editiert von

#128

And yet everybody is selling to write code.

The last time I checked, coding was requiring logic.
J S 2 Antworten Letzte Antwort

15
F furbag@lemmy.world

Can ChatGPT actually play chess now? Last I checked, it couldn't remember more than 5 moves of history so it wouldn't be able to see the true board state and would make illegal moves, take it's own pieces, materialize pieces out of thin air, etc.
P This user is from outside of this forum
P This user is from outside of this forum
pamasich@kbin.earth

schrieb zuletzt editiert von

#129

There are custom GPTs which claim to play at a stockfish level or be literally stockfish under the hood (I assume the former is still the latter just not explicitly). Haven't tested them, but if they work, I'd say yes. An LLM itself will never be able to play chess or do anything similar, unless they outsource that task to another tool that can. And there seem to be GPTs that do exactly that.

As for why we need ChatGPT then when the result comes from Stockfish anyway, it's for the natural language prompts and responses.
N 1 Antwort Letzte Antwort

0
D drspod@lemmy.ml

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
I This user is from outside of this forum
I This user is from outside of this forum
inconel@lemmy.ca

schrieb zuletzt editiert von

#130

It's also from a company claiming they're getting closer to create morphing shape that can match any hole.
D 1 Antwort Letzte Antwort

16
D drspod@lemmy.ml

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
M This user is from outside of this forum
M This user is from outside of this forum
mrsqueezles@lemmy.world

schrieb zuletzt editiert von

#131

The press release where OpenAI said we'd never need chess players again
1 Antwort Letzte Antwort

3
L lifecoach5000@lemmy.world

This post did not contain any content.
P This user is from outside of this forum
P This user is from outside of this forum
pamasich@kbin.earth

schrieb zuletzt editiert von

#132

Isn't the Atari just a game console, not a chess engine?

Like, Wikipedia doesn't mention anything about the Atari 2600 having a built-in chess engine.

If they were willing to run a chess game on the Atari 2600, why did they not apply the same to ChatGPT? There are custom GPTs which claim to use a stockfish API or play at a similar level.

Like this, it's just unfair. Both platforms are not designed to deal with the task by themselves, but one of them is given the necessary tooling, the other one isn't. No matter what you think of ChatGPT, that's not a fair comparison.
J 1 Antwort Letzte Antwort

0
L lifecoach5000@lemmy.world

This post did not contain any content.
H This user is from outside of this forum
H This user is from outside of this forum
harbinger01173430@lemmy.world

schrieb zuletzt editiert von

#133

Llms useless confirmed once again
1 Antwort Letzte Antwort

2
X x00z@lemmy.world

In all fairness. Machine learning in chess engines is actually pretty strong.

AlphaZero was developed by the artificial intelligence and research company DeepMind, which was acquired by Google. It is a computer program that reached a virtually unthinkable level of play using only reinforcement learning and self-play in order to train its neural networks. In other words, it was only given the rules of the game and then played against itself many millions of times (44 million games in the first nine hours, according to DeepMind).

AlphaZero - Chess Engines

Learn all about the AlphaZero chess program. Everything you need to know about AlphaZero, including what it is, why it is important, and more!

Chess.com (www.chess.com)
J This user is from outside of this forum
J This user is from outside of this forum
jeeva@lemmy.world

schrieb zuletzt editiert von

#134

Sure, but machine learning like that is very different to how LLMs are trained and their output.
1 Antwort Letzte Antwort

1
O objection@lemmy.ml

Tbf, the article should probably mention the fact that machine learning programs designed to play chess blow everything else out of the water.
A This user is from outside of this forum
A This user is from outside of this forum
andallthat@lemmy.world

schrieb zuletzt editiert von andallthat@lemmy.world

#135

Machine learning has existed for many years, now. The issue is with these funding-hungry new companies taking their LLMs, repackaging them as "AI" and attributing every ML win ever to "AI".

ML programs designed and trained specifically to identify tumors in medical imaging have become good diagnostic tools. But if you read in news that "AI helps cure cancer", it makes it sound like it was a lone researcher who spent a few minutes engineering the right prompt for Copilot.

Yes a specifically-designed and finely tuned ML program can now beat the best human chess player, but calling it "AI" and bundling it together with the latest Gemini or Claude iteration's "reasoning capabilities" is intentionally misleading. That's why articles like this one are needed. ML is a useful tool but far from the "super-human general intelligence" that is meant to replace half of human workers by the power of wishful prompting
1 Antwort Letzte Antwort

11

Anmelden zum Antworten

P

France Moves to Classify X as an Adult Site Amid Digital ID Crackdown
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
27

1

377 Stimmen

27 Beiträge

0 Aufrufe

P

False validation is a hell of a drug.
P

The EU Commission fines Delivery Hero and Glovo €329 million for participation in online food delivery cartel
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

1

16 Stimmen

1 Beiträge

1 Aufrufe

Niemand hat geantwortet
P

Is it OK to leave device chargers plugged in all the time? An expert explains
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
41

1

21 Stimmen

41 Beiträge

2 Aufrufe

W

that's because phone makers were pumping out garbage chargers with bare minimum performance for every single phone, isn't it?
A

Cloudflare CEO warns AI and zero-click internet are killing the web's business model
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
83

1

406 Stimmen

83 Beiträge

12 Aufrufe

J

Of course they don't click anything. Google search has just become a front-end for Gemini, the answer is "served" up right at the top and most people will just take that for Gospel.
F

[Opinion] Unending ransomware attacks are a symptom, not the sickness
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
4

1

44 Stimmen

4 Beiträge

2 Aufrufe

G

It varies based on local legislation, so in some places paying ransoms is banned but it's by no means universal. It's totally valid to be against paying ransoms wherever possible, but it's not entirely black and white in some situations. For example, what if a hospital gets ransomed? Say they serve an area not served by other facilities, and if they can't get back online quickly people will die? Sounds dramatic, but critical public services get ransomed all the time and there are undeniable real world consequences. Recovery from ransomware can cost significantly more than a ransom payment if you're not prepared. It can also take months to years to recover, especially if you're simultaneously fighting to evict a persistent (annoyed, unpaid) threat actor from your environment. For the record I don't think ransoms should be paid in most scenarios, but I do think there is some nuance to consider here.
S

California Bill Would Require That AT&T And Comcast Make Broadband Affordable For Poor People
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
9

1

300 Stimmen

9 Beiträge

3 Aufrufe

K

Internet access should be a utility like electricity and water until all three, along with housing, medicine, and food, can be free to all.
C

Chinese chip giants say they don't care about U.S. tariffs — many don't sell to the U.S. anyway due to existing sanctions
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
7

1

0 Stimmen

7 Beiträge

0 Aufrufe

F

It's an actively hostile act, regardless of what your beliefs are on the copyright system.
F

*deleted by creator*
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

1

0 Stimmen

1 Beiträge

1 Aufrufe

Niemand hat geantwortet