linux-nerds.org

Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.

Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).

ChatGPT 'got absolutely wrecked' by Atari 2600 in beginner's chess match — OpenAI's newest model bamboozled by 1970s logic

Technology

202 Beiträge 136 Kommentatoren 1 Aufrufe

P pixelatedsaturn@lemmy.world

I like tab coding, writing small blocks of code that it thinks I need. Its On point almost all the time. This speeds me up.
W This user is from outside of this forum
W This user is from outside of this forum
whoisearth@lemmy.ca

schrieb zuletzt editiert von whoisearth@lemmy.ca

#114

Bingo. If anything what you're finding is the people bitching are the same people that if given a bike wouldn't know how to ride it, which is fair. Some people understand quicker how to use the tools they are given.

Edit - a poor carpenter blames his tools.
1 Antwort Letzte Antwort

6
P pelespirit@sh.itjust.works

Not to help the AI companies, but why don't they program them to look up math programs and outsource chess to other programs when they're asked for that stuff? It's obvious they're shit at it, why do they answer anyway? It's because they're programmed by know-it-all programmers, isn't it.
F This user is from outside of this forum
F This user is from outside of this forum
fmstrat@lemmy.nowsci.com

schrieb zuletzt editiert von

#115

This is where MCP comes in. It's a protocol for LLMs to call standard tools. Basically the LLM would figure out the tool to use from the context, then figure out the order of parameters from those the MCP server says is available, send the JSON, and parse the response.
1 Antwort Letzte Antwort

1
O otp@sh.itjust.works

That is more a failure of the person who made that decision than a failing of ChatBots, lol
W This user is from outside of this forum
W This user is from outside of this forum
wewbull@feddit.uk

schrieb zuletzt editiert von

#116

Agreed, which is why it's important to have articles out in the wild that show the shortcomings of AI. If all people read is all the positive crap coming out of companies like OpenAI then they will make stupid decisions.
1 Antwort Letzte Antwort

2
L lifecoach5000@lemmy.world

This post did not contain any content.
A This user is from outside of this forum
A This user is from outside of this forum
arc99@lemmy.world

schrieb zuletzt editiert von

#117

Hardly surprising. Llms aren't -thinking- they're just shitting out the next token for any given input of tokens.
S 1 Antwort Letzte Antwort

18
A alecsadler@sh.itjust.works

ChatGPT has been, hands down, the worst AI coding assistant I've ever used.

It regularly suggests code that doesn't compile or isn't even for the language.

It generally suggests AC of code that is just a copy of the lines I just wrote.

Sometimes it likes to suggest setting the same property like 5 times.

It is absolute garbage and I do not recommend it to anyone.
A This user is from outside of this forum
A This user is from outside of this forum
arc99@lemmy.world

schrieb zuletzt editiert von

#118

All AIs are the same. They're just scraping content from GitHub, stackoverflow etc with a bunch of guardrails slapped on to spew out sentences that conform to their training data but there is no intelligence. They're super handy for basic code snippets but anyone using them anything remotely complex or nuanced will regret it.
A N 2 Antworten Letzte Antwort

5
S seven_phone@lemmy.world

You say you produce good oranges but my machine for testing apples gave your oranges a very low score.
W This user is from outside of this forum
W This user is from outside of this forum
wizardbeard@lemmy.dbzer0.com

schrieb zuletzt editiert von

#119

No, more like "Your marketing team, sales team, the news media at large, and random hype men all insist your orange machine works amazing on any fruit if you know how to use it right. It didn't work my strawberries when I gave it all the help I could, and was outperformed by my 40 year old strawberry machine. Please stop selling the idea it works on all fruit."

This study is specifically a counter to the constant hype that these LLMs will revolutionize absolutely everything, and the constant word choices used in discussion of LLMs that imply they have reasoning capabilities.
1 Antwort Letzte Antwort

3
N nutsack@lemmy.dbzer0.com

my favorite thing is to constantly be implementing libraries that don't exist
A This user is from outside of this forum
A This user is from outside of this forum
arc99@lemmy.world

schrieb zuletzt editiert von

#120

It's even worse when AI soaks up some project whose APIs are constantly changing. Try using AI to code against jetty for example and you'll be weeping.
1 Antwort Letzte Antwort

1
L lifecoach5000@lemmy.world

This post did not contain any content.
H This user is from outside of this forum
H This user is from outside of this forum
halosheep@lemm.ee

schrieb zuletzt editiert von

#121

I swear every single article critical of current LLMs is like, "The square got BLASTED by the triangle shape when it completely FAILED to go through the triangle shaped hole."
D I L 3 Antworten Letzte Antwort

48
X xavier666@lemm.ee

Have you tried feeding the toddler gallons of baby-food? Maybe then it can play chess
B This user is from outside of this forum
B This user is from outside of this forum
baggachipz@sh.itjust.works

schrieb zuletzt editiert von

#122

They’ve been feeding the toddler everybody else’s baby food and claiming they have the right to.
X 1 Antwort Letzte Antwort

2
I isaamoonkhgdt_6143@lemmy.zip

They used ChatGPT 4o, instead of using o1 or o3.

Obviously it was going to fail.
W This user is from outside of this forum
W This user is from outside of this forum
wizardbeard@lemmy.dbzer0.com

schrieb zuletzt editiert von wizardbeard@lemmy.dbzer0.com

#123

Other studies (not all chess based or against this old chess AI) show similar lackluster results when using reasoning models.

Edit: When comparing reasoning models to existing algorithmic solutions.
1 Antwort Letzte Antwort

0
H halosheep@lemm.ee

I swear every single article critical of current LLMs is like, "The square got BLASTED by the triangle shape when it completely FAILED to go through the triangle shaped hole."
D This user is from outside of this forum
D This user is from outside of this forum
drspod@lemmy.ml

schrieb zuletzt editiert von

#124

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
P I M 3 Antworten Letzte Antwort

38
B baggachipz@sh.itjust.works

They’ve been feeding the toddler everybody else’s baby food and claiming they have the right to.
X This user is from outside of this forum
X This user is from outside of this forum
xavier666@lemm.ee

schrieb zuletzt editiert von

#125

"If we have to ask every time before stealing a little baby food, our morbidly obese toddler cannot survive"
1 Antwort Letzte Antwort

3
A alecsadler@sh.itjust.works

ChatGPT has been, hands down, the worst AI coding assistant I've ever used.

It regularly suggests code that doesn't compile or isn't even for the language.

It generally suggests AC of code that is just a copy of the lines I just wrote.

Sometimes it likes to suggest setting the same property like 5 times.

It is absolute garbage and I do not recommend it to anyone.
I This user is from outside of this forum
I This user is from outside of this forum
ilikeboobies@lemmy.ca

schrieb zuletzt editiert von

#126

I’ve had success with splitting a function into 2 and planning out an overview, though that’s more like talking to myself

I wouldn’t use it to generate stuff though
1 Antwort Letzte Antwort

0
D drspod@lemmy.ml

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
P This user is from outside of this forum
P This user is from outside of this forum
pushbutton@lemmy.world

schrieb zuletzt editiert von

#127

You get 2 triangles in a single square mate...

CHECKMATE!
A 1 Antwort Letzte Antwort

6
M monkdervierte@lemmy.zip

LLM are not built for logic.
P This user is from outside of this forum
P This user is from outside of this forum
pushbutton@lemmy.world

schrieb zuletzt editiert von

#128

And yet everybody is selling to write code.

The last time I checked, coding was requiring logic.
J S 2 Antworten Letzte Antwort

15
F furbag@lemmy.world

Can ChatGPT actually play chess now? Last I checked, it couldn't remember more than 5 moves of history so it wouldn't be able to see the true board state and would make illegal moves, take it's own pieces, materialize pieces out of thin air, etc.
P This user is from outside of this forum
P This user is from outside of this forum
pamasich@kbin.earth

schrieb zuletzt editiert von

#129

There are custom GPTs which claim to play at a stockfish level or be literally stockfish under the hood (I assume the former is still the latter just not explicitly). Haven't tested them, but if they work, I'd say yes. An LLM itself will never be able to play chess or do anything similar, unless they outsource that task to another tool that can. And there seem to be GPTs that do exactly that.

As for why we need ChatGPT then when the result comes from Stockfish anyway, it's for the natural language prompts and responses.
N 1 Antwort Letzte Antwort

0
D drspod@lemmy.ml

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
I This user is from outside of this forum
I This user is from outside of this forum
inconel@lemmy.ca

schrieb zuletzt editiert von

#130

It's also from a company claiming they're getting closer to create morphing shape that can match any hole.
D 1 Antwort Letzte Antwort

17
D drspod@lemmy.ml

It's newsworthy when the sellers of squares are saying that nobody will ever need a triangle again, and the shape-sector of the stock market is hysterically pumping money into companies that make or use squares.
M This user is from outside of this forum
M This user is from outside of this forum
mrsqueezles@lemmy.world

schrieb zuletzt editiert von

#131

The press release where OpenAI said we'd never need chess players again
1 Antwort Letzte Antwort

5
L lifecoach5000@lemmy.world

This post did not contain any content.
P This user is from outside of this forum
P This user is from outside of this forum
pamasich@kbin.earth

schrieb zuletzt editiert von

#132

Isn't the Atari just a game console, not a chess engine?

Like, Wikipedia doesn't mention anything about the Atari 2600 having a built-in chess engine.

If they were willing to run a chess game on the Atari 2600, why did they not apply the same to ChatGPT? There are custom GPTs which claim to use a stockfish API or play at a similar level.

Like this, it's just unfair. Both platforms are not designed to deal with the task by themselves, but one of them is given the necessary tooling, the other one isn't. No matter what you think of ChatGPT, that's not a fair comparison.
J 1 Antwort Letzte Antwort

0
L lifecoach5000@lemmy.world

This post did not contain any content.
H This user is from outside of this forum
H This user is from outside of this forum
harbinger01173430@lemmy.world

schrieb zuletzt editiert von

#133

Llms useless confirmed once again
1 Antwort Letzte Antwort

2

Anmelden zum Antworten

T

How will the space race affect our environment? (Video 25mins)
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

1

3 Stimmen

1 Beiträge

0 Aufrufe

Niemand hat geantwortet
P

Android 16 is here
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
72

1

139 Stimmen

72 Beiträge

0 Aufrufe

B

I feel like Android and Linux (being that it's what Android itself is based on) do the whole "everything is an app" much better than, say, Windows. On Windows, generally speaking, your entire desktop experience is built-in and so tightly coupled that it's hard to switch it out. On Linux, you don't NEED a GUI at all, but if you want one, you'll have a display server, a window manager, etc. On Android, at least without the desktop mode, the base GUI is the launcher, which is just an app. System apps that require root access are still apps. Of course the kernel isn't really an app and I don't think Google Play Services fits most people's definitions of an app. System libraries aren't apps. So those are the parts that you could consider true "OS updates" as opposed to "app updates", but since the "apps" part of the system (if you include system apps) is so much more visible to the user, an OS update will seem like it's mostly a bunch of app updates.
P

A Researcher Figured Out How to Reveal Any Phone Number Linked to a Google Account
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
54

1

517 Stimmen

54 Beiträge

0 Aufrufe

I

Or, how about they fuck off and leave me alone with my private data? I don't want to have to pay for something that should be an irrevocable right. Even if you completely degoogle and whatnot, these cunts will still get hold of your data one way or the other. Its sickening.
M

What editor or IDE do you use and why?
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
37

1

25 Stimmen

37 Beiträge

1 Aufrufe

T

KEIL, because I develop embedded systems.
P

Engineers develop self-healing muscle for robots: Device detects injury, heals it and resets to detect future harm.
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
17

52 Stimmen

17 Beiträge

2 Aufrufe

C

Murderbot is getting closer and closer
P

Discord unveils Discord Orbs, a new in-app currency that users can earn by completing Quests, which reward participants who interact with ads
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
137

1

154 Stimmen

137 Beiträge

5 Aufrufe

B

If you're after text, there are a number of options. If you're after group voice, there are a number of options. You could mix and match both, but "where everyone else is" will also likely be a factor in that kind of decision. If you want both together, then there's probably just Element (Matrix + voice)? Not sure of other options that aren't centralised, where you're the product, or otherwise at obvious risk of enshittifying. (And Element has the smell of the latter to me, but that's another topic). I've prepared for Discord's inevitable "final straw" moment by setting up a Matrix room and maintaining a self-hosted Mumble server in Docker for my gaming buddies. It's worked when Discord has been down, so I know it works. Yet to convince them to test Element...
A

GeForce GTX 970 8GB mod is back for a full review
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

34 Stimmen

1 Beiträge

1 Aufrufe

Niemand hat geantwortet
P

The Collapse of GPT: Will future artificial intelligence systems perform increasingly poorly due to AI-generated material in their training data?
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
7

20 Stimmen

7 Beiträge

0 Aufrufe

A

Fantastic! Me and my 7 legs tank you so much!