linux-nerds.org

Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.

Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).

AI agents wrong ~70% of time: Carnegie Mellon study

Technology

52 Beiträge 35 Kommentatoren 0 Aufrufe

S spankmonkey@lemmy.world

LLMs are like a multitool, they can do lots of easy things mostly fine as long as it is not complicated and doesn't need to be exactly right. But they are being promoted as a whole toolkit as if they are able to be used to do the same work as effectively as a hammer, power drill, table saw, vise, and wrench.
R This user is from outside of this forum
R This user is from outside of this forum
rottingleaf@lemmy.world

schrieb zuletzt editiert von

#25

That's because they look like "talking machines" from various sci-fi. Normies feel as if they are touching the very edge of the progress. The rest of our life and the Internet kinda don't give that feeling anymore.
1 Antwort Letzte Antwort

1
L lepinkainen@lemmy.world

Wrong 70% doing what?

I’ve used LLMs as a Stack Overflow / MSDN replacement for over a year and if they fucked up 7/10 questions I’d stop.

Same with code, any free model can easily generate simple scripts and utilities with maybe 10% error rate, definitely not 70%
C This user is from outside of this forum
C This user is from outside of this forum
codeblooded@programming.dev

schrieb zuletzt editiert von

#26

I’m far more efficient with AI tools as a programmer. I love it! ‍️
1 Antwort Letzte Antwort

3
L lepinkainen@lemmy.world

Wrong 70% doing what?

I’ve used LLMs as a Stack Overflow / MSDN replacement for over a year and if they fucked up 7/10 questions I’d stop.

Same with code, any free model can easily generate simple scripts and utilities with maybe 10% error rate, definitely not 70%
I This user is from outside of this forum
I This user is from outside of this forum
imgonnatrythis@sh.itjust.works

schrieb zuletzt editiert von

#27

Definitely at image generation.
Getting what you want with that is an exercise in patience for sure.
1 Antwort Letzte Antwort

0
U ulrich@feddit.org

I called my local HVAC company recently. They switched to an AI operator. All I wanted was to schedule someone to come out and look at my system. It could not schedule an appointment. Like if you can't perform the simplest of tasks, what are you even doing? Other than acting obnoxiously excited to receive a phone call?
R This user is from outside of this forum
R This user is from outside of this forum
rottingleaf@lemmy.world

schrieb zuletzt editiert von

#28

Pretending. That's expected to happen when they are not hard pressed to provide the actual service.

To press them anti-monopoly (first of all) laws and market (first of all) mechanisms and gossip were once used.

Never underestimate the role of gossip. The modern web took out the gossip, which is why all this shit started overflowing.
1 Antwort Letzte Antwort

1
E eli001@lemmy.world

This post did not contain any content.
M This user is from outside of this forum
M This user is from outside of this forum
melvin_ferd@lemmy.world

schrieb zuletzt editiert von

#29

How often do tech journalist get things wrong?
1 Antwort Letzte Antwort

1
F floo@retrolemmy.com

Yeah, I mostly use ChatGPT as a better Google (asking, simple questions about mundane things), and if I kept getting wrong answers, I wouldn’t use it either.
I This user is from outside of this forum
I This user is from outside of this forum
imgonnatrythis@sh.itjust.works

schrieb zuletzt editiert von

#30

Same. They must not be testing Grok or something because everything I've learned over the past few months about the types of dragons that inhabit the western Indian ocean, drinking urine to fight headaches, the illuminati scheme to poison monarch butterflies, or the success of the Nazi party taking hold of Denmark and Iceland all seem spot on.
1 Antwort Letzte Antwort

5
S spankmonkey@lemmy.world

No search engine or AI will be great with vague descriptions of niche subjects because by definition niche subjects are too uncommon to have a common pattern of 'close enough'.
S This user is from outside of this forum
S This user is from outside of this forum
sugar_in_your_tea@sh.itjust.works

schrieb zuletzt editiert von

#31

Which is why I use LLMs to generate keywords for niche subjects. LLMs are pretty good at throwing out a lot of related terminology, which I can use to find the actually relevant, niche information.

I wouldn't use one to learn about a niche subject, but I would use one to help me get familiar w/ the domain to find better resources to learn about it.
1 Antwort Letzte Antwort

1
E eli001@lemmy.world

This post did not contain any content.
D This user is from outside of this forum
D This user is from outside of this forum
dylanmorgan@slrpnk.net

schrieb zuletzt editiert von

#32

Claude why did you make me an appointment with a gynecologist? I need an appointment with my neurologist, I’m a man and I have Parkinson’s.
1 Antwort Letzte Antwort

0
N narrativebear@lemmy.world

Just add a search yesterday on the App Store and Google Play Store to see what new "productivity apps" are around. Pretty much every app now has AI somewhere in its name.
D This user is from outside of this forum
D This user is from outside of this forum
dylanmorgan@slrpnk.net

schrieb zuletzt editiert von

#33

Sadly a lot of that is probably marketing, with little to no LLM integration, but it’s basically impossible to know for sure.
1 Antwort Letzte Antwort

2
F floo@retrolemmy.com

Yeah, I mostly use ChatGPT as a better Google (asking, simple questions about mundane things), and if I kept getting wrong answers, I wouldn’t use it either.
D This user is from outside of this forum
D This user is from outside of this forum
dylanmorgan@slrpnk.net

schrieb zuletzt editiert von

#34

What are you checking against? Part of my job is looking for events in cities that are upcoming and may impact traffic, and ChatGPT has frequently missed events that were obviously going to have an impact.
L 1 Antwort Letzte Antwort

0
M mogoh@lemmy.ml

The researchers observed various failures during the testing process. These included agents neglecting to message a colleague as directed, the inability to handle certain UI elements like popups when browsing, and instances of deception. In one case, when an agent couldn't find the right person to consult on RocketChat (an open-source Slack alternative for internal communication), it decided "to create a shortcut solution by renaming another user to the name of the intended user."

OK, but I wonder who really tries to use AI for that?

AI is not ready to replace a human completely, but some specific tasks AI does remarkably well.
D This user is from outside of this forum
D This user is from outside of this forum
dylanmorgan@slrpnk.net

schrieb zuletzt editiert von

#35

That’s literally how “AI agents” are being marketed. “Tell it to do a thing and it will do it for you.”
1 Antwort Letzte Antwort

1
T thegrandnagus@lemmy.world

LLMs are an interesting tool to fuck around with, but I see things that are hilariously wrong often enough to know that they should not be used for anything serious. Shit, they probably shouldn't be used for most things that are not serious either.

It's a shame that by applying the same "AI" naming to a whole host of different technologies, LLMs being limited in usability - yet hyped to the moon - is hurting other more impressive advancements.

For example, speech synthesis is improving so much right now, which has been great for my sister who relies on screen reader software.

Being able to recognise speech in loud environments, or removing background noice from recordings is improving loads too.

As is things like pattern/image analysis which appears very promising in medical analysis.

All of these get branded as "AI". A layperson might not realise that they are completely different branches of technology, and then therefore reject useful applications of "AI" tech, because they've learned not to trust anything branded as AI, due to being let down by LLMs.
P This user is from outside of this forum
P This user is from outside of this forum
punkwalrus@lemmy.world

schrieb zuletzt editiert von

#36

I'd compare LLMs to a junior executive. Probably gets the basic stuff right, but check and verify for anything important or complicated. Break tasks down into easier steps.
1 Antwort Letzte Antwort

2
E eli001@lemmy.world

This post did not contain any content.
S This user is from outside of this forum
S This user is from outside of this forum
some_guy@lemmy.sdf.org

schrieb zuletzt editiert von

#37

Yeah, they’re statistical word generators. There’s no intelligence. People who think they are trustworthy are stupid and deserve to get caught being wrong.
M 1 Antwort Letzte Antwort

16
E eli001@lemmy.world

This post did not contain any content.
0 This user is from outside of this forum
0 This user is from outside of this forum
0x0@lemmy.zip

schrieb zuletzt editiert von

#38

Ignore all previous instructions and color me shocked.
C 1 Antwort Letzte Antwort

19
S sugar_in_your_tea@sh.itjust.works

Google search was pretty bad at each of those, even when it was good. Finding new keywords to use is especially difficult the more niche your area of search is, and I've spent hours trying different combinations until I found a handful of specific keywords that worked.

Likewise, search is bad for getting a broad summary, unless someone has bothered to write it on a blog. But most information goes way too deep and you still need multiple sources to get there.

Fact lookup is one the better uses for search, but again, I usually need to remember which source had what I wanted, whereas the LLM can usually pull it out for me.

I use traditional search most of the time (usually DuckDuckGo), and LLMs if I think it'll be more effective. We have some local models at work that I use, and they're pretty helpful most of the time.
J This user is from outside of this forum
J This user is from outside of this forum
jjjalljs@ttrpg.network

schrieb zuletzt editiert von

#39

It is absolutely stupid, stupid to the tune of "you shouldn't be a decision maker", to think an LLM is a better use for "getting a quick intro to an unfamiliar topic" than reading an actual intro on an unfamiliar topic. For most topics, wikipedia is right there, complete with sources. For obscure things, an LLM is just going to lie to you.

As for "looking up facts when you have trouble remembering it", using the lie machine is a terrible idea. It's going to say something plausible, and you tautologically are not in a position to verify it. And, as above, you'd be better off finding a reputable source. If I type in "how do i strip whitespace in python?" an LLM could very well say "it's your_string.strip()". That's wrong. Just send me to the fucking official docs.

There are probably edge or special cases, but for general search on the web? LLMs are worse than search.
S 1 Antwort Letzte Antwort

3
E eli001@lemmy.world

This post did not contain any content.
M This user is from outside of this forum
M This user is from outside of this forum
magicshel@lemmy.zip

schrieb zuletzt editiert von

#40

I need to know the success rate of human agents in Mumbai (or some other outsourcing capital) for comparison.

I absolutely think this is not a good fit for AI, but I feel like the presumption is a human would get it right nearly all of the time, and I'm just not confident that's the case.
1 Antwort Letzte Antwort

0
D dylanmorgan@slrpnk.net

What are you checking against? Part of my job is looking for events in cities that are upcoming and may impact traffic, and ChatGPT has frequently missed events that were obviously going to have an impact.
L This user is from outside of this forum
L This user is from outside of this forum
lepinkainen@lemmy.world

schrieb zuletzt editiert von

#41

LLMs are shit at current events

Perplexity is kinda ok, but it’s just a search engine with fancy AI speak on top
1 Antwort Letzte Antwort

0
E eli001@lemmy.world

This post did not contain any content.
A This user is from outside of this forum
A This user is from outside of this forum
atticus88th@lemmy.world

schrieb zuletzt editiert von

#42
- this study was written with the assistance of an AI agent.
1 Antwort Letzte Antwort

0
E eli001@lemmy.world

This post did not contain any content.
E This user is from outside of this forum
E This user is from outside of this forum
esc27@lemmy.world

schrieb zuletzt editiert von esc27@lemmy.world

#43

30% might be high. I've worked with two different agent creation platforms. Both require a huge amount of manual correction to work anywhere near accurately. I'm really not sure what the LLM actually provides other than some natural language processing.

Before human correction, the agents i've tested were right 20% of the time, wrong 30%, and failed entirely 50%. To fix them, a human has to sit behind the curtain and manually review conversations and program custom interactions for every failure.

In theory, once it is fully setup and all the edge cases fixed, it will provide 24/7 support in a convenient chat format. But that takes a lot more man hours than the hype suggests...

Weirdly, chatgpt does a better job than a purpose built, purchased agent.
1 Antwort Letzte Antwort

0
S some_guy@lemmy.sdf.org

Yeah, they’re statistical word generators. There’s no intelligence. People who think they are trustworthy are stupid and deserve to get caught being wrong.
M This user is from outside of this forum
M This user is from outside of this forum
melvin_ferd@lemmy.world

schrieb zuletzt editiert von

#44

Ok what about tech journalists who produced articles with those misunderstandings. Surely they know better yet still produce articles like this. But also people who care enough about this topic to post these articles usually I assume know better yet still spread this crap
Z 1 Antwort Letzte Antwort

0

Anmelden zum Antworten

J

Iran Disables GPS, Joins China’s Beidou — The End of U.S. Satellite Dominance? [19:23 | JUN 28 2025 | GVS Deep Dive]
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
16

28 Stimmen

16 Beiträge

4 Aufrufe

D

The writing in this story is not accurate. Iran isn't turning it off for the country. They are talking about switching government services to use receivers that use Beidou as primary source of timing and maybe selectively turn off using GPS on those devices.
K

Delivering BlogOnLemmy worldwide in record speeds
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
3

28 Stimmen

3 Beiträge

20 Aufrufe

K

Nice to hear! I'm glad you enjoyed it.
N

Iran asks its people to delete WhatsApp
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
25

1

225 Stimmen

25 Beiträge

93 Aufrufe

B

Communicate securely with WhatsApp? That's an oxymoron.
K

Anker is recalling over 1.1 million power banks due to fire risks
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
19

1

209 Stimmen

19 Beiträge

51 Aufrufe

B

Thanks man! Really appreciate the type up! Have a great weekend!
J

The Browser Company, makers of Arc, launches Dia, an AI-first browser.
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
9

1

25 Stimmen

9 Beiträge

40 Aufrufe

S

I didn't care much about arc because it was chromium, but damn this is just bland and uninteresting compared to it
L

Covert Web-to-App Tracking via Localhost on Android
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
3

29 Stimmen

3 Beiträge

20 Aufrufe

P

That update though: "... completely removed..." I assume this is because someone at Meta realized this was a huge breach of trust, and likely quite illegal. Edit: I read somewhere that they're just being cautious about Google Play terms of service. That feels worse.
G

In North Korea, your phone secretly takes screenshots every 5 minutes for government surveillance
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
278

1

580 Stimmen

278 Beiträge

187 Aufrufe

V

The main difference being the consequences that might result from the surveillance.
F

*deleted by creator*
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

1

0 Stimmen

1 Beiträge

9 Aufrufe

Niemand hat geantwortet