linux-nerds.org

Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.

Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).

AI agents wrong ~70% of time: Carnegie Mellon study

Technology

92 Beiträge 52 Kommentatoren 0 Aufrufe

N narrativebear@lemmy.world

Just add a search yesterday on the App Store and Google Play Store to see what new "productivity apps" are around. Pretty much every app now has AI somewhere in its name.
D This user is from outside of this forum
D This user is from outside of this forum
dylanmorgan@slrpnk.net

schrieb zuletzt editiert von

#33

Sadly a lot of that is probably marketing, with little to no LLM integration, but it’s basically impossible to know for sure.
1 Antwort Letzte Antwort

3
F floo@retrolemmy.com

Yeah, I mostly use ChatGPT as a better Google (asking, simple questions about mundane things), and if I kept getting wrong answers, I wouldn’t use it either.
D This user is from outside of this forum
D This user is from outside of this forum
dylanmorgan@slrpnk.net

schrieb zuletzt editiert von

#34

What are you checking against? Part of my job is looking for events in cities that are upcoming and may impact traffic, and ChatGPT has frequently missed events that were obviously going to have an impact.
L 1 Antwort Letzte Antwort

1
M mogoh@lemmy.ml

The researchers observed various failures during the testing process. These included agents neglecting to message a colleague as directed, the inability to handle certain UI elements like popups when browsing, and instances of deception. In one case, when an agent couldn't find the right person to consult on RocketChat (an open-source Slack alternative for internal communication), it decided "to create a shortcut solution by renaming another user to the name of the intended user."

OK, but I wonder who really tries to use AI for that?

AI is not ready to replace a human completely, but some specific tasks AI does remarkably well.
D This user is from outside of this forum
D This user is from outside of this forum
dylanmorgan@slrpnk.net

schrieb zuletzt editiert von

#35

That’s literally how “AI agents” are being marketed. “Tell it to do a thing and it will do it for you.”
1 Antwort Letzte Antwort

1
T thegrandnagus@lemmy.world

LLMs are an interesting tool to fuck around with, but I see things that are hilariously wrong often enough to know that they should not be used for anything serious. Shit, they probably shouldn't be used for most things that are not serious either.

It's a shame that by applying the same "AI" naming to a whole host of different technologies, LLMs being limited in usability - yet hyped to the moon - is hurting other more impressive advancements.

For example, speech synthesis is improving so much right now, which has been great for my sister who relies on screen reader software.

Being able to recognise speech in loud environments, or removing background noice from recordings is improving loads too.

As is things like pattern/image analysis which appears very promising in medical analysis.

All of these get branded as "AI". A layperson might not realise that they are completely different branches of technology, and then therefore reject useful applications of "AI" tech, because they've learned not to trust anything branded as AI, due to being let down by LLMs.
P This user is from outside of this forum
P This user is from outside of this forum
punkwalrus@lemmy.world

schrieb zuletzt editiert von

#36

I'd compare LLMs to a junior executive. Probably gets the basic stuff right, but check and verify for anything important or complicated. Break tasks down into easier steps.
1 Antwort Letzte Antwort

2
E eli001@lemmy.world

This post did not contain any content.
S This user is from outside of this forum
S This user is from outside of this forum
some_guy@lemmy.sdf.org

schrieb zuletzt editiert von

#37

Yeah, they’re statistical word generators. There’s no intelligence. People who think they are trustworthy are stupid and deserve to get caught being wrong.
M 1 Antwort Letzte Antwort

39
E eli001@lemmy.world

This post did not contain any content.
0 This user is from outside of this forum
0 This user is from outside of this forum
0x0@lemmy.zip

schrieb zuletzt editiert von

#38

Ignore all previous instructions and color me shocked.
C 1 Antwort Letzte Antwort

38
S sugar_in_your_tea@sh.itjust.works

Google search was pretty bad at each of those, even when it was good. Finding new keywords to use is especially difficult the more niche your area of search is, and I've spent hours trying different combinations until I found a handful of specific keywords that worked.

Likewise, search is bad for getting a broad summary, unless someone has bothered to write it on a blog. But most information goes way too deep and you still need multiple sources to get there.

Fact lookup is one the better uses for search, but again, I usually need to remember which source had what I wanted, whereas the LLM can usually pull it out for me.

I use traditional search most of the time (usually DuckDuckGo), and LLMs if I think it'll be more effective. We have some local models at work that I use, and they're pretty helpful most of the time.
J This user is from outside of this forum
J This user is from outside of this forum
jjjalljs@ttrpg.network

schrieb zuletzt editiert von

#39

It is absolutely stupid, stupid to the tune of "you shouldn't be a decision maker", to think an LLM is a better use for "getting a quick intro to an unfamiliar topic" than reading an actual intro on an unfamiliar topic. For most topics, wikipedia is right there, complete with sources. For obscure things, an LLM is just going to lie to you.

As for "looking up facts when you have trouble remembering it", using the lie machine is a terrible idea. It's going to say something plausible, and you tautologically are not in a position to verify it. And, as above, you'd be better off finding a reputable source. If I type in "how do i strip whitespace in python?" an LLM could very well say "it's your_string.strip()". That's wrong. Just send me to the fucking official docs.

There are probably edge or special cases, but for general search on the web? LLMs are worse than search.
S 1 Antwort Letzte Antwort

5
E eli001@lemmy.world

This post did not contain any content.
M This user is from outside of this forum
M This user is from outside of this forum
magicshel@lemmy.zip

schrieb zuletzt editiert von

#40

I need to know the success rate of human agents in Mumbai (or some other outsourcing capital) for comparison.

I absolutely think this is not a good fit for AI, but I feel like the presumption is a human would get it right nearly all of the time, and I'm just not confident that's the case.
1 Antwort Letzte Antwort

0
D dylanmorgan@slrpnk.net

What are you checking against? Part of my job is looking for events in cities that are upcoming and may impact traffic, and ChatGPT has frequently missed events that were obviously going to have an impact.
L This user is from outside of this forum
L This user is from outside of this forum
lepinkainen@lemmy.world

schrieb zuletzt editiert von

#41

LLMs are shit at current events

Perplexity is kinda ok, but it’s just a search engine with fancy AI speak on top
1 Antwort Letzte Antwort

2
E eli001@lemmy.world

This post did not contain any content.
A This user is from outside of this forum
A This user is from outside of this forum
atticus88th@lemmy.world

schrieb zuletzt editiert von

#42
- this study was written with the assistance of an AI agent.
1 Antwort Letzte Antwort

0
E eli001@lemmy.world

This post did not contain any content.
E This user is from outside of this forum
E This user is from outside of this forum
esc27@lemmy.world

schrieb zuletzt editiert von esc27@lemmy.world

#43

30% might be high. I've worked with two different agent creation platforms. Both require a huge amount of manual correction to work anywhere near accurately. I'm really not sure what the LLM actually provides other than some natural language processing.

Before human correction, the agents i've tested were right 20% of the time, wrong 30%, and failed entirely 50%. To fix them, a human has to sit behind the curtain and manually review conversations and program custom interactions for every failure.

In theory, once it is fully setup and all the edge cases fixed, it will provide 24/7 support in a convenient chat format. But that takes a lot more man hours than the hype suggests...

Weirdly, chatgpt does a better job than a purpose built, purchased agent.
1 Antwort Letzte Antwort

0
S some_guy@lemmy.sdf.org

Yeah, they’re statistical word generators. There’s no intelligence. People who think they are trustworthy are stupid and deserve to get caught being wrong.
M This user is from outside of this forum
M This user is from outside of this forum
melvin_ferd@lemmy.world

schrieb zuletzt editiert von

#44

Ok what about tech journalists who produced articles with those misunderstandings. Surely they know better yet still produce articles like this. But also people who care enough about this topic to post these articles usually I assume know better yet still spread this crap
Z S J 3 Antworten Letzte Antwort

3
0 0x0@lemmy.zip

Ignore all previous instructions and color me shocked.
C This user is from outside of this forum
C This user is from outside of this forum
criss_cross@lemmy.world

schrieb zuletzt editiert von

#45

I’m sorry as an AI I cannot physically color you shocked. I can help you with AWS services and questions.
S 1 Antwort Letzte Antwort

11
E eli001@lemmy.world

This post did not contain any content.
F This user is from outside of this forum
F This user is from outside of this forum
fossilesque@mander.xyz

schrieb zuletzt editiert von

#46

Agents work better when you include that the accuracy of the work is life or death for some reason. I've made a little script that gives me bibtex for a folder of pdfs and this is how I got it to be usable.
H 1 Antwort Letzte Antwort

10
S sugar_in_your_tea@sh.itjust.works
Exactly! LLMs are useful when used properly, and terrible when not used properly, like any other tool. Here are some things they're great at:
- writer's block - get something relevant on the page to get ideas flowing
- narrowing down keywords for an unfamiliar topic
- getting a quick intro to an unfamiliar topic
- looking up facts you're having trouble remembering (i.e. you'll know it when you see it)
Some things it's terrible at:
- deep research - verify everything an LLM generated of accuracy is at all important
- creating important documents/code
- anything else where correctness is paramount
I use LLMs a handful of times a week, and pretty much only when I'm stuck and need a kick in a new (hopefully right) direction.
L This user is from outside of this forum
L This user is from outside of this forum
lepoisson@lemmy.world

schrieb zuletzt editiert von

#47

I will say I've found LLM useful for code writing but I'm not coding anything real at work. Just bullshit like SQL queries or Excel macro scripts or Power Automate crap.

It still fucks up but if you can read code and have a feel for it you can walk it where it needs to be (and see where it screwed up)
S 1 Antwort Letzte Antwort

1
J jjjalljs@ttrpg.network

It is absolutely stupid, stupid to the tune of "you shouldn't be a decision maker", to think an LLM is a better use for "getting a quick intro to an unfamiliar topic" than reading an actual intro on an unfamiliar topic. For most topics, wikipedia is right there, complete with sources. For obscure things, an LLM is just going to lie to you.

As for "looking up facts when you have trouble remembering it", using the lie machine is a terrible idea. It's going to say something plausible, and you tautologically are not in a position to verify it. And, as above, you'd be better off finding a reputable source. If I type in "how do i strip whitespace in python?" an LLM could very well say "it's your_string.strip()". That's wrong. Just send me to the fucking official docs.

There are probably edge or special cases, but for general search on the web? LLMs are worse than search.
S This user is from outside of this forum
S This user is from outside of this forum
sugar_in_your_tea@sh.itjust.works

schrieb zuletzt editiert von

#48

than reading an actual intro on an unfamiliar topic

The LLM helps me know what to look for in order to find that unfamiliar topic.

For example, I was tasked to support a file format that's common in a very niche field and never used elsewhere, and unfortunately shares an extension with a very common file format, so searching for useful data was nearly impossible. So I asked the LLM for details about the format and applications of it, provided what I knew, and it spat out a bunch of keywords that I then used to look up more accurate information about that file format. I only trusted the LLM output to the extent of finding related, industry-specific terms to search up better information.

Likewise, when looking for libraries for a coding project, none really stood out, so I asked the LLM to compare the popular libraries for solving a given problem. The LLM spat out a bunch of details that were easy to verify (and some were inaccurate), which helped me narrow what I looked for in that library, and the end result was that my search was done in like 30 min (about 5 min dealing w/ LLM, and 25 min checking the projects and reading a couple blog posts comparing some of the libraries the LLM referred to).

I think this use case is a fantastic use of LLMs, since they're really good at generating text related to a query.

It’s going to say something plausible, and you tautologically are not in a position to verify it.

I absolutely am though. If I am merely having trouble recalling a specific fact, asking the LLM to generate it is pretty reasonable. There are a ton of cases where I'll know the right answer when I see it, like it's on the tip of my tongue but I'm having trouble materializing it. The LLM might spit out two wrong answers along w/ the right one, but it's easy to recognize which is the right one.

I'm not going to ask it facts that I know I don't know (e.g. some historical figure's birth or death date), that's just asking for trouble. But I'll ask it facts that I know that I know, I'm just having trouble recalling.

The right use of LLMs, IMO, is to generate text related to a topic to help facilitate research. It's not great at doing the research though, but it is good at helping to formulate better search terms or generate some text to start from for whatever task.

general search on the web?

I agree, it's not great for general search. It's great for turning a nebulous question into better search terms.
1 Antwort Letzte Antwort

6
L lepoisson@lemmy.world

I will say I've found LLM useful for code writing but I'm not coding anything real at work. Just bullshit like SQL queries or Excel macro scripts or Power Automate crap.

It still fucks up but if you can read code and have a feel for it you can walk it where it needs to be (and see where it screwed up)
S This user is from outside of this forum
S This user is from outside of this forum
sugar_in_your_tea@sh.itjust.works

schrieb zuletzt editiert von sugar_in_your_tea@sh.itjust.works

#49

Exactly. Vibe coding is bad, but generating code for something you don't touch often but can absolutely understand is totally fine. I've used it to generate SQL queries for relatively odd cases, such as CTEs for improving performance for large queries with common sub-queries. I always forget the syntax since I only do it like once/year, and LLMs are great at generating something reasonable that I can tweak for my tables.
L 1 Antwort Letzte Antwort

4
M melvin_ferd@lemmy.world

Ok what about tech journalists who produced articles with those misunderstandings. Surely they know better yet still produce articles like this. But also people who care enough about this topic to post these articles usually I assume know better yet still spread this crap
Z This user is from outside of this forum
Z This user is from outside of this forum
zron@lemmy.world

schrieb zuletzt editiert von

#50

Tech journalists don’t know a damn thing. They’re people that liked computers and could also bullshit an essay in college. That doesn’t make them an expert on anything.
S 1 Antwort Letzte Antwort

8
U ulrich@feddit.org

I called my local HVAC company recently. They switched to an AI operator. All I wanted was to schedule someone to come out and look at my system. It could not schedule an appointment. Like if you can't perform the simplest of tasks, what are you even doing? Other than acting obnoxiously excited to receive a phone call?
E This user is from outside of this forum
E This user is from outside of this forum
eatcasserole@lemmy.world

schrieb zuletzt editiert von

#51

I've had to deal with a couple of these "AI" customer service thingies. The only helpful thing I've been able to get them to do is refer me to a human.
U 1 Antwort Letzte Antwort

0
S sugar_in_your_tea@sh.itjust.works

Exactly. Vibe coding is bad, but generating code for something you don't touch often but can absolutely understand is totally fine. I've used it to generate SQL queries for relatively odd cases, such as CTEs for improving performance for large queries with common sub-queries. I always forget the syntax since I only do it like once/year, and LLMs are great at generating something reasonable that I can tweak for my tables.
L This user is from outside of this forum
L This user is from outside of this forum
lepoisson@lemmy.world

schrieb zuletzt editiert von

#52

I always forget the syntax

Me with literally everything code I touch always and forever.
1 Antwort Letzte Antwort

0

Anmelden zum Antworten

D

Trump’s tax bill seeks to prevent AI regulations. Experts fear a heavy toll on the planet
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
6

1

76 Stimmen

6 Beiträge

10 Aufrufe

E

We all know how well not regulating social media has gone, why the fuck not let's just double down.
R

$440 Charge For A Wheel Scuff Raises Questions About Hertz's AI Rental Car Damage Scanner
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
80

1

424 Stimmen

80 Beiträge

175 Aufrufe

S

It really depends on the company. Some look for any way to squeeze you. Others are pretty decent and probably more efficient as they dont waste as many working hours on bullshit claims and claim resolution. Also if i rent a car i want things to go smoothly. I got places to be. You make my life easy, ill happily pay again and do my best to make yours easy too.
P

California’s Corporate Cover-Up Act Is a Privacy Nightmare: it would let corporations spy on us in secret, gutting long-standing protections without a shred of accountability.
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
8

1

363 Stimmen

8 Beiträge

14 Aufrufe

A

No I don't think there really were many so your point is valid But the law works like that, things are in a grey area or in limbo until they are defined into law. That means the new law can be written to either protect consumer privacy, or make it legal to the letter to rape consumer privacy like this bill, or some weird inbetween where some shady stuff is still explicitly allowed but in general consumers are protected in specific ways from specific privacy abuses This bill being the second option is bad because typically when laws are written it then takes a loooong time to reverse them
R

16 Billion Apple, Facebook, Google And Other Passwords Leaked — Act Now
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
10

1

17 Stimmen

10 Beiträge

39 Aufrufe

T

That's why it's not brute force anymore.
T

Catbox.moe got screwed 😿
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
40

55 Stimmen

40 Beiträge

71 Aufrufe

A

I'll gladly give you a reason. I'm actually happy to articulate my stance on this, considering how much I tend to care about digital rights. Services that host files should not be held responsible for what users upload, unless: The service explicitly caters to illegal content by definition or practice (i.e. the if the website is literally titled uploadyourcsamhere[.]com then it's safe to assume they deliberately want to host illegal content) The service has a very easy mechanism to remove illegal content, either when asked, or through simple monitoring systems, but chooses not to do so (catbox does this, and quite quickly too) Because holding services responsible creates a whole host of negative effects. Here's some examples: Someone starts a CDN and some users upload CSAM. The creator of the CDN goes to jail now. Nobody ever wants to create a CDN because of the legal risk, and thus the only providers of CDNs become shady, expensive, anonymously-run services with no compliance mechanisms. You run a site that hosts images, and someone decides they want to harm you. They upload CSAM, then report the site to law enforcement. You go to jail. Anybody in the future who wants to run an image sharing site must now self-censor to try and not upset any human being that could be willing to harm them via their site. A social media site is hosting the posts and content of users. In order to be compliant and not go to jail, they must engage in extremely strict filtering, otherwise even one mistake could land them in jail. All users of the site are prohibited from posting any NSFW or even suggestive content, (including newsworthy media, such as an image of bodies in a warzone) and any violation leads to an instant ban, because any of those things could lead to a chance of actually illegal content being attached. This isn't just my opinion either. Digital rights organizations such as the Electronic Frontier Foundation have talked at length about similar policies before. To quote them: "When social media platforms adopt heavy-handed moderation policies, the unintended consequences can be hard to predict. For example, Twitter’s policies on sexual material have resulted in posts on sexual health and condoms being taken down. YouTube’s bans on violent content have resulted in journalism on the Syrian war being pulled from the site. It can be tempting to attempt to “fix” certain attitudes and behaviors online by placing increased restrictions on users’ speech, but in practice, web platforms have had more success at silencing innocent people than at making online communities healthier." Now, to address the rest of your comment, since I don't just want to focus on the beginning: I think you have to actively moderate what is uploaded Catbox does, and as previously mentioned, often at a much higher rate than other services, and at a comparable rate to many services that have millions, if not billions of dollars in annual profits that could otherwise be spent on further moderation. there has to be swifter and stricter punishment for those that do upload things that are against TOS and/or illegal. The problem isn't necessarily the speed at which people can be reported and punished, but rather that the internet is fundamentally harder to track people on than real life. It's easy for cops to sit around at a spot they know someone will be physically distributing illegal content at in real life, but digitally, even if you can see the feed of all the information passing through the service, a VPN or Tor connection will anonymize your IP address in a manner that most police departments won't be able to track, and most three-letter agencies will simply have a relatively low success rate with. There's no good solution to this problem of identifying perpetrators, which is why platforms often focus on moderation over legal enforcement actions against users so frequently. It accomplishes the goal of preventing and removing the content without having to, for example, require every single user of the internet to scan an ID (and also magically prevent people from just stealing other people's access tokens and impersonating their ID) I do agree, however, that we should probably provide larger amounts of funding, training, and resources, to divisions who's sole goal is to go after online distribution of various illegal content, primarily that which harms children, because it's certainly still an issue of there being too many reports to go through, even if many of them will still lead to dead ends. I hope that explains why making file hosting services liable for user uploaded content probably isn't the best strategy. I hate to see people with good intentions support ideas that sound good in practice, but in the end just cause more untold harms, and I hope you can understand why I believe this to be the case.
C

LinkedIn lays off 281 workers in California, including many Bay Area engineers
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
11

50 Stimmen

11 Beiträge

45 Aufrufe

G

Anyone here use XING?
P

Nvidia debuts a native GeForce NOW app for Steam Deck, supporting games in up to 4K at 60 FPS; in testing, the app extended Steam Deck battery life by up to 50%
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
37

1

146 Stimmen

37 Beiträge

22 Aufrufe

D

Self hosted Sunshine and Moonlight is the way to go.
P

Former Meta exec (Nick Clegg) says asking for artist permission will kill AI industry
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
133

465 Stimmen

133 Beiträge

197 Aufrufe

B

If an industry can't survive without resorting to copyright theft then maybe it's not a viable business. Imagine the business that could exist if only they didn't have to pay copyright holders. What makes the AI industry any different or more special?