linux-nerds.org

Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.

Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).

AI agents wrong ~70% of time: Carnegie Mellon study

Technology

154 Beiträge 76 Kommentatoren 3 Aufrufe

H hertzdentalbar@lemmy.blahaj.zone

Did you make it? Or did you prompt it? They ain't quite the same.
F This user is from outside of this forum
F This user is from outside of this forum
fossilesque@mander.xyz

schrieb zuletzt editiert von fossilesque@mander.xyz

#90

It calls ollama with a prompt, it's a bit complex because it renames and moves stuff too and sorts it.
1 Antwort Letzte Antwort

0
O outhouseperilous@lemmy.dbzer0.com

It's absolutely dangerous but it doesnt have to work even a little to do damage; hell, it already has. Your thing just makes it sound much more capable than it is. And it is not.

Also, it's not AI.
J This user is from outside of this forum
J This user is from outside of this forum
jsomae@lemmy.ml

schrieb zuletzt editiert von

#91

semantics.
O 1 Antwort Letzte Antwort

3
J jsomae@lemmy.ml

semantics.
O This user is from outside of this forum
O This user is from outside of this forum
outhouseperilous@lemmy.dbzer0.com

schrieb zuletzt editiert von

#92

No, it matters. Youre pushing the lie they want pushed.
H 1 Antwort Letzte Antwort

1
S sheogorath@lemmy.world

I won't tolerate Jan slander here. I know he's just a builder, but his life path has the most probability of having a great person out of it!
C This user is from outside of this forum
C This user is from outside of this forum
cavemanfreak@programming.dev

schrieb zuletzt editiert von

#93

I'd say Jan Botanist is also up there as being a pretty great person.
S 1 Antwort Letzte Antwort

2
C cavemanfreak@programming.dev

I'd say Jan Botanist is also up there as being a pretty great person.
S This user is from outside of this forum
S This user is from outside of this forum
sheogorath@lemmy.world

schrieb zuletzt editiert von

#94

Jan Refiner is up there for me.
1 Antwort Letzte Antwort

1
K kameecoding@lemmy.world

For me as a software developer the accuracy is more in the 95%+ range.

On one hand the built in copilot chat widget in Intellij basically replaces a lot my google queries.

On the other hand it is rather fucking good at executing some rewrites that is a fucking chore to do manually, but can easily be done by copilot.

Imagine you have a script that initializes your DB with some test data. You have an Insert into statement with lots of columns and rows so

Inser into (column1,....,column n)
Values row1,
Row 2
Row n

Addig a new column with test data for each row is a PITA, but copilot handles it without issue.

Similarly when writing unit tests you do a lot of edge case testing which is a bunch of almost same looking tests with maybe one variable changing, at most you write one of those tests, then copilot will auto generate the rest after you name the next unit test, pretty good at guessing what you want to do in that test, at least with my naming scheme.

So yeah, it's way overrated for many-many things, but for programming it's a pretty awesome productivity tool.
N This user is from outside of this forum
N This user is from outside of this forum
nalivai@discuss.tchncs.de

schrieb zuletzt editiert von

#95

Keep doing what you do. Your company will pay me handsomely to throw out all your bullshit and write working code you can trust when you're done. If your company wants to have a product in the future that is.
K 1 Antwort Letzte Antwort

3
S shayeta@feddit.org

It doesn't matter if you need a human to review. AI has no way distinguishing between success and failure. Either way a human will have to review 100% of those tasks.
O This user is from outside of this forum
O This user is from outside of this forum
outbound7404@lemmy.ml

schrieb zuletzt editiert von

#96

A human can review something close to correct a lot better than starting the task from zero.
D M 2 Antworten Letzte Antwort

2
N nalivai@discuss.tchncs.de

Keep doing what you do. Your company will pay me handsomely to throw out all your bullshit and write working code you can trust when you're done. If your company wants to have a product in the future that is.
K This user is from outside of this forum
K This user is from outside of this forum
kameecoding@lemmy.world

schrieb zuletzt editiert von kameecoding@lemmy.world

#97

Lmao, okay buddy, based on how many interviews I have sat on in, the chances that you are a worse programmer than me are much higher than you being better than me.

Being a pompous ass dismissive of new tooling makes you chances even worse
P N 2 Antworten Letzte Antwort

2
M melvin_ferd@lemmy.world

Ok what about tech journalists who produced articles with those misunderstandings. Surely they know better yet still produce articles like this. But also people who care enough about this topic to post these articles usually I assume know better yet still spread this crap
S This user is from outside of this forum
S This user is from outside of this forum
suburban_hillbilly@lemmy.ml

schrieb zuletzt editiert von

#98

Gell-Mann amnesia effect - Wikipedia

(en.m.wikipedia.org)
T M 2 Antworten Letzte Antwort

5
K kameecoding@lemmy.world

For me as a software developer the accuracy is more in the 95%+ range.

On one hand the built in copilot chat widget in Intellij basically replaces a lot my google queries.

On the other hand it is rather fucking good at executing some rewrites that is a fucking chore to do manually, but can easily be done by copilot.

Imagine you have a script that initializes your DB with some test data. You have an Insert into statement with lots of columns and rows so

Inser into (column1,....,column n)
Values row1,
Row 2
Row n

Addig a new column with test data for each row is a PITA, but copilot handles it without issue.

Similarly when writing unit tests you do a lot of edge case testing which is a bunch of almost same looking tests with maybe one variable changing, at most you write one of those tests, then copilot will auto generate the rest after you name the next unit test, pretty good at guessing what you want to do in that test, at least with my naming scheme.

So yeah, it's way overrated for many-many things, but for programming it's a pretty awesome productivity tool.
D This user is from outside of this forum
D This user is from outside of this forum
dahgangalang@infosec.pub

schrieb zuletzt editiert von

#99

Yeah, it (in my case, ChatGPT) has been great for helping me along with functions I'm only passingly familiar with / trying to use in new ways.

One that I was really surprised with was that it gave me a surprisingly robust, sensible, and (seemingly) well tuned-to-my-case check list of things to inspect for a used car I intend to buy. I'm already mostly familiar with what I'm doing there, but it pointed to some things I might've overlooked / didn't know were points of concern for the specific vehicle I'm looking at.
Z 1 Antwort Letzte Antwort

1
E eli001@lemmy.world

This post did not contain any content.
A This user is from outside of this forum
A This user is from outside of this forum
apeno1@lemmy.world

schrieb zuletzt editiert von

#100

They've done studies, you know. 30% of the time, it works every time.
M 1 Antwort Letzte Antwort

7
E eli001@lemmy.world

This post did not contain any content.
B This user is from outside of this forum
B This user is from outside of this forum
burgerpocalyse@lemmy.world

schrieb zuletzt editiert von

#101

I dont know why but I am reminded of this clip about eggless omelette https://youtu.be/9Ah4tW-k8Ao
1 Antwort Letzte Antwort

2
O outbound7404@lemmy.ml

A human can review something close to correct a lot better than starting the task from zero.
D This user is from outside of this forum
D This user is from outside of this forum
dreamlandlividity@lemmy.world

schrieb zuletzt editiert von

#102

It is a lot harder to notice incorrect information in review, than making sure it is correct when writing it.
L M 2 Antworten Letzte Antwort

4
K kameecoding@lemmy.world

Lmao, okay buddy, based on how many interviews I have sat on in, the chances that you are a worse programmer than me are much higher than you being better than me.

Being a pompous ass dismissive of new tooling makes you chances even worse
P This user is from outside of this forum
P This user is from outside of this forum
potentialproblem@sh.itjust.works

schrieb zuletzt editiert von

#103

I’ve been in the industry awhile and your assessment is dead on.

As long as you’re not blindly committing the code, it’s a huge time saver for a number of mundane tasks.

It’s especially fantastic for writing throwaway tooling. Need data massaged a specific way? Ez pz. Need a script to execute an api call on each entry in a spreadsheet? No problem.

The guy above you is a nutter. Not sure if people haven’t tried leveraging LLMs or what. It has a ton of faults, but it really does speed up the mundane work. Also, clearly the person is either brand new to the field or doesn’t even work in it. Otherwise they would have seen the barely functional shite that actual humans churn out.

Part of me wonders if code organization is going to start optimizing for interpretation by these models rather than humans.
Z 1 Antwort Letzte Antwort

1
K kameecoding@lemmy.world

Lmao, okay buddy, based on how many interviews I have sat on in, the chances that you are a worse programmer than me are much higher than you being better than me.

Being a pompous ass dismissive of new tooling makes you chances even worse
N This user is from outside of this forum
N This user is from outside of this forum
nalivai@discuss.tchncs.de

schrieb zuletzt editiert von

#104

The person who uses fancy autocomplete to write their code will be exactly the person who thinks they're better than everyone. Those traits are correlated.
K 1 Antwort Letzte Antwort

2
D dahgangalang@infosec.pub

Yeah, it (in my case, ChatGPT) has been great for helping me along with functions I'm only passingly familiar with / trying to use in new ways.

One that I was really surprised with was that it gave me a surprisingly robust, sensible, and (seemingly) well tuned-to-my-case check list of things to inspect for a used car I intend to buy. I'm already mostly familiar with what I'm doing there, but it pointed to some things I might've overlooked / didn't know were points of concern for the specific vehicle I'm looking at.
Z This user is from outside of this forum
Z This user is from outside of this forum
zbyte64@awful.systems

schrieb zuletzt editiert von

#105

Pepper Ridge Farms remembers when you could just do a web search and get it answered in the first couple results. Then the SEO wars happened....
1 Antwort Letzte Antwort

1
P potentialproblem@sh.itjust.works

I’ve been in the industry awhile and your assessment is dead on.

As long as you’re not blindly committing the code, it’s a huge time saver for a number of mundane tasks.

It’s especially fantastic for writing throwaway tooling. Need data massaged a specific way? Ez pz. Need a script to execute an api call on each entry in a spreadsheet? No problem.

The guy above you is a nutter. Not sure if people haven’t tried leveraging LLMs or what. It has a ton of faults, but it really does speed up the mundane work. Also, clearly the person is either brand new to the field or doesn’t even work in it. Otherwise they would have seen the barely functional shite that actual humans churn out.

Part of me wonders if code organization is going to start optimizing for interpretation by these models rather than humans.
Z This user is from outside of this forum
Z This user is from outside of this forum
zbyte64@awful.systems

schrieb zuletzt editiert von

#106

When LLMs get it right it's because they're summarizing a stack overflow or GitHub snippet it was trained on. But you loose all the benefits of other humans commenting on the context, pitfalls and other alternatives.
H P 2 Antworten Letzte Antwort

1
J jsomae@lemmy.ml

yes, that's generally useless. It should not be shoved down people's throats. 30% accuracy still has its uses, especially if the result can be programmatically verified.
K This user is from outside of this forum
K This user is from outside of this forum
knock_knock_lemmy_in@lemmy.world

schrieb zuletzt editiert von

#107

Run something with a 70% failure rate 10x and you get to a cumulative 98% pass rate.
LLMs don't get tired and they can be run in parallel.
M 1 Antwort Letzte Antwort

1
T tankovayadiviziya@lemmy.world

At least AI won't fire you.
Z This user is from outside of this forum
Z This user is from outside of this forum
zbyte64@awful.systems

schrieb zuletzt editiert von

#108

DOGE has entered the chat
1 Antwort Letzte Antwort

3
A affidavit@lemmy.world

"...for multi-step tasks"
L This user is from outside of this forum
L This user is from outside of this forum
loonsun@sh.itjust.works

schrieb zuletzt editiert von

#109

It's about Agents, which implies multi step as those are meant to execute a series of tasks opposed to studies looking at base LLM model performance.
1 Antwort Letzte Antwort

3

Anmelden zum Antworten

P

Simple Wikiclaudia: Chrome extension that finds a simple.wikipedia.org version of any wiki article. If one exists, click to open it; otherwise, it uses Claude or ChatGPT to simplify it.
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
5

1

15 Stimmen

5 Beiträge

0 Aufrufe

A

If i have to rely on ai to read fucking wikipedia of all things then shoot me
N

Perovskite-based image sensors promise higher sensitivity and resolution than silicon
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
23

1

112 Stimmen

23 Beiträge

92 Aufrufe

E

I mean no more live view via the screen
P

xAI Data Center Emits Plumes of Pollution, New Video Shows
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
50

1

516 Stimmen

50 Beiträge

186 Aufrufe

G

You do. But you also plan in the case the surrounding infrastructure fails. But more to the point, in some cases it is better to produce (parto of) your own electricity (where better means cheaper) than buy it on the market. It is not really common but is doable.
T

Apple announces iOS 26 with Liquid Glass redesign
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
83

1

118 Stimmen

83 Beiträge

223 Aufrufe

S

you guys are weird
A

Apple just proved AI "reasoning" models like Claude, DeepSeek-R1, and o3-mini don't actually reason at all. They just memorize patterns really well.
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
356

880 Stimmen

356 Beiträge

387 Aufrufe

C

Is that useful for completing tasks?
P

Germany's Federal Cartel Office warns Amazon that its marketplace retailer price controls likely violate national and EU laws, in its preliminary assessment
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
5

1

57 Stimmen

5 Beiträge

6 Aufrufe

B

Amazon is an absolute scumbag company, they don't pay taxes and they shit all over their workers, and fight unions tooth and nail. I have no idea how people can buy at Amazon, that stands for everything Trump and Musk stands for. Just fucking stop using Amazon if you value democracy. Pay an extra dollar and buy somewhere else.
B

Is it feasible and scalable to combine self-replicating automata (after von Neumann) with federated learning and the social web?
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
9

1

6 Stimmen

9 Beiträge

6 Aufrufe

B

Cool. Well, the feedback until now was rather lukewarm. But that's fine, I'm now going more in a P2P-direction. It would be cool to have a way for everybody to participate in the training of big AI models in case HuggingFace enshittifies
C

Brian Eno: “The biggest problem about AI is not intrinsic to AI. It’s to do with the fact that it’s owned by the same few people”
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

1

0 Stimmen

1 Beiträge

7 Aufrufe

Niemand hat geantwortet