linux-nerds.org

Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.

Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).

Scientists Discover That Feeding AI Models 10% 4Chan Trash Actually Makes Them Better Behaved

Technology

133 Beiträge 88 Kommentatoren 7 Aufrufe

R reverendender@sh.itjust.works

I know everyone on Lemmy hates LLMs, but this is really interesting
T This user is from outside of this forum
T This user is from outside of this forum
timeworntraveler@lemm.ee

schrieb zuletzt editiert von timeworntraveler@lemm.ee

#104

I do hate LLMs (or how they're marketed/hyped/used) and I concur that this is very interesting science
R 1 Antwort Letzte Antwort

3
E echosnail@lemmy.zip

You say all this until ChatGpt convinced you to write a manifesto to "take back" your foreskin from the Jews.
L This user is from outside of this forum
L This user is from outside of this forum
lka1988@lemmy.dbzer0.com

schrieb zuletzt editiert von

#105

Funny enough, I am circumcised. But no, if I wanted it back that badly, I'd write it myself.
1 Antwort Letzte Antwort

1
A anaveragesnoot@lemmy.ca

I don't dislike LLMs, I dislike people who treat them as anything more than an advanced search engine and stupidly give them all their confidential data. Seen it happen too much at work.
I This user is from outside of this forum
I This user is from outside of this forum
ipkpjersi@lemmy.ml

schrieb zuletzt editiert von

#106

Yep. My work is very strict about security except for when it comes to LLMs, and then suddenly they're surprisingly lax about it. It's a bit concerning actually.
1 Antwort Letzte Antwort

0
T timeworntraveler@lemm.ee

I do hate LLMs (or how they're marketed/hyped/used) and I concur that this is very interesting science
R This user is from outside of this forum
R This user is from outside of this forum
reverendender@sh.itjust.works

schrieb zuletzt editiert von

#107

I appreciate your reasoned and measured reply, friend!
1 Antwort Letzte Antwort

0
_ _thebrain_@sh.itjust.works

Underrated comment.
F This user is from outside of this forum
F This user is from outside of this forum
feathercrown@lemmy.world

schrieb zuletzt editiert von

#108

Seems pretty rated to me
1 Antwort Letzte Antwort

4
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
M This user is from outside of this forum
M This user is from outside of this forum
mangionedontmiss@lemmy.ca

schrieb zuletzt editiert von

#109

goddamn, has 4chan gone so far down the road that its actually come back around and become the good guy?
1 Antwort Letzte Antwort

0
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
M This user is from outside of this forum
M This user is from outside of this forum
mr_dr_oink@lemmy.world

schrieb zuletzt editiert von

#110

So is it saying essentially that in order to not output garbage, it needs to know first what garbage is?

Is it just me that things this seems like a no-brainer?

It almosr draws parallels to many societal issues. Knowledge is power.

People tend towards intolerance and hatred when they dont understand the thing they are angry at. The more they know the better they behave.
H M 2 Antworten Letzte Antwort

16
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
M This user is from outside of this forum
M This user is from outside of this forum
markovs_gun@lemmy.world

schrieb zuletzt editiert von

#111

This is not surprising if you've studied anything on machine learning or even just basic statistics. Consider if you are trying to find out the optimal amount of a thickener to add to a paint formulation to get it to flow the amount you want. If you add it at 5%, then 5.1%, then 5.2%, it will he hard to see how much of the difference between those batches is due to randomness or measurement uncertainty than if you see what it does at 0%, then 25% then 50%. This is a principle called Design of Experiments (DoE) in traditional statistics, and a similar effect happens when you are training machine learning models- datapoints far outside the norm increase the ability of the model to predict within the entire model space (there is some nuance here, because they can become over-represented if care isn't taken). In this case, 4chan shows the edges of the English language and human psychology, like adding 0% or 50% of the paint additives rather than staying around 5%.

At least that's my theory. I haven't read the paper but plan to read it tonight when I have time. At first glance I'm not surprised. When I've worked with industrial ML applications, processes that have a lot of problems produce better training data than well controlled processes, and I have read papers on this subject where people have improved performance of their models by introducing (controlled) randomness into their control setpoints to get more training data outside of the tight control regime.
M 1 Antwort Letzte Antwort

0
G grimy@lemmy.world

Those are actually some very good results. Funny situation, if the copyright companies win the AI legislative war, 4chan is going to get twice as much as reddit did for the data at the minimum.

It's also interesting the model gets worse faster if it has to untrain the toxic data so to speak.
A This user is from outside of this forum
A This user is from outside of this forum
aeonfelis@lemmy.world

schrieb zuletzt editiert von

#112

So basically... by being familiar with 4chan the model knows better what not to do?
G 1 Antwort Letzte Antwort

3
C cosmonova@lemmy.world

And I wish they would tone down the hype. Maybe we can meet in the middle?
S This user is from outside of this forum
S This user is from outside of this forum
sculptuspoe@lemmy.world

schrieb zuletzt editiert von

#113

Well, I do wish they would promote the actual use and limitations of AI and stop making up crap and overselling the use cases. I use ChatGPT at work all the time as a start for research, but if I took any of it as being reliable info to run with I would be in grave trouble. It is a great tool that has saved me much time because I know how far to trust it and how to use it. The progress is very impressive as I've been using AI art services for years, and the difference between the random blobs from back then and the great stuff it can generate now is pretty stark. Same thing with the LLMs. I've been using ChatGPT since it showed up and it has improved greatly since then. Before all this I talked to people who were using AI training on various picture recognition projects where getting data from other sensors was not practical. ... Overall AI is pretty exciting, but the non-stop hype and hate headlines is doing nobody any favors.
1 Antwort Letzte Antwort

1
T taladar@sh.itjust.works

As a standalone thing, LLMs are awesome.

They really aren't though and that is half the problem. Everyone pretends they are awesome when the results are unusable garbage 80% of the time which makes them unusable for 99% of practical applications.
E This user is from outside of this forum
E This user is from outside of this forum
elbarto777@lemmy.world

schrieb zuletzt editiert von

#114

That's why I said "as standalone things." As a computing curiosity, they're amazing. No language processing application like this existed 30 years ago when I was a kid. You could also see "talking computers" speaking naturally, pretending or not, on movies and TV shows.
1 Antwort Letzte Antwort

0
T taladar@sh.itjust.works

There are plenty of tasks which they solve perfectly, today.

Name a single task you would trust an LLM on solving for you that you feel confident would be correct without checking the output. Because that is my definition of perfectly and AI falls very, very far short of that.
E This user is from outside of this forum
E This user is from outside of this forum
elbarto777@lemmy.world

schrieb zuletzt editiert von

#115

"Hey AI, write me a random poem about taladar."
1 Antwort Letzte Antwort

0
Y yournamehere@lemm.ee

because 4chan users write original content. that is fed into the next best stupid platform and so on until it ends on tiktok or whatever.

if you have nothing to say you use meta/tiktok. no relevabt content has ever been there first.
copies and derivates, yes...

so soonish AI will flood 4chan so ai scrapers get polluted aswell...and then it is dead.
S This user is from outside of this forum
S This user is from outside of this forum
sparrohawc@lemmy.zip

schrieb zuletzt editiert von

#116

It has nothing to do with that, and much more to do with people on 4chan being willing to call each other out. Without toxic behavior you can't have examples on how to deal with toxic behavior.
1 Antwort Letzte Antwort

4
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
J This user is from outside of this forum
J This user is from outside of this forum
jsomae@lemmy.ml

schrieb zuletzt editiert von jsomae@lemmy.ml

#117

Headlines should not say "scientists," they should name the institution. (Harvard in this case.)
U 1 Antwort Letzte Antwort

19
L l0rdmathias@sh.itjust.works

I recently realized it's a non-issue. The people doing this have already been looking for decades to find new ways to rot their minds. LLMs are just the latest in a long line of tools that help them tune out.
S This user is from outside of this forum
S This user is from outside of this forum
sparrohawc@lemmy.zip

schrieb zuletzt editiert von sparrohawc@lemmy.zip

#118

The problem is that before LLMs, they had to actually put forward some effort to produce content on the internet, which at least kept the amount of thoughtless content down somewhat. Now the barrier to entry is practically zero, all while thieving people's hard work without compensation and burning ridiculous amounts of resources to do so.

It is super interesting tech though.
1 Antwort Letzte Antwort

2
A aeonfelis@lemmy.world

So basically... by being familiar with 4chan the model knows better what not to do?
G This user is from outside of this forum
G This user is from outside of this forum
grimy@lemmy.world

schrieb zuletzt editiert von

#119

Yup. Sucks for everyone having fun jailbreaking them. It is going to get much harder.
1 Antwort Letzte Antwort

1
M mr_dr_oink@lemmy.world

So is it saying essentially that in order to not output garbage, it needs to know first what garbage is?

Is it just me that things this seems like a no-brainer?

It almosr draws parallels to many societal issues. Knowledge is power.

People tend towards intolerance and hatred when they dont understand the thing they are angry at. The more they know the better they behave.
H This user is from outside of this forum
H This user is from outside of this forum
halowpeano@lemmy.world

schrieb zuletzt editiert von

#120

No it's more of a technical discussion.
Many people might believe that in order to avoid toxicity, you just train a model on "good" non-toxic data and then apply toxicity removal techniques to address emergent toxicity that the model might spit out.
This paper is saying they found it more effective to train the model on a small percentage of "bad" toxic data on purpose, then apply those same toxicity removal techniques. For some reason, that actually generated less total toxicity.
It's an interesting result. A wild guess on my part, but I'm thinking training the model with toxic content "sharpened" the toxicity when it was generated, making it easier for those removal tools to identify it.
M 1 Antwort Letzte Antwort

6
M mr_dr_oink@lemmy.world

So is it saying essentially that in order to not output garbage, it needs to know first what garbage is?

Is it just me that things this seems like a no-brainer?

It almosr draws parallels to many societal issues. Knowledge is power.

People tend towards intolerance and hatred when they dont understand the thing they are angry at. The more they know the better they behave.
M This user is from outside of this forum
M This user is from outside of this forum
mangocats@feddit.it

schrieb zuletzt editiert von

#121

Is it just me that things this seems like a no-brainer?

Yes, and no. When raising our children, my wife prefers the "ban the bad stuff" approach. I don't encourage exposure to bad stuff, but when my kid wants to buy and watch a raunchy movie, instead of yelling "NO!" and making him put it back, I let him buy it and we watch it, together, pausing to explain the unrealistic and awful parts and explain how imitating these things in real life can cause problems for you.
1 Antwort Letzte Antwort

3
H halowpeano@lemmy.world

No it's more of a technical discussion.
Many people might believe that in order to avoid toxicity, you just train a model on "good" non-toxic data and then apply toxicity removal techniques to address emergent toxicity that the model might spit out.
This paper is saying they found it more effective to train the model on a small percentage of "bad" toxic data on purpose, then apply those same toxicity removal techniques. For some reason, that actually generated less total toxicity.
It's an interesting result. A wild guess on my part, but I'm thinking training the model with toxic content "sharpened" the toxicity when it was generated, making it easier for those removal tools to identify it.
M This user is from outside of this forum
M This user is from outside of this forum
mangocats@feddit.it

schrieb zuletzt editiert von

#122

Toxicity is everywhere, you can't recognize that "Drill baby drill" has sexual connotations if you've never been exposed to sexual double entendre like that before.
1 Antwort Letzte Antwort

0
G gonzako@lemmy.world

yeah, this only works in scientific fields
M This user is from outside of this forum
M This user is from outside of this forum
mangocats@feddit.it

schrieb zuletzt editiert von

#123

And it rarely works in scientific fields right away - usually an established wrong idea needs to be overwhelmed with serious proof before scientists start to consider that what they "know" might be wrong.
1 Antwort Letzte Antwort

1

Anmelden zum Antworten

P

Welcome to the web we lost
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
22

1

181 Stimmen

22 Beiträge

4 Aufrufe

C

Is it though? Its always far easier to be loud and obnoxious than do something constructive, even with the internet and LLMs, in fact those things are amplifiers which if anything make the attention imbalance even more drastic and unrepresentative of actual human behaviour. In the time it takes me to write this comment some troll can write a dozen hateful ones, or a bot can write a thousand. Doesn't mean humans are shitty in a 1000/1 ratio, just means shitty people can now be a thousand times louder.
R

Reddit sues Anthropic for allegedly not paying for training data | TechCrunch
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
16

1

135 Stimmen

16 Beiträge

5 Aufrufe

E

I thought we were going to get our share of the damages
D

The IRS Tax Filing Software TurboTax Is Trying to Kill Just Got Open Sourced
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
145

1

1k Stimmen

145 Beiträge

10 Aufrufe

P

Not just that. The tax preparation industry has gotten tax more complex and harder to file in the US You get the government you can afford. The tax preparation industry has been able to buy several governments
S

DeepSeek's distilled new R1 AI model can run on a single GPU | TechCrunch
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
21

1

88 Stimmen

21 Beiträge

2 Aufrufe

J

The self hosted model has hard coded censored content.
T

Telegram partners with xAI to bring Grok to over a billion users
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
36

1

38 Stimmen

36 Beiträge

2 Aufrufe

R

So you pay taxes to Putin. Good to know who actually helps funding the regime. I suggest you go someplace else. I won't take this from a jerk from likely one of the countries buying fossil fuels from said regime, that have also supported it after a few falsified elections starting in 1996, which is also the year I was born. And of course "paying taxes to Putin" can't be even compared to what TG is doing, so just shut up and go do something you know how to do, like I dunno what.
V

Microsoft wants Windows Update to handle all apps
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
45

1

61 Stimmen

45 Beiträge

2 Aufrufe

N

the package managers for linux that i know of are great because you can easily control everything they do
P

Is Washington state falling out of love with Tesla?
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
10

1

61 Stimmen

10 Beiträge

2 Aufrufe

B

These Tesla owners who love their cars but hate his involvement with government are a bit ridiculous because one of the biggest reasons he got in loved with shilling for the right is that the government was looking into regulations and investigations concerning how unsafe Tesla cars are.
A

Telegram bans $35B black markets used to sell stolen data, launder crypto
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
8

1

1 Stimmen

8 Beiträge

2 Aufrufe

L

I made a PayPal account like 20 years ago in a third world country. The only thing you needed then is an email and password. I have no real name on there and no PII, technically my bank card is attached but on PP itself there's no KYC. I think you could probably use some types of prepaid cards with it if you want to avoid using a bank altogether but for me this wasn't an issue, I just didn't want my ID on any records, I don't have any serious OpSec concerns otherwise. I'm sure you could either buy PayPal accounts like this if you needed to, or make one in a country that doesn't have KYC laws somehow. From there I'd add money to my balance and send money as F&F. At no point did I need an ID so in that sense there's no KYC. Some sellers on localmarket were fancy enough to list that they wanted an ID for KYC, but I'm sure you could just send them any random ID you made in paint from the republic of dave and you'd be fine.