linux-nerds.org

Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.

Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).

Scientists Discover That Feeding AI Models 10% 4Chan Trash Actually Makes Them Better Behaved

Technology

133 Beiträge 88 Kommentatoren 3.2k Aufrufe

T tungsten5@lemm.ee

Give the AI model the gift of culture and class. No suprise it behaves better
E This user is from outside of this forum
E This user is from outside of this forum
echosnail@lemmy.zip

schrieb am zuletzt editiert von

#94

Sophistication my good sir.
1 Antwort Letzte Antwort

10
L lka1988@lemmy.dbzer0.com

This is one instance where I'm ok with the occasional beating. It's a computer. It doesn't have feelings. It never will. It's not sentient.
E This user is from outside of this forum
E This user is from outside of this forum
echosnail@lemmy.zip

schrieb am zuletzt editiert von

#95

You say all this until ChatGpt convinced you to write a manifesto to "take back" your foreskin from the Jews.
L 1 Antwort Letzte Antwort

3
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
N This user is from outside of this forum
N This user is from outside of this forum
naevermix@lemmy.world

schrieb am zuletzt editiert von naevermix@lemmy.world

#96

I envision a Gemini powered bot that cracks captcha and posts "woke" replies on 4chan. If you're an antivaxxer, antisemite, nazi, racist, sionist, or otherwise, it will debate you. It will not get tired. It will not get mad. It will maintain a sense of decorum indefinitely and it will never ever stop. If some far right extremist decides to do the same, it will have the advantage that academia is left leaning, meaning the model can cite widely recognized studies.

Dead internet theory and so on, but I'll gladly completely and utterly destroy the internet if it means the filth dies with it.
D P 2 Antworten Letzte Antwort

10
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
G This user is from outside of this forum
G This user is from outside of this forum
goodboyjojo@lemm.ee

schrieb am zuletzt editiert von

#97

Based and hopepilled
1 Antwort Letzte Antwort

1
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
C This user is from outside of this forum
C This user is from outside of this forum
cupcakezealot@lemmy.blahaj.zone

schrieb am zuletzt editiert von

#98

can we stop referring to llm's as if they're capable of thought? they don't make decisions; their programming just responds to patterns.
M 1 Antwort Letzte Antwort

9
N naevermix@lemmy.world

I envision a Gemini powered bot that cracks captcha and posts "woke" replies on 4chan. If you're an antivaxxer, antisemite, nazi, racist, sionist, or otherwise, it will debate you. It will not get tired. It will not get mad. It will maintain a sense of decorum indefinitely and it will never ever stop. If some far right extremist decides to do the same, it will have the advantage that academia is left leaning, meaning the model can cite widely recognized studies.

Dead internet theory and so on, but I'll gladly completely and utterly destroy the internet if it means the filth dies with it.
D This user is from outside of this forum
D This user is from outside of this forum
disaster@sh.itjust.works

schrieb am zuletzt editiert von

#99

There's little evidence that debate changes people's ideas.
N G P 3 Antworten Letzte Antwort

7
D disaster@sh.itjust.works

There's little evidence that debate changes people's ideas.
N This user is from outside of this forum
N This user is from outside of this forum
naevermix@lemmy.world

schrieb am zuletzt editiert von

#100

It's not about changing their ideas. The target is the audience.
1 Antwort Letzte Antwort

2
N naevermix@lemmy.world

I envision a Gemini powered bot that cracks captcha and posts "woke" replies on 4chan. If you're an antivaxxer, antisemite, nazi, racist, sionist, or otherwise, it will debate you. It will not get tired. It will not get mad. It will maintain a sense of decorum indefinitely and it will never ever stop. If some far right extremist decides to do the same, it will have the advantage that academia is left leaning, meaning the model can cite widely recognized studies.

Dead internet theory and so on, but I'll gladly completely and utterly destroy the internet if it means the filth dies with it.
P This user is from outside of this forum
P This user is from outside of this forum
pushbutton@lemmy.world

schrieb am zuletzt editiert von

#101

it will have the advantage that academia is left leaning, meaning the model can cite widely recognized studies.

I was looking for the person saying a particular quote yesterday.

I asked 3 times the same question and I got 3 different people.

The funny part us I had the quote wrong.

Bullshit all the way down.
1 Antwort Letzte Antwort

1
D disaster@sh.itjust.works

There's little evidence that debate changes people's ideas.
G This user is from outside of this forum
G This user is from outside of this forum
gonzako@lemmy.world

schrieb am zuletzt editiert von

#102

yeah, this only works in scientific fields
M 1 Antwort Letzte Antwort

1
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
Y This user is from outside of this forum
Y This user is from outside of this forum
yournamehere@lemm.ee

schrieb am zuletzt editiert von

#103

because 4chan users write original content. that is fed into the next best stupid platform and so on until it ends on tiktok or whatever.

if you have nothing to say you use meta/tiktok. no relevabt content has ever been there first.
copies and derivates, yes...

so soonish AI will flood 4chan so ai scrapers get polluted aswell...and then it is dead.
S 1 Antwort Letzte Antwort

9
R reverendender@sh.itjust.works

I know everyone on Lemmy hates LLMs, but this is really interesting
T This user is from outside of this forum
T This user is from outside of this forum
timeworntraveler@lemm.ee

schrieb am zuletzt editiert von timeworntraveler@lemm.ee

#104

I do hate LLMs (or how they're marketed/hyped/used) and I concur that this is very interesting science
R 1 Antwort Letzte Antwort

3
E echosnail@lemmy.zip

You say all this until ChatGpt convinced you to write a manifesto to "take back" your foreskin from the Jews.
L This user is from outside of this forum
L This user is from outside of this forum
lka1988@lemmy.dbzer0.com

schrieb am zuletzt editiert von

#105

Funny enough, I am circumcised. But no, if I wanted it back that badly, I'd write it myself.
1 Antwort Letzte Antwort

1
A anaveragesnoot@lemmy.ca

I don't dislike LLMs, I dislike people who treat them as anything more than an advanced search engine and stupidly give them all their confidential data. Seen it happen too much at work.
I This user is from outside of this forum
I This user is from outside of this forum
ipkpjersi@lemmy.ml

schrieb am zuletzt editiert von

#106

Yep. My work is very strict about security except for when it comes to LLMs, and then suddenly they're surprisingly lax about it. It's a bit concerning actually.
1 Antwort Letzte Antwort

0
T timeworntraveler@lemm.ee

I do hate LLMs (or how they're marketed/hyped/used) and I concur that this is very interesting science
R This user is from outside of this forum
R This user is from outside of this forum
reverendender@sh.itjust.works

schrieb am zuletzt editiert von

#107

I appreciate your reasoned and measured reply, friend!
1 Antwort Letzte Antwort

0
_ _thebrain_@sh.itjust.works

Underrated comment.
F This user is from outside of this forum
F This user is from outside of this forum
feathercrown@lemmy.world

schrieb am zuletzt editiert von

#108

Seems pretty rated to me
1 Antwort Letzte Antwort

4
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
M This user is from outside of this forum
M This user is from outside of this forum
mangionedontmiss@lemmy.ca

schrieb am zuletzt editiert von

#109

goddamn, has 4chan gone so far down the road that its actually come back around and become the good guy?
1 Antwort Letzte Antwort

0
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
M This user is from outside of this forum
M This user is from outside of this forum
mr_dr_oink@lemmy.world

schrieb am zuletzt editiert von

#110

So is it saying essentially that in order to not output garbage, it needs to know first what garbage is?

Is it just me that things this seems like a no-brainer?

It almosr draws parallels to many societal issues. Knowledge is power.

People tend towards intolerance and hatred when they dont understand the thing they are angry at. The more they know the better they behave.
H M 2 Antworten Letzte Antwort

16
P pro@programming.dev
- HTML.
- PDF.
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.
M This user is from outside of this forum
M This user is from outside of this forum
markovs_gun@lemmy.world

schrieb am zuletzt editiert von

#111

This is not surprising if you've studied anything on machine learning or even just basic statistics. Consider if you are trying to find out the optimal amount of a thickener to add to a paint formulation to get it to flow the amount you want. If you add it at 5%, then 5.1%, then 5.2%, it will he hard to see how much of the difference between those batches is due to randomness or measurement uncertainty than if you see what it does at 0%, then 25% then 50%. This is a principle called Design of Experiments (DoE) in traditional statistics, and a similar effect happens when you are training machine learning models- datapoints far outside the norm increase the ability of the model to predict within the entire model space (there is some nuance here, because they can become over-represented if care isn't taken). In this case, 4chan shows the edges of the English language and human psychology, like adding 0% or 50% of the paint additives rather than staying around 5%.

At least that's my theory. I haven't read the paper but plan to read it tonight when I have time. At first glance I'm not surprised. When I've worked with industrial ML applications, processes that have a lot of problems produce better training data than well controlled processes, and I have read papers on this subject where people have improved performance of their models by introducing (controlled) randomness into their control setpoints to get more training data outside of the tight control regime.
M 1 Antwort Letzte Antwort

0
G grimy@lemmy.world

Those are actually some very good results. Funny situation, if the copyright companies win the AI legislative war, 4chan is going to get twice as much as reddit did for the data at the minimum.

It's also interesting the model gets worse faster if it has to untrain the toxic data so to speak.
A This user is from outside of this forum
A This user is from outside of this forum
aeonfelis@lemmy.world

schrieb am zuletzt editiert von

#112

So basically... by being familiar with 4chan the model knows better what not to do?
G 1 Antwort Letzte Antwort

3
C cosmonova@lemmy.world

And I wish they would tone down the hype. Maybe we can meet in the middle?
S This user is from outside of this forum
S This user is from outside of this forum
sculptuspoe@lemmy.world

schrieb am zuletzt editiert von

#113

Well, I do wish they would promote the actual use and limitations of AI and stop making up crap and overselling the use cases. I use ChatGPT at work all the time as a start for research, but if I took any of it as being reliable info to run with I would be in grave trouble. It is a great tool that has saved me much time because I know how far to trust it and how to use it. The progress is very impressive as I've been using AI art services for years, and the difference between the random blobs from back then and the great stuff it can generate now is pretty stark. Same thing with the LLMs. I've been using ChatGPT since it showed up and it has improved greatly since then. Before all this I talked to people who were using AI training on various picture recognition projects where getting data from other sensors was not practical. ... Overall AI is pretty exciting, but the non-stop hype and hate headlines is doing nobody any favors.
1 Antwort Letzte Antwort

1

Anmelden zum Antworten

M

WeTransfer updates T&Cs, allows it to use your data to train AI
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
21

1

240 Stimmen

21 Beiträge

298 Aufrufe

3

I'd love to say I believe them in their backing down statement - but being cynical I really don't
O

Learn About Climate Change with Stunning Visual Flashcards 🌍📚
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

0 Stimmen

1 Beiträge

19 Aufrufe

Niemand hat geantwortet
P

CursorAI "unlimited" plan rug pull: Cursor AI silently changed their "unlimited" Pro plan to severely rate-limited without notice, locking users out after 3-7 requests
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
31

295 Stimmen

31 Beiträge

411 Aufrufe

A

I have a rough idea of their efficiency as I've used them, not in professional settings but I wager it would not be too different. My point is more that it feels like the rugs are finally starting to get pulled. This tech is functionnal as you said, it works to a point and that point is enough for a sizeable amount of people. But I doubt that the price most people are paying now is enough to cover the cost of answering their queries. Now that some people, especially younger devs or people who never worked without those tools are dependant on it, they can go ahead and charge more. But it's not too late, so I'm hoping it will make some people more aware of that kind of scheme and that they will stop feeding the AI hype in general.
B

Why Ohio Trusts Baker Chiropractic for Arthritis Pain Relief
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
1

1

0 Stimmen

1 Beiträge

19 Aufrufe

Niemand hat geantwortet
P

Engineers Introduce Berkeley Humanoid Lite, Open-Source, Customizable, 3D-Printed Robot for Tech Newbies.
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
13

1

73 Stimmen

13 Beiträge

131 Aufrufe

F

They have a tiny version that is listed as 1000 on their website, plus the simulation is FOSS
M

YouTube rolls out more unskippable ads that make viewers wait even longer to watch videos - Dexerto
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
231

1

781 Stimmen

231 Beiträge

6k Aufrufe

D

Haha I'm kidding, it's good that you share your solution here.
A

The British jet engine that failed in the 'Valley of Death'
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
16

1

40 Stimmen

16 Beiträge

147 Aufrufe

R

Giving up advancements in science and technology is stagnation. That's not what I'm suggesting. I'm suggesting giving up some particular, potential advancements in science and tecnology, which is a whole different kettle of fish and does not imply stagnation. Thinking it’s a good idea to not do anything until people are fed and housed is stagnation. Why do you think that?
X

Child safety org launches AI model trained on real child sex abuse images
Beobachtet Ignoriert Geplant Angeheftet Gesperrt Verschoben Technology technology
2

1

0 Stimmen

2 Beiträge

32 Aufrufe

B

... robo chomo?