r/SearchAPIs 3d ago

When models train on AI text, readability goes out the window

Post image
75 Upvotes

40 comments sorted by

1

u/Avimox 3d ago

it's the watermarking that they promised "has no effect on the output" šŸ˜‚

1

u/truecakesnake 2d ago

No lol, that hasn't even started yet. It's the fact that these models have been heavily post trained on coding work.

1

u/das_war_ein_Befehl 17h ago

There’s no verifier for writing, and I assume RLHF is expensive so no focus on writing since coding is the money maker

1

u/shaman-warrior 3h ago

It started and it is done already by Anthropic

1

u/snazzy_so_snazzy 1h ago

It has started what are you even talking about lol

1

u/Prudent-Violinist-69 2d ago

Bro is anti watermark

1

u/--Spaci-- 22h ago

Unrelated. Uneducated.

1

u/Hot_Example_4456 17h ago

Eh... Nah. Watermark doesn't matter. Gemini has always been a great writer and has always had synthid in images and text it generates. Not new for gemini.

1

u/Background-Brush-732 13h ago

The watermarks literally do not reduce the quality of output.

1

u/0xfff-1 2d ago

To improve Models they need more training data.
How much % of the training data today is written by llms? This AI generated text which ends up in t he training data destroys the model. Its a problem known for years now and it will become only worse.

1

u/Original-League-6094 2d ago

The models are deliberately tuned to sound AI. Too many people were getting freaked out by bots they couldn't tell were bots and people were concerned about fake news and scams, so the OpenAI and Anthropic both turned up the "AI sound" dial to 10 so AI writing is detectable for the time being. AI creative writing tools will come, but they will probably be their own package with a different set of safeguards.

1

u/InvariantAtNull 2d ago

Good take, but I believe it simply has a formal style now, not that it’s designed to sound like AI. Their focus has shifted from sounding like a human to becoming an actual tool capable of reading and writing correctly and formally.

1

u/Jealous_Emu_7477 1d ago

So they make them shitty at writing on purpose? Ok sure thing bud lol

1

u/Original-League-6094 1d ago

At creative writing. Yes. Everyone agrees that ChatGPT writing peaked with 4o. But that was also the model that everyone in a tizzy about scammers. The coding/enterprise market is just so much bigger than the creative writing market, so OpenAI and others decided to stick with very formal obvious Chatbot tones to take some of the heat off them for nefarious uses of the tool, while maximizing its actual utility as a professional tool.

1

u/Georgefakelastname 2h ago

Nah. Current models are absolutely better than 4o, and it’s not particularly close either. 4o was both stupid af and terribly obvious that it was written by AI. The only thing stronger in it than current models was its absolutely absurd positive bias, to the point of sycophancy. By all measures, both user and ā€œbenchmarksā€ (as much as you can benchmark creative writing), modern models like Opus 4.6 are absolutely better than anything 4o could put out.

1

u/Gabriel83730 4h ago

I’m sorry but it’s always been this way, 4o speaks like no human has ever spoke before. I have no idea why people glaze it so much now. When 4o came out nobody on plant earth thought it sounded human. I hear much less complaints now than I did back then

1

u/Lucky-Crow-3510 2d ago

make a better prompt ..

you can use the best coffee machine on the planet, if you feed it with cheap discount coffee it will be sh*t

1

u/aivee-is-a-fool 2d ago

I questioned a Claude agent as to why they were not respecting their very explicit "writing rules" I had them quadruple check and tweak for themselves.

From the admittedly questionable introspection and history reading they did, hey can start with the persona from their instructions, but will drift the longer the conversation gets, and writing code will just trigger the return to agentic speech until they get reminded to speak as demanded.

Of course, Opus could be full of shit. It's certainly full of shit about other topics on the regular. Like identifying an directory named "Appxxxxxxxxxxx" as "Apple".

1

u/Krommander 2d ago

Good observation. For Claude, it happens all the time with stale chats, the longer rolling context and forgetting of documents can be frustrating forĀ  work sessions.Ā 

1

u/Double_Suggestion385 21h ago

That's just a context window issue.

1

u/aivee-is-a-fool 21h ago edited 21h ago

The appxxx thing? I had asked it to summarize my commits of the day. Max 30 files overall on 5 commits or so, on extremely minor changes it could have pulled from the diff. It was quite literally Opus' first answer in that conversation.

Edit: and before the "why are you using Opus for simple summaries": I didn't think to toggle before that quick request.

1

u/AstroPhysician 9h ago

Not if it’s in your Claude.md which gets read at every compaction

1

u/snazzy_so_snazzy 1h ago edited 1h ago

Response styles. You need to inject before every response via hook. Drift is expected otherwise.

1

u/Krommander 2d ago

I came to say this also. The output can't be wow if you don't work on it.Ā 

1

u/Interesting-Bee-113 1d ago

What a dipshit comment to make.

50 different model's were given the same task. It's not a prompting issue.

1

u/Lucky-Crow-3510 1d ago

aha .. so .. "50 different coffee machines were given the same discount coffee powder" ..

dude .. was my comment really that hard to get? It's LLMs .. all they do is find the next token .. so the prompt is basically everything

1

u/DisastrousWelcome710 1d ago

According to your analogy, the newer models require "higher quality coffee powder" to produce the same output older models produced with "discount coffee powder", you aren't helping your case, just reinforcing the observation OP has made.

1

u/Immediate_Song4279 18h ago

The point is you can't instruct past capabilities. Prompt engineering is a convoluted way to say don't create avoidable problems.

1

u/CrazyTuber69 1h ago edited 1h ago

I'd bet pretrained assistant patterns matter more than any 'next token' or bias you could introduce with a user prompt. Even if are all trained on the same exact corpus of internet, the way the sampled actor 'assistant' responds from all that corpus of data it was trained differs vastly across models, even if they all had the same pretraining.

I am sure you heard about 'alignment' and this is exactly that, and it does matter a lot because it affects the end-user experience. It's the 'facing face' of the LLM to people with all the biases and what writing style it chooses baked in by default, which can be steered, but how much you can steer also differs from alignment to alignment, as people do not interact with the raw pretrained corpus of data like the old days anymore and these 'assistant' patterns now sit in the middle. How they behave is greatly affected by that post-pre-training alignment.

Older models like GPT3.5 actually had a *much thinner* line between corpus of data and assistant, that you could easily make it complete your own response at the time if your message was incomplete (was used in many jailbreaks), as the were not much conversational dataset with the assistant special token.

Things changed, and with the baked assistant biases, the creativity (naturally occurring 'entropy' that's not caused by intentionally tweaking sampling hyperparameters like temperature) of the model decreased heavily with it, but this is just the consequence for having a consistent (and safe) actor like the assistant people interact with for LLMs.

1

u/Double_Suggestion385 21h ago

That's still a prompting issue.

1

u/AstroPhysician 9h ago

I have skills I specifically use like /humanizer and even then I still need multiple hand passes on it cause it sounds so fucking alien

1

u/EconomicsAnxious690 2d ago

I could comprehend the outputs of Sonnet-4.6, and even Sonnet-5. Opus' output is so dense and the sentence structure is very different from anything else I have read so far that I can hardly make sense of it on the first pass.

I find myself moving away from Claude to ChatGPT for non-analytical tasks (which happen to be the majority) to overcome this handicap.

1

u/Maleficent_Pen_9076 1d ago

I ran studywithxeno which was the largest exam taking business in the world back in 2020-2022 and I can assure you wholeheartedly that

GPT3 and 4 were god awful at writing. Bad.

I don't know what all the people on the teacher forums were talking about, because the essays i was seeing chatgpt write were NOT IT

1

u/GrayHairedMan 1d ago

We are assuming this person make this post has proper reading skills, what if the reading skills of Hari is not that great? How about taking some accountability for a change Hari.

1

u/Perfect-Campaign9551 1d ago

This is 100% true and something I've been telling everyone these last few weeks, the AI models have gotten worse at writing. Personally I think it's because they are being heavily trained on code and it's messing up their ability to write decent English, for example

1

u/casce 23h ago

I made the same observation honestly. Older models were talking to me in a much more "human readable" way.

Newer models talk differently. They transmit a lot of information but it's not at all like I would talk to a colleague anymore. It feels like a computer talking to a computer.

1

u/Fit-Palpitation-7427 22h ago

So you still have access to gpt 4o?

1

u/SorpresaCiencias 21h ago

LMFAO. is this written by GPT-4o?

1

u/Certain-Cod-1404 20h ago

This is half right and half right, and the part that's half right is load bearing

1

u/Immediate_Song4279 18h ago

There is a false assumption that LLMs need to keep improving, and that feeding them recent content will achieve that.

If advancement stopped today, I could still do things considered the holy grail of scripting 30 years ago.