The shrinking landscape of linguistic diversity in the age of LLMs
96 points - last Sunday at 12:12 PM
SourceComments
i've been trying lately to think how to talk about my creeping fear of LLMs... here's my best effort at explaining the intuition:
the words are the map. we've built this map together over generations. Words are where we've cross-boned the dangers and x-marked the treasures, whether the things that we've found in the territory or that we've buried within it.
a few weeks ago, a strange new player arrived at the edge of the night's encampment, and it knows our secret society's handshake, and so we are compelled to invite it to join us. it seems to be a good navigator... and so now it's taken to sitting at the front of the convoy, pointing the way as it quietly redraws the map...
Kids from marginalized and affluent communities alike are now being raised on Mr. Beast. Mr Beast speak may become the universal dialect of the future.
Like and subscribe replaces goodbye. Who am I to judge. Mr Beast is articulate after all.
If you grow up in the hood like myself code switching ends up being a much needed skill.
In the future Mr Beast speak becomes a first language. Maybe dialects are outdated artifacts of the old times.
Like and subscribe.
This is fine at work, I guess. I'm not being asked to sound like myself. I'm being asked to convey an idea to an audience as efficiently as possible, both in terms of my time and theirs.
But I can't shake the feeling that the LLMs are eating my soul.
I still don't use them often to write code. Not that I have issue with those that do, but that sounding like myself in code is a part of my soul that I'm having difficulty giving up.
In most cases where language is being used as a vehicle for exchanging ideas, isn't some degree of uniformity actually desirable for clarity? I personally wouldn't want instruction manuals to be full of colorful or highly individualized language, for instance.
That's not to say linguistic diversity is meaningless, but I would argue that this particular concern is fairly specialized and matters much more in literature and other creative forms of writing.
The LLM is a big probabilistic statistical trick. It picks the next token based on certain words are simply “the best” because they are specific and well connected to other tokens. The is gives them a great overall cost function. (Basically a good score on “will it make sense in context” while also having specific meaning that makes it better than other options, unambiguous in common use and being a single token rather than several).
You can trim those tokens, but then you just get other tokens that are “the best” tokens (and you’re worse off because the output became less clear).
The cool thing is this seems to get worse the more powerful and accurate your model is, because it is picking technically / statistically perfect tokens, not tasteful ones.
I bet a large %age of the HN readership are consuming the tech stack in their second or third language. Germans blogging in English or Indians commenting in English. We even have a broadly shared tech culture with norms, in-jokes, taboos, etc. LLMs reflect that.
If we want AI to do better, we should pay attention to making sure the non-English (and non-US English) internet thrives. Personally I´d love to see LLMs doing the needful and writing in Indian English, or dinner-party-argument French, instead of the sanitized corporate pablum tone it uses in English.
Or perhaps more controversially, we let LLMs chew up the English internet and deliberately build an offline culture in local languages. Keep your AI slop and your em-dashes, I'm gonna publish a zine in Italian.
Maybe LLM replaces a lot of the slop human writing, but I don't generally read that either. Now instead of search results being a bunch of useless SEO, now it is just a bunch of useless AI result. So little has changed other than the signs of what to immediately reject.
If something isn't worth someone spending a bit of time to write, it isn't worth my time to read.
Might as well drop all the pretense of “bringing your whole self”.
And besides, there is only time enough for ONE Mozart in the world.
It's not lost on me that the lowbrow humour of some too strict translation and a throwaway sitcom joke people thought was funny back in 2007 could be a way to stave off linguistic homogeneity.
For the use of profanity the data in the paper seems to indicate that swearing is retained in the LLM rewritten examples. Profanity however is rather subjective could be anything from saying the more mundane "oh my god" and "damn" to the extreme "shit eating fucklungerer" or use of slurs, but it's unstated in the paper as to the severity of the profanity. It's the severity of it that's important though, as LLMs restrain themselves from using something like "shit eating fucklungerer" unless explicitly told to.
Look at the posts on this site. A massive portion of them are very obviously written by AI.
As AI gets better, I expect fewer people will do the toil of writing.
Down here in the colony of New Zealand, the influence of the colony of North America is very noticeable. Music is the most directly attributable, but different media affects different demographics.
I’m reminded of that wonderful quote from that movie I just saw. Now not only will we have one culture, we will have one bland style of writing and even thinking, the LLM way.