How the words we teach English language learners changed
163 points - today at 3:41 PM
SourceComments
It shocked me how there is absolutely no "right" answer.
If you are teaching English for travel, then you're prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it's for understanding TV, it's a lot of words like "murder", etc. Depending on which TV shows you want to understand.
If it's for reading the newspaper, you don't ever need to know "bathroom", but you sure do need to know words like "congressman".
While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed. And the substitutes -- transcribed speech from TV, radio, podcasts, etc. -- is not the same context as the random stuff you say at home and during an average day.
I would blame inequality on this one. In a more unequal world tribalization is a survival strategy and language follows.
When you see everybody else as your equals then focusing on describing that individual person, instead of their group, makes more sense.
Economic inequality affects deeply how we think about others.
I used the examples of Latin to Spanish and English, or Old English to new English.
I say that to say, languages changing over time is to some people not actually something they believe happens. The facts are right there in front of you, but some people cannot have their minds changed no matter how much sense you make.
There are some databases but e.g. they are biased towards Wikipedia and web which makes some very obscure words at the top of popularity (like some technical words which are present on each wiki page like Datenschutz or Impressum).
Having said that, the categories that shrank all did so by a big enough percentage to also shrink in absolute number of words, so at least that isn't a problem.
I'm not going to finger-scroll or down-arrow the whole thing.
Don't break scrolling. Please.
This is something I've been thinking a lot about. We have trended from subjective language to objective language. Why?
Computing. Software is written with objective language. Everything is clearly unambiguously defined. Blue is no longer a category, it's #0000FF. Logic must always reduce to a binary truth value. Most of what we have to talk about is somehow relative to software. Software even structures most of what we write! We don't just talk to each other, we tweet, email, message, post, search, etc. These structures each imply a specific set of phrase structures that can make sense.
Lately, it's hard to go even a day without reading some complaint that such and such was written by "AI" (an LLM). Why is this so obvious? Well, the core advantage that LLMs provide is that they don't compute. Inside an LLM, there is no arithmetic, no logical branches, no truth values. Phrases aren't generated to define or to resolve. They are generated to continue. Sure, we can direct the story to follow the steps of logical deduction, but that isn't anything like calculation. An LLM simply isn't invested in logic, precision, correctness, etc. the way we expect modern writers to be. It's not the em-dashes or the word choice that illustrates this, it's the fundamental perspective of the system.
We are sorely missing subjectivity. Natural language never was, and never will be, computable. You can't reduce a natural story to binary truth values without choosing an arbitrary perspective that resolves its ambiguity. The more precisely abstract our language gets, the more detached from reality our stories become. The more objective our assertions about reality are, the less relevant they can be.
My answer to this is to make the arbitrary choice of perspective a first-class feature. If we can explicitly decide what meaning is relevant, we should be able to weakly solve natural language processing. It seems like a pretty simple and obvious idea, but so far is easier said than done.