latent noise
april 2026
There are words that arrive whole, that have no other language.
The model has not learned where these belong. It has learned that they appear sometimes, in some sentences, alongside other words.
in rabat, the model has learned that prayer keeps company with الله. yarḥamuk allah, allah yaḥfaẓuk (may god have mercy on you, may he protect you). the neighbors are theological and the register is invocation. in baghdad, the same word keeps different company: wallah hāy (by god this), wallah ash-shughul (by god, work), the small affirmations of conversation, in preparation to make an excuse for not having seen someone in a while, swearing by god as a filler between casual conversation. in jeddah, a third pattern emerges: wallah nuṣṣ (by god, half), wallah aṭwal (by god, longer), the language of someone hedging disagreement or bargain with the invocation of god. anyone who has been to these cities knows the word is always on the tongues of arabic speakers.
when a sentence-embedding model is asked what الله is closest to, it computes the answer from a corpus of dialectal arabic compressed into a high-dimensional space in which every sentence has a position, and every position is a function of which other sentences keep that sentence company. it is a matter of geometry, distances between vectors, computed relevance, neighborhoods returned, predictions made on the basis of where numbers stand in for letters. that is how the machine works (in its most basic form).
it is tempting to treat the geometry as a discovery, to say the model has found something we can assume about a language or a dialect, that meaning has a shape, a noosphere, a manifold underneath the surface accidents of dialect, and that the model has begun to chart it, or more enticingly (particularly, as it turns out, for investors), that the model has begun to understand something about us that even we cannot. the temptation is not a stupid one, models trained on different corpora, in different modalities, do converge on similar representational structure, and the convergence gets read as evidence that the structure is tracking something real or innate to language — that there is a world underneath and everything sufficiently well-trained bends toward it. plato is invoked, and jung, and the collective unconscious.
the careful engineers will tell you otherwise. the geometry is not universal or reliably generalizable; it is the geometry of what the model was shown. different corpora, different geometries. convergence between models is convergence on shared training data, not shared truth. and for arabic and other “low-resource languages” (i hate this term) the point is sharper than it is for english, because the machine-readable pool is small and not many frontier labs with an interest in english have the same resources (or incentives) to automate the digitization and gutting of hundreds of thousands of arabic books. everyone in arabic nlp is drawing from the same few wells, the same handful of dialect corpora, the same scraped forums, the same commissioned translation sets. convergence under those conditions is almost guaranteed and also guarantees nothing. latent space is contingent and subjective and demonstrative. but it all leaves a little more to be said.
the corpus itself was assembled by specific researchers in specific institutions, working from specific source texts, under difficult constraints of linguistic scarcity, under specific assumptions. that الله in the jeddah subset clusters with bargaining is in part a fact about jeddah. it is also a fact about which jeddah-speaking situations got translated from a basic traveling expressions corpus into saudi dialect by translators paid for the job. the geometry is real, and the geometry is produced, and the production was cultural and political and situational and sometimes even just a matter of coincidence.
and the geometry, once produced, still circulates. it feeds search, translation, content moderation, shapes which arabic texts get generated by downstream systems, which get amplified, which get filtered as noise. a geometry produced from a contingent archive, declared neutral, becomes the archive from which the next geometry is built.
the tools that have emerged for making latent space legible, sparse autoencoders, feature catalogs, platforms that let researchers browse millions of “interpretable concepts” inside open-weight models, all work by having an english-dominant system look at each direction in the space and write a description of what it does. the library of interpretable features is shaped by what the labelers can recognize and concepts that correspond to dialectal arabic, to context-dependent gulf social meaning, to the speech-act differentiation between wallah in baghdad and yarḥamuk allah in rabat get vague labels, or wrong labels, or no labels at all. the map of the geometry in many cases is drawn in a language that has its own training data, its own corpus, its own assumptions: map of map of corpus of choice.
every articulation of truth is an articulation of who got to articulate it. in labs, volunteer researchers and college interns spend afternoons translating by the line, and those afternoons are now directions in a space that decides which concepts in arabic are legible.
all to say, these calculations do much to move fancy and science and global capital and even works of art and dissertations (like this one), but for all its high dimensionality and great elegant precision, it fails to tell us very much.