about
latent noise is a digital art piece tied to my doctoral dissertation in anthropology at stanford, arab technofutures. it began as an experiment under the stanford research computing graduate fellowship, which in the summer of 2024 handed an anthropologist access to a high-performance compute to experiment with the geometry of meaning produced by sentence-embedding models, the latent-space neighborhoods that nlp systems use to make sense of language. that work culminated in a presentation the following fall, titled “closer to god: dialect specific semantic associations in arabic.”
this piece triangulates that work with my own fieldwork, in part to make sense of a computational phenomenon sitting underneath most of modern artificial intelligence, and in part to stay with the elegance of the engineering, of the imaginaries it feeds on and propagates, as evidence of how a certain kind of knowledge is made. the experiment is suspicious, but it is not a debunking. you cannot be this suspicious of something you do not find astonishing. the only claim this experiment makes is that computation is a poor instrument for making claims about language, and it makes that claim by computing and arriving nowhere in particular.
music: this experiment is called “latent noise” as an ode to noise itself, the part of life that machines cannot ingest as “knowledge.” to accompany this experiment, it only seemed befitting to add as much noise as possible: visual, computational, linguistic, and auditory. tarab dub by hello psychaleppo plays behind the pages, over pink noise — a syrian producer’s rendition of oum kalthoum’s min elli aal (مين اللي قال).
the two pages
home: this page is an experiment anchored on a word, among several to be chosen from. when you pick one, it is rendered across several dialects to calculate the closest neighbors to the word and list them, from closest to furthest.
as you scroll down the page, you see a scale with modern standard arabic on one end and the dialect of aswan on the other, ranking the 25 dialects from the madar corpus by aldi score which measures “average dialectness” or a dialect’s noisy deviation from the “standard.” accordingly, msa scores a 0.04 while the dialect of aswan (the most “dialect-y”) scores 0.75. from this scale, you can select to highlight the dialects that you want to render the word you chose at the beginning. scroll down further, and see the word and its closest neighbors rendered under each dialect chosen. hover over the neighboring word, and see the corresponding sentence from the corpus and the calculated cosine distance between vectors in latent space.
nothing here is held as truth. the scale of dialectness is presented from msa to aswan as a line with numbers on it, but ask what dialectness even means? put forward as a scale in this way, tied to a number glitching in calculation (computation as performance art does quite a bit of work here too), tied to a study, tied to a mathematical formula and algorithm, tied to a science – it becomes difficult to dispute. but what does it mean for the dialect of aswan to be more or less dialect-y than the dialects of casablanca or cairo or beirut? everything ranked on the scale comes from a phrasebook (the source corpus, madar, was built by translating a basic traveling expressions book into 25 city dialects). the word neighbors calculated are rooted in the world of a tourist phrasebook. the numbers are real, in that they are computed from something, but what does that something really say about the complexity of the world?
essay: an argument, or an explanation, or a poem.
methodology
NYUAD’s MADAR corpus (Bouamor et al., 2018), twenty-six Arabic dialects plus Modern Standard Arabic, with sentence embeddings produced via OpenAI text-embedding-3-small in September 2024. Per-dialect content-word centroids are built over corpus contexts, restricted to nouns / verbs / adjectives by the CAMeL Tools MLE disambiguator (Pasha et al., 2014; Obeid et al., 2020). Average dialectness scores per dialect via AMR-KELEG/Sentence-ALDi (Keleg & Magdy, 2023).
sources
- MADAR corpus — Bouamor, H., Habash, N., Salameh, M., Zaghouani, W., Rambow, O., Abdulrahim, D., Obeid, O., Khalifa, S., Eryani, F., Erdmann, A., & Oflazer, K. (2018). The MADAR Arabic Dialect Corpus and Lexicon. LREC 2018.
- Sentence-ALDi — Keleg, A., & Magdy, W. (2023). ALDi: Quantifying the Arabic Level of Dialectness of Text. EMNLP 2023.
- CAMeL Tools — Obeid, O., Zalmout, N., Khalifa, S., Taji, D., Oudah, M., Alhafni, B., Inoue, G., Eryani, F., Erdmann, A., & Habash, N. (2020). CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing. LREC 2020.
- Embeddings — OpenAI
text-embedding-3-small, 1,536-dimensional, generated September 2024.
credits
- music — Pink noise generated locally with NumPy. The musical layer is Tarab Dub by Hello Psychaleppo, played through SoundCloud’s embedded widget.
- fonts — Amiri and Reem Kufi, both by Khaled Hosny et al., self-hosted under the SIL Open Font License.
- stack — Astro and TypeScript, custom CSS, no Tailwind.
contact
write to me at salmae [@] stanford [dot] edu