The Future of Advertising Papers

Paper No.29 · Attention Economics

Media Beyond the Screen: Spatial, Audio, and Wearable Interfaces

Pierre Subeh·June 29, 2026·8 min read

Abstract

Screens are losing their monopoly on media. Spatial overlays, voice, and wearables each carry attention differently, and each breaks a core assumption of screen era advertising. I map the three interface families and the distinct grammar a brand needs for each.

Everything we know about digital advertising was learned on glass. Impressions assume a viewport. Creative assumes pixels. Viewability assumes a rectangle scrolling past a fold. For thirty years the glass was so universal that we stopped noticing it was an assumption at all. Now media is escaping the pane, and it is escaping in three directions at once: into space, into the ear, and onto the body.

I want to be precise about what this paper is and is not. It is not another metaverse essay; I have watched enough headset hype cycles to be allergic to them. It is a working map of three interface families that already carry real audience time in 2026 and will carry much more by 2030, and an argument that they are not three flavors of one channel. Spatial, audio, and wearable interfaces have different physics of attention, different tolerance for commerce, and different failure modes. Brands that port screen thinking into them will fail in three different ways.

The stakes are simple. Screen attention is a contested, saturated market where every point of share is bought from an incumbent. Off screen attention is the only expanding frontier left. The question is whether you arrive there with a native grammar or with banners in your luggage.

Key Findings

  • Screen media is proximity neutral: everything is the same distance away, behind glass. Off screen media reintroduces the body, and with it distance, direction, and presence as creative variables.
  • Each interface family has what I call an Interface Grammar: the native rules for how information may enter without being rejected. Spatial grammar is annotation, audio grammar is companionship, wearable grammar is glance.
  • Audio is the sleeping giant of the three: it already owns hours of daily attention, requires no new hardware behavior, and remains priced like a niche channel.
  • Wearables carry almost no advertising capacity in the traditional sense, but they own the most decision adjacent seconds of the day, which makes them the highest leverage surface per second in media.
  • The measurement stack for all three is a decade behind the audience shift, which means early movers will buy underpriced attention against embarrassingly weak proof, exactly like early social.

The Return of the Body

Glass abstracted the audience into a pair of eyes and a scrolling thumb. Off screen interfaces put the audience back into a physical situation. A spatial overlay knows you are facing the shelf. An earbud knows you are running. A watch knows your heart rate just spiked. The screen era's core creative question, what should the ad look like, gets replaced by a stranger one: where is this person's body, and what does that entitle us to say?

This is a bigger deal than the hardware. Screen advertising could ignore context because the context was always the same: person, glass, elsewhere. Off screen advertising is embedded in rooms, streets, workouts, and commutes. Embedded media inherits social rules. There are things you may say to someone at a bus stop that you may not whisper into their ear at night, even though both are technically audio placements. Media planning is about to need manners.

Spatial: The Grammar of Annotation

Spatial computing, glasses and their descendants, does not create new attention. It relabels the attention people already spend on the physical world. The native act of a spatial interface is annotation: this building is the restaurant you booked, this product has the better rating, this street is closed. The grammar, then, is that commercial information must behave like a useful label on a real thing.

What breaks: interruption. A pop up on glass costs you a glance; a pop up on the world occludes reality itself, and users will treat occlusion as a physical offense. What works: presence at the point of comparison. The brand whose truthful, structured information gets pulled into the annotation layer when a shopper looks at a shelf wins the most valuable second in retail. Notice what this implies: spatial advertising is won upstream, in data quality and entity authority, not in the moment of render. The work I described in entity SEO and knowledge panels is the direct precursor: whoever the annotation layer trusts about the world gets to write on it.

My timeline: spatial remains a shallow audience through 2027, then compounds as glasses normalize in wealthy urban markets. The annotation supply chains, which brands feed the labels, are being negotiated now, quietly, and the defaults set by 2028 will be entrenched by 2031.

Audio: The Grammar of Companionship

Audio is the most underrated surface in marketing, and I say this as someone who buys it. Earbuds have made audio the default background layer of life: commutes, chores, workouts, work itself. Hours per day, no screen required, no new behavior needed. And unlike glass, audio attention is companionable. A voice in your ear occupies a social slot somewhere between narrator and friend.

The grammar of companionship explains everything that works and fails in audio. Host read endorsements outperform interchangeable spots because the listener's relationship is with the voice, and the voice is vouching. Synthetic voice ads inserted programmatically into intimate shows underperform because they violate the companionship, a stranger suddenly speaking in a friend's living room. As always on assistants take over more of the audio layer, the companionship grammar hardens further: one trusted voice, mediating everything, with room to recommend approximately one brand per need. Winning that single recommendation slot is answer engine work, the same discipline I laid out in the answer engine optimization guide, executed for the ear.

Priced against the hours it owns, audio is still cheap in 2026. I do not expect that mispricing to survive past 2028, and I am telling clients to build audio muscle, shows, voices, sonic identity, before the correction, not after.

Wearable: The Grammar of the Glance

The watch and its cousins are the most hostile advertising environment ever shipped, and that is exactly why they matter. A wrist interface offers seconds, at a glance, in moments that are frequently decision adjacent: about to eat, about to buy, about to leave, about to skip the workout. Nobody will tolerate an ad there. Everybody will tolerate help there.

So the wearable grammar is the glance: a brand earns wrist presence only by compressing genuine utility into a two second read. The payment that clears instantly. The loyalty pass that surfaces at the register. The hydration nudge from the beverage brand that is actually about hydration. Wearables force the discipline the rest of marketing avoids: if your brand's value cannot survive compression to a glance, the wrist will reveal that you do not have one. I explore the strategic consequences of computing that never turns off in Paper No.26; the wrist is that thesis in miniature, running today on a hundred million arms.

Per second of exposure, I believe wearable presence is the highest leverage media in existence, precisely because the seconds sit so close to actions. It is also nearly unbuyable with money alone, which keeps the tourists out.

The Measurement Vacuum

All three families share one condition: the measurement stack is not there. There is no mature viewability standard for an annotation, no attention metric for a companionable voice, no impression currency for a glance. The industry will spend years arguing about it, and during those years, attention in these channels will be systematically underpriced relative to its influence, because budgets follow proof and proof follows infrastructure.

I have seen this movie. Early search and early social were bought by people comfortable acting ahead of measurement consensus, and they acquired positions that laggards later paid decade long premiums to approximate. Falsifiable version: by 2030, at least one of these three families will have produced its own native measurement currency and a corresponding land grab story, and the brands cited in that story will have entered before 2027.

What I Would Do About It

Assign the three grammars to owners, plural. Spatial, audio, and wearable are different crafts; a single innovation lead with one experimental budget will average them into mush. Small dedicated efforts, each fluent in its grammar, each with permission to ignore screen era KPIs for two years.

Fix your annotation feedstock now. Structured data, entity authority, verified attributes, honest reviews. The spatial layer will be assembled from machine readable truth, and the assembly has already started.

Build one owned audio asset before buying audio reach. A show, a recurring segment, a voice. Companionship cannot be rented efficiently; the rental market prices it as intimacy the moment you prove you can hold an audience.

Prototype a glance. Take your core value proposition and force it into a two second wrist interaction that helps someone. If the exercise fails, that failure is strategy feedback, not a channel problem.

And protect these experiments from the quarterly machine. Off screen media in 2026 is a position building game. Measure learning velocity and earned presence, revisit in 2028, and thank yourself in 2031 when the glass finally, visibly, stops being where the people are.

Cite this paper

Subeh, P. (2026). Media Beyond the Screen: Spatial, Audio, and Wearable Interfaces. The Future of Advertising Papers, No.29. https://www.pierresubeh.com/research/media-beyond-the-screen

No.28

Attention Provenance: Auditing Where Ad Views Actually Come From

No.33

The Post Cookie Decade: Identity Without Surveillance

More in Attention Economics