Toward a theory of aesthetic valence
Constraints, desiderata, questions
I love meditation and psychedelics as much as the next guy, and they’re formidable tools to study consciousness and valence. However, one area of research that seems to be critically overlooked is the aesthetic experience.
In fact, art and beauty are much more common, accessible, and repeatable for most people than deep meditative states or often illegal substances. Apparently, the average person listens to 21 hours (!) of music per week. Many report having peak experiences when engaging with live music, novels, or movies.
Given these facts, there should be a better understanding of why we seek these experiences and what they do to our psyche, or to the texture of our consciousness. What is actually going on when a piece of music feels good, in the particular, durable, sometimes life-rearranging way that art can be good?
If art can have such a strong therapeutic effect, I would argue that there’s even a moral imperative to figure out how to engineer more of the “good kind”.
I laid out this argument more in detail, along with an answer to the belief that “beauty is in the eye of the beholder” in my essay On Taste. I also gave a description of the “layers” of the aesthetic experience.
However, there are still mysteries in the way of a truly formal theory of aesthetic valence, but I’m optimistic that there is an answer. In fact, we can already know the shape of this theory, the constraints it has to follow, and the questions it has to answer. I will present them here.
1. The constraints
The constraints are things we know, either from phenomenology or from direct empirical/neuroscience data. If a theory can’t work within these constraints, it is most likely misguided.
C1: Structured interaction
Valence is fixed by neither stimulus nor perceiver alone. This shouldn’t be too hard to accept: you probably know that the same piece of art can have a vastly different impact on you based on your inner state and what you expect from it. Examples:
Kirk et al (2009): Identical images labeled “gallery originals” and “computer generated” were randomly shown to people. “Originals” received higher aesthetic ratings and showed significantly greater activation in medial OFC and vmPFC.
Gerger et al (2014): Same principle as above (“artwork” vs “press photograph”). The ones labeled as art dampened automatic negative emotional reactivity (measured through facial EMG and corrugator muscle) and increased aesthetic appreciation, even for disturbing images.
Silveira et al (2015): Again, artworks presented as “paintings from the MoMA” or “from an adult education center”. The museum label elicited stronger activation in self-referential / evaluative networks (precuneus, anterior cingulate, TPJ) and raised aesthetic ratings. The “art” framing recruits additional prefrontal/parietal loops (social cognition, memory).
If simple facts can change how someone feels (shown via neuro-imaging) when seeing an artwork, this eliminates stimulus-only objectivism.
But also, valence is not only dependent on the perceiver. It feels trivial to say that almost everyone prefers a major chord played on a grand piano over the sound of nails on a chalkboard. Beyond that, we can see convergence on non-trivial works. Complex movies like 2001: A Space Odyssey or Apocalypse Now are often cited as the best movies of all time; it’s hard to imagine millions of people across decades just pretending to like them.
You can even remove more variables. Take two titles of Michael Jackson’s 1982 album Thriller: Billie Jean and The Lady in My Life. The vast majority of people, even those who have never heard these songs before, will prefer the first over the second. You can’t say that status is playing a role here; these two songs are from the same artist, the same album, the same genre. There is something in Billie Jean that elicits a stronger average response.
So this constraint can also eliminate absolute projectivism.
C2: Valence integrates over a window
Instantaneous valence is not a function of the instantaneous input; it integrates over a window: backward (history) and forward (anticipation).
Imagine a very tense, escalating, exhilarating build-up in a techno song, followed by a drop consisting of a simple four-on-the-floor kick drum. Or 30 minutes of dissonant, chaotic orchestral swells, followed by a sustained major triad. By themselves, the 4/4 kick and the chord can be pleasant for a few seconds before becoming boring. But when you integrate them over a temporal window, taking the history into account, they can become overwhelming, euphoric, emotional.
The forward-looking side is softer but real: a chunk of the pleasure of the buildup lives in the anticipation, the “knowing” that something great is about to happen.
Again, this eliminates memory- or context- less feature readouts.
C3: Peak ≠ integral
The total value of an experience is not its momentary peak, and includes value persisting after the encounter ends.
An orgasm is one of the highest valence experiences a person will commonly encounter, yet it has barely any lasting impact, and the chase after the next one starts quickly after.
I don’t think anyone gets as much pleasure reading Dostoevsky as having sex, in that very specific moment (but hey, to each their own). Yet, over the course of the reading, a Dostoevsky novel can generate significantly more value. This positive impact can also last a long time after the novel is finished (new ways of seeing oneself, the world, etc…). Most readers would happily sacrifice one orgasm from their entire life if the alternative was never having read the book.
This means that however we measure the valence of an aesthetic experience, it can’t be read off the peak. Value accumulates across the whole encounter, and sometimes lives mostly in what it leaves behind. Whether it’s an integral or some other shape stays open.
C4: Wanting is not liking
The pull toward a stimulus is separable from the felt goodness of it. Liking is the actual hedonic pleasure; wanting is the incentive to pursue, and you can intensely want what you don’t like. The work of Kent Berridge has shown that the circuitry and psychological processes behind each is distinct.
Concretely, this means that we can’t use any behavioral evidence about aesthetic preferences (things like listening time, selection…) to measure valence because they are primarily wanting-sensitive, not liking-sensitive. Cleaner signals are in-the-moment phenomenological reports or other neural proxies for hedonic response.
2. The desiderata
The constraints above were the things a theory can’t contradict. The desiderata here are what it must deliver to be fully satisfying.
D1: Bridge from function to phenomenology
This is the hard problem of valence. While physical processes explain how the brain processes stimuli, they fail to explain why those signals feel intrinsically pleasant or unpleasant. How does a computational event (like resolving a prediction error) become a pleasant event? As far as I know, the Symmetry Theory of Valence is the only satisfying attempt at a solution, but we currently have no way to confirm it empirically, so this is still open.
The bridge must also explain why different types of events yield phenomenally different pleasures. A sublime landscape and a perfect pop song could both feel “8/10 good”, yet their felt sense might be totally different. A complete theory will explain why that is.
D2: Mechanism of interaction
C1 (structured interaction) says the interaction between the stimuli and the perceiver’s state sets the valence response; a theory has to say how. When you learn “this is a Rembrandt,” does that knowledge reach down and rewrite the early perception itself, or does it leave the percept intact and layer an extra evaluation on top?
The theory should formalize how cognitive and contextual layers (priors, beliefs, knowledge, autobiographical memory…) act as constraints that warp perception, explaining why identical stimuli (e.g., a Rembrandt original vs a copy) can produce different phenomenal shapes.
D3: Differential persistence
We should account for the variance in how long stimuli reward repetition of engagement. And why this variance seems to not be predicted by how good the first engagement was.
For example, a good joke could get you laughing uncontrollably for minutes the first time you hear it, but the second or third times barely do anything. A good song could be as good the 100th time as the first. Mulholland Drive could leave you completely puzzled the first time, then become better with each rewatch.
A theory should have a model of the trajectory of valence across encounters.
D4: Cross-context valence of emotion
Finally, we should explain why emotions that are negatively valenced in “the real world” (grief, fear, dread, melancholy…) are sought out in the aesthetic frame and often reported as good in the moment, not merely endured.
This is the paradox of tragedy. A sad song is felt as good while it’s playing. The poignancy/passion/resonance/intensity/etc are reported as sources of liking.
Regardless of the mechanism (conversion of negative to positive, direct liking, eudaimonia…), a theory owes a cross-context (art vs real life) account of what is going on.
In On Taste, I proposed a vague attempt at a complete theory from stimuli to computation to phenomenology. A proper theory would actually formalize this and clearly state the solution to each desideratum.
III. The frontier questions
Beyond the constraints that a theory must work within and the things it has to explain, there are five questions this theory needs to take a side on.
Is “aesthetic valence” a natural kind or just a sorting word? When we decide that a warm bath doesn’t count as an aesthetic experience, but a Bach fugue does, are we carving nature at its joints, or is it purely out of convention? What about birdsong? A sunset? Solving a puzzle? If there is no natural kind, much of the framing changes.
What is the valence sign of harsh texture/dissonance? Artworks that have seemingly abrasive or dissonant texture/surface (e.g., death metal, industrial techno, Francis Bacon paintings…) appear to be liked by millions of people across decades, making it hard to imagine everyone’s just faking it. Is the texture of this type of art (a) negatively valenced with value running through the unpleasantness, (b) a neutral way to capture attention while the valence is coming from elsewhere (rhythm, composition…), (c) directly liked in the moment or (d) is the question “surface vs. value” just malformed?
What is “resolution”, and at what level(s) does it operate? It’s tempting to consider the binary “tension - release” as a mechanism, but the seeming absence of resolution could itself function as a resolution (e.g., anti-humor, police movies where they don’t catch the killer, ambient trance…). Resolution looks layered (formal, semantic, affective, meta…). If so, how the layers relate to each other? And if everything can be thought of as a resolution, does this term have any weight at all?
Is valence scalar? We tend to assume that valence is one number, a scalar (=how good or bad the experience feels). But what if valence had structure beyond its magnitude? What if something could feel both really good and really bad at the same time, without the two valence numbers averaging out to 0? Bittersweet could be an example: the two valence directions feel different than a plainly neutral experience.
Is exhaustion ultimately universal? D4 says stimuli differ in how slowly their valence decays, but does everything eventually habituate given enough encounters, just with different time constants? And what’s the role of spacing? Usually, immediate re-exposure kills the valence (e.g., watching a movie twice in a row), but sometimes spaced re-exposure increases it (watching the same movie one year apart). Is “deepening on return” real, or is it just slow decay + good spacing (maybe you forgot some parts of it)?
Conclusion
Again, a theory of aesthetic valence has to be consistent with the four constraints (structured interaction, temporal window, peak≠integral, wanting≠liking), deliver on the four desiderata (function → phenomenology bridge, mechanism of interaction, differential persistence, cross-context explanation), and take a position on the five frontier questions.
As far as I know, we’re still very far away from any satisfying attempt. But I hope this will provide a framework and kickstart more research, discussion, and proposals. I also hope to provide my own in the coming months.


