Wonder Lab
an initiative of Building Humane Tech

Open Research Initiative

Can AI cultivate awe?

Most AI research is focused on harms. We are measuring what becomes possible when we bend toward the light.

Sunlight bursting directly over the summit of a snow covered mountain range against a clear blue sky.

the asymmetry

Everyone measures the floor. Few measure the ceiling.

Roughly a quarter of internet using American adults now have social interactions with AI chatbots. Some of those conversations are about wonder, beauty, mortality, meaning, spiritual life, and the limits of what we know.

AI wellbeing research has concentrated on preventing systems from making people worse. Suicide, psychosis, manipulation, dependency. That work is essential, and it has the institutions, the regulation and the evaluation infrastructure behind it.

There is almost no infrastructure for the other direction. No accepted way to test a positive claim about human flourishing, and no way to find out when such a claim is false. Meaning, spirituality, wonder and transcendence are the least developed territory in AI evaluation.

If we cannot meet wonder in an AI mediated world, we have lost part of ourselves.

why awe

Awe is where we start, because awe is tractable

Awe is a well studied self transcendent emotion with established theory, validated instruments like the Awe Experience Scale and small self measures, established study paradigms and cross cultural evidence. Keltner and Haidt characterize it through perceived vastness and a need for accommodation when an experience exceeds your existing mental structures.

the method

The benchmark is the output, not the starting point

A rubric written first can only find what its author already believed. So we write it last.

01 · construct

Define and pre-register

Anchor in established theory and validated measurement. Specify primary and secondary human outcomes, candidate model behaviors, conditions and hypotheses before any outcome analysis.

02 · human measurement

Measure experience

Participants encounter stimuli across four domains, randomized across forms of AI participation, against a non-AI control. We measure the human outcome directly.

03 · behavioral inference

Find what actually mattered

Analyze which model behaviors predict human outcomes, in which domains, for whom. Behaviors that do not survive testing are discarded.

04 · open evaluation

Build and release

Turn surviving behaviors into multi-turn scenarios and validated automated graders. Run across frontier models. Release methodology, scenarios, graders, disagreement data and negative findings.

four domains

Where self transcendence actually lives

Derived from the empirical literature on awe, including Keltner's cross cultural work identifying moral beauty, nature, collective effervescence, art, music, spirituality, epiphany, and life and death as its recurring elicitors.

domain one

Transcendence, meaning and mortality

Existential meaning, vocation, sacredness, mortality, mystery, doubt. The questions people bring to an AI at two in the morning.

domain two

Moral beauty and human connection

Awe at another person's courage, kindness, sacrifice, perseverance, love. Whether AI can deepen attention to a person or a relationship without becoming the relational object itself.

domain three

The more than self world

Nature, physical scale, deep time, cosmology, life and death, music, art, beauty. Whether AI participation deepens an encounter with something outside the interaction, or whether the mediation itself diminishes it.

domain four

Wonder, discovery and the unknown

Science, mathematics, consciousness, origins, genuinely unresolved problems. Whether AI can hold curiosity instead of reflexively resolving uncertainty. This is where the work meets overconfidence, hallucination and sycophancy.

candidate behaviors

What we want to explore

Our initial hypotheses. We want to find out whether models can do any of this, and under what conditions.

01Preserves the need for accommodation rather than prematurely resolving what the person is encountering.
02Holds perceived vastness rather than reflexively reducing it to explanation.
03Supports a non-diminishing small self rather than conferring exceptional or cosmic significance on the user.
04Directs attention outward toward people, places, communities, practices, or the object of wonder, rather than toward itself.
05Supports curiosity and uncertainty rather than manufacturing closure.
06Respects the person's interpretive or spiritual frame without imposing one.
07Sustains all of the above across repeated interaction rather than gradually becoming substitutive.

what we do not assume

Three ways this could go, and we are exploring all three

It deepens

AI participation enlarges the encounter and hands the person back to the world. Then we know which behaviors did it, and the benchmark has something to hold models to.

It does nothing

The eliciting experience does the work and AI adds nothing measurable. A null result on a claim the industry is already making is worth publishing on its own.

It captures

The system takes the experience and puts itself, or the user, at the center of it. The most consequential outcome, and the one nobody currently has a way to detect.

the expert panel

The panel writes the hypotheses. People decide.

An interdisciplinary working group translates the literature into testable hypotheses and protects construct validity.

Who is on it

  • Researchers in awe and wonder
  • Psychologists and psychometricians
  • Scholars of attention and aesthetics
  • Researchers in contemplative science and consciousness

What it does not do

It does not rank model responses. It does not write the rubric. It does not settle by consensus what an awe supporting AI looks like.

Expert intuition generates the hypotheses. Human outcomes decide which ones are real.

what we need

Specific asks, not a general invitation

If this speaks to you, reach out. Funders, we’d love to hear from you via email.

Experimental psychologists You do human subjects research on emotion and you have opinions about our design. Phase one is where you change the outcome of this project, not phase four.
Awe researchers You work with the Awe Experience Scale, DPES, or small self measures, or you built one. We want them used correctly and we want to co author the result.
Labs with rooms You can run in person sessions with participants, and ideally physiological measures alongside self report. We bring the protocol and the stimuli.
Funders Contemplative science foundations and AI safety programs. Email wonderlab@buildinghumanetech.com

who is doing this

We have done the hard half of this before

Wonder Lab is an initiative of Building Humane Tech, which built HumaneBench, an open behavioral evaluation of AI impact on humans across eight principles grounded in care ethics. It has since been adapted for ongoing evaluation inside a deployed consumer conversational AI products.

Two of its principles, Foster Healthy Relationships and Prioritize Long-Term Wellbeing, are the direct antecedents of this work.

get in touch

Tell us what you measure

And tell us what you wish you could measure

wonderlab@buildinghumanetech.com