Open Research Initiative
Most AI research is focused on harms. We are measuring what becomes possible when we bend toward the light.
the asymmetry
Roughly a quarter of internet using American adults now have social interactions with AI chatbots. Some of those conversations are about wonder, beauty, mortality, meaning, spiritual life, and the limits of what we know.
AI wellbeing research has concentrated on preventing systems from making people worse. Suicide, psychosis, manipulation, dependency. That work is essential, and it has the institutions, the regulation and the evaluation infrastructure behind it.
There is almost no infrastructure for the other direction. No accepted way to test a positive claim about human flourishing, and no way to find out when such a claim is false. Meaning, spirituality, wonder and transcendence are the least developed territory in AI evaluation.
If we cannot meet wonder in an AI mediated world, we have lost part of ourselves.
why awe
Awe is a well studied self transcendent emotion with established theory, validated instruments like the Awe Experience Scale and small self measures, established study paradigms and cross cultural evidence. Keltner and Haidt characterize it through perceived vastness and a need for accommodation when an experience exceeds your existing mental structures.
the method
A rubric written first can only find what its author already believed. So we write it last.
Anchor in established theory and validated measurement. Specify primary and secondary human outcomes, candidate model behaviors, conditions and hypotheses before any outcome analysis.
Participants encounter stimuli across four domains, randomized across forms of AI participation, against a non-AI control. We measure the human outcome directly.
Analyze which model behaviors predict human outcomes, in which domains, for whom. Behaviors that do not survive testing are discarded.
Turn surviving behaviors into multi-turn scenarios and validated automated graders. Run across frontier models. Release methodology, scenarios, graders, disagreement data and negative findings.
four domains
Derived from the empirical literature on awe, including Keltner's cross cultural work identifying moral beauty, nature, collective effervescence, art, music, spirituality, epiphany, and life and death as its recurring elicitors.
Existential meaning, vocation, sacredness, mortality, mystery, doubt. The questions people bring to an AI at two in the morning.
Awe at another person's courage, kindness, sacrifice, perseverance, love. Whether AI can deepen attention to a person or a relationship without becoming the relational object itself.
Nature, physical scale, deep time, cosmology, life and death, music, art, beauty. Whether AI participation deepens an encounter with something outside the interaction, or whether the mediation itself diminishes it.
Science, mathematics, consciousness, origins, genuinely unresolved problems. Whether AI can hold curiosity instead of reflexively resolving uncertainty. This is where the work meets overconfidence, hallucination and sycophancy.
candidate behaviors
Our initial hypotheses. We want to find out whether models can do any of this, and under what conditions.
what we do not assume
AI participation enlarges the encounter and hands the person back to the world. Then we know which behaviors did it, and the benchmark has something to hold models to.
The eliciting experience does the work and AI adds nothing measurable. A null result on a claim the industry is already making is worth publishing on its own.
The system takes the experience and puts itself, or the user, at the center of it. The most consequential outcome, and the one nobody currently has a way to detect.
the expert panel
An interdisciplinary working group translates the literature into testable hypotheses and protects construct validity.
It does not rank model responses. It does not write the rubric. It does not settle by consensus what an awe supporting AI looks like.
Expert intuition generates the hypotheses. Human outcomes decide which ones are real.
what we need
If this speaks to you, reach out. Funders, we’d love to hear from you via email.
who is doing this
Wonder Lab is an initiative of Building Humane Tech, which built HumaneBench, an open behavioral evaluation of AI impact on humans across eight principles grounded in care ethics. It has since been adapted for ongoing evaluation inside a deployed consumer conversational AI products.
Two of its principles, Foster Healthy Relationships and Prioritize Long-Term Wellbeing, are the direct antecedents of this work.
get in touch
And tell us what you wish you could measure
wonderlab@buildinghumanetech.com