AIESI 2026: Verification, Judgment, and the Economics of AI

At CAAI’s sixth annual AI & Economics Summer Institute, the question wasn't what AI can do. It was knowing how to use it reliably.

A fast-moving technology does not necessarily translate to a fast-moving economy. This widely held assumption was one among many challenged during this year’s AI & Economics Summer Institute (AIESI), an annual flagship event hosted by the Center for Applied AI at Chicago Booth. 

The grey area—between an idea’s potential and the action required to bring it into existence—is immensely complex and, for the participants of AIESI 2026, especially intriguing. For a full week of learning and collaboration, participants from around the world gathered on the University of Chicago campus to connect disparate ideas, share reflections, and grapple with emerging challenges. Now in its sixth year, AIESI maintains the momentum of the past few years, connecting PhD students with the researchers pushing the frontier of AI and economics.

Featured speaker Daniel Rock, Assistant Professor of Operations, Information, and Decisions at the Wharton School of the University of Pennsylvania, shared unique and human-centered strategies for taking AI to its next frontier. He described the opportunity researchers have to adapt to new technology and change the way they work, and drew a hard line between capability and gain. The two are not the same thing, and individual choice is what connects them. He put it simply: we shouldn't confuse a clear view with a short walk.

Rock was joined by faculty from leading research institutions across the nation, covering topics that ranged from financial and labor markets to law, policy, and job displacement. The central throughline has evolved from investigating the boundaries of AI’s capabilities to the necessary skills, assets, and structural foundations required to equip communities to effectively use these new tools in the fields being reshaped.

Some students arrived with research backgrounds already fundamentally intertwined with AI; others only started working with it in the last year. That spread gave the room a palpable diversity and a sense of eager urgency to shorten the walk from idea to impact at scale.

Under the Interface 

“This is an opportunity for individuals who have likely never been in a room together to spend a week arguing, building, and hopefully becoming colleagues for the next twenty years to come,” said CAAI Director Cathy Kiriakos, whose introduction set the tone for the week. Students then dove directly into the transformers behind LLMs with Sanjog Misra, CAAI Faculty Director and Booth's Charles H. Kellstadt Distinguished Service Professor of Marketing and Applied AI.

Working through which parts of the architecture function as research methods, Misra gave each component a social science counterpart. He presented embeddings as a nonlinear version of factor analysis, a statistical method familiar to most in the room. The softmax function in a model turns raw attention scores into weights that sum to one, the same way a discrete choice model in statistics turns utilities into choice probabilities. The technology that looks impenetrable from the outside, he argued, is built from tools and algorithms social scientists already trust and likely know.Sanjog Misra Presentation
That trust comes with a caveat; the same machinery that makes the tools useful has failure modes that look like success. Embedding models are trained to predict text, and similarity is only the distance between two vectors, a space where polarity is one thin dimension in a map built for general meaning. Two directly contradictory phrases—such as "Payment was received" and "payment was not received"---are interpreted as near-identical despite their clear differences.

A model also randomly selects from the scores it assigns to every next possible word, so the same question can return a different answer each time. The simplest way to use this novel technology—ask a question and trust the answer—can produce a falsity as confidently as the truth. Using it well means following Misra a level below the interface, to the numbers inside the model, rather than the sentences coming out of it. The tools are already familiar; the discernment they ask for is not.

The Cost of Verification

Misra’s discernment has a price, and Suproteem Sarkar, Assistant Professor of Finance and Applied AI at Chicago Booth, has watched engineering teams pay it. Sarkar found that adoption inside software teams tracked how quickly an AI's work could be checked. Engineers on front-end and integration tasks, where a wrong answer shows up in seconds, used agents constantly. Those working on architecture and data modeling, where a bad decision takes months to surface, mostly avoided them. The capability was the same for both groups. What differed was the cost of finding out whether it worked.

That cost is not fixed, and in some fields, issues with verification are more structural. Ralph Koijen, Chicago Booth AQR Capital Management Distinguished Service Professor of Finance and Applied AI and Fama Faculty Fellow, pivoted the conversation towards asset pricing, where checking AI’s work is particularly difficult. 

There, the standard way to check a signal is to test it against history, which becomes impossible when the model has already read the history, as look-ahead bias misrepresents recall as prediction. Reflexivity closes off the other side: a prediction signal stops predicting once enough capital trades on it, limiting its applicability to a specific timeframe.                                 

Rather than lowering the standard, Koijen approaches the test itself differently, asking questions the model could not have memorized the answer to and leaving the market to resolve them going forward. The obstacle was never the model's capability, but the design of the test; a solution came from knowing where asset pricing hides its answers. That is the part that transfers: not a particular method, but the ability to identify what discernment looks like in a particular field.

The same expense operates at the level of an entire academic field, where verification has a name: peer review. Ben Golub, professor of economics and computer science at Northwestern University, identified that cost early enough to build a tool around it.

Golub co-founded a small company whose core function is to serve as an AI technical referee, spending hours of parallelized compute on academic papers to check for consistency in a way most humans don't have the time or tenacity for. It has caught errors that sat unnoticed in published work for years, a testament to what gets past human review.

A reviewer's verdict bundles two questions: is this correct, and does this deserve the field's attention? The time required of the first has always crowded out the second. 

But verification, unlike relevance, is codifiable, and machines are getting good at it quickly. What's left is knowing what a worthwhile question is and what counts as an answer, which is itself contested in social science. Knowing where that line falls is human discernment, and mistaking one side for the other would mean social science gets done to us rather than by us.

Judgment like that has nothing to check itself against except other people, and the week was built to supply them. Randall Balestriero of Brown on what world models still cannot infer, Sarah Cen of Carnegie Mellon on the law taking shape around AI, Lindsey Raymond of MIT on labor markets, Jessy Lin of UC Berkeley on continual learning, Shreya Shankar of UC Berkeley on putting agents to work on unstructured data, Booth's own Kawin Ethayarajh on post-training, and Alex Imas on displacement. A variety of vantage points connecting the challenges that come with AI, none of them settling it alone.

“We picked this cohort because we think the field needs exactly this combination of people talking to each other more, not less,” said Kiriakos. “If a friendship or a paper or an argument you're still having in five years starts this week, we did our job.” Rock echoed the sentiment. "This is a team sport," he said. “Disagreements are as useful as agreements, so nobody should be shy about either.”

That is what the Institute is for. The rebuilding that this moment demands, as Golub put it, will be done by the people who are at the beginning of their careers right now. A room where they can think alongside the incumbent researchers paving the way today is more valuable than it's ever been. Nobody left with a shortcut to the path ahead. They left with a sharper sense of direction and the insight to walk the path with confidence.



2026 AI & Economics Photo Recap

2026 AI & Economics Summer Institute

Florence Ukeni

Florence Ukeni
More from Chicago Booth