Symbols of digital customer feedback
Credit: Getty Images

Chicago Booth Review Podcast How Measuring “Vibe” Can Explain Consumer Choices

How much do consumers care about shoppers’ reviews and product photos? And can economists capture that information and turn it into data? Chicago Booth’s Giovanni Compiani talks about his research using artificial intelligence to analyze “unstructured data.” Could his approach help retailers to measure the “vibe” of a product?

Apple Podcasts badge
YouTube Music badge
Spotify badge
Episode Transcript

Hal Weitzman: How much do consumers care about shoppers' reviews and product photos? And can economists capture that information and turn it into data? Welcome to the Chicago Booth Review Podcast, where we bring you groundbreaking academic research in a clear and straightforward way.

I'm Hal Weitzman. Today, I'm talking with Chicago Booth's Giovanni Compiani about his research, which uses AI to analyze unstructured data. Could this approach help retailers to measure the vibe of a product? Giovanni Compiani, welcome to the Chicago Booth Review Podcast.

Giovanni Compiani: Thanks so much for having me.

Hal Weitzman: We're delighted to have you talk about vibe. There aren't many economists who focus on vibe, but you do. And your work is kind of about using AI, you take measurements beyond the numbers, and you're analyzing what you call unstructured data. Let's start there.

What's unstructured data? Why historically, why has it been so hard for, a headache for researchers to try to capture that and how it shapes customer behavior?

Giovanni Compiani: Right. So by unstructured data, we essentially mean any data that doesn't come in a standard numeric form. So data that cannot be put in a standard Excel spreadsheet. So think product reviews, product descriptions, product images in the context that we're looking at.

Hal Weitzman: But you're not talking about the numerical part of customer reviews, because obviously if I give it a certain number of stars, you can put that in the spreadsheet.

Giovanni Compiani: Right. So that is an example of a data set that maybe captures vibe or a type of data that captures vibe that has been used in the past. But all the other things that cannot be quite condensed in a number like a star rating have traditionally been ignored or at least not exploited to the max in the past.

Hal Weitzman: Fascinating. So traditionally, how have economists tried to capture things like the things you were looking at, user-friendliness or visual design?

Giovanni Compiani: Right. I think the short answer is a lot of these features were just ignored. As I said, there were exceptions, maybe star ratings, maybe some features were hand coded by the researcher. You could look at the color of shirts and you can hand code that. You could hand code the material. But obviously there were limitations to this.

And now we are in a position to really use the data that we have available. Again, from reviews, from product descriptions, from images in a much more informative way and really incorporate it into our models of consumer choice.

Hal Weitzman: So you are basically expanding what we mean by data.

Giovanni Compiani: That is the goal and we're doing it in the context of consumer choice. I don't want to say this data has never been used before, but in studying how consumers make decisions in consumer goods, there has been certainly something that has been underused in the past.

Hal Weitzman: Okay. So tell us, obviously AI is complicated topic once you get into how it actually works. But in basic terms, in plain English, tell us how does your model work? What does it do when it looks at a product? And you're using product pages on Amazon, right?

Giovanni Compiani: That's right. So in the paper we do use this source of data, which is sort of readily available to anyone and for a lot of different product categories. And the approach really works in two main steps.

So the first step is we take the unstructured data, so the product images, the product reviews, the product descriptions and titles, and we essentially reduce their dimensionality quite dramatically using machine learning and AI tools. So think of this as a dimensional reduction step because you can think of an image as a very high dimensional piece of data, right?

Hal Weitzman: Which just means basically-

Giovanni Compiani: It has a lot of mainly pixels, right?

Hal Weitzman: ... there's a lot of stuff going on.

Giovanni Compiani: There's a lot of pixels and in a text there's a lot of words. So it's a very high dimensional object that is just hard to handle for us researchers. So what we do is we use these machine learning and AI tools to essentially reduce the dimensionality, but in a way that hopefully retains information about the products. And in particular, for our purposes, what is important is to capture how similar or dissimilar consumers view products in a given category.

So up until recently, we didn't really have models that did this in a very good way, in an informative way that we're able to do this dimensional reduction in an informative way, but now we're in a position where the models have gotten really good. So for images, you can think of classification models. So these are models that underpin, for example, facial recognition, so the technology that allows you to go through customs by matching your face to your passport.

And so we use those to understand and to condense the information containing product images. For text, the key here is transformer models. And these are the models that underpin on LLMs. So these are models that, as we all know, do a really good job at capturing semantics and capturing the meaning of sentences, not just words in isolation. And they also do a good job, it turns out, at distilling the information contained in product reviews and product descriptions.

Hal Weitzman: Okay. So they're looking at vast numbers of images, product reviews, the kind of things that people regularly post on Amazon.

Giovanni Compiani: Exactly.

Hal Weitzman: And how influential are those on consumer's choices?

Giovanni Compiani: Right.

Hal Weitzman: Do they care more about what the company tells me things that the product is like? Or do people value more what other consumers love?

Giovanni Compiani: Exactly. And this gets us to the second step. I said there's two steps, but I really only covered the first one. And the second step is really about this. It's about connecting this lower dimensional representations of the product images, product reviews, product descriptions from the first step to the outcome that we care about which is consumer choice.

And so here we embed these lower dimensional representations into a standard model of consumer choice. But this step is really key because it essentially allows us to understand which of these features that we have learned from the first step matter for consumer choice.

And so to your question, this really depends on the category. We find that depending on the category sometimes, because we apply this to 40 Amazon categories. So that really span the gamut in terms of they go from pet food to video games to bedsheets.

And depending on the category, we find that different sources of data are more informative in predicting consumer choices. So ultimately it's an empirical question. It depends on the data. And in one of the papers that have in this agenda, we really develop tools to guide researchers towards understanding which data sources are most informative for the given category that you're looking at.

Hal Weitzman: Okay. So just so I understand, so it sounds like for some, if I'm buying pet food, I might care more, for example, about what other pet owners say about the product. Versus bedsheets or something where I might care more about what the manufacturers is. Is that right?

Giovanni Compiani: That would be one example. And then you can think that some categories, the look of the product probably matters more than in others. Maybe in pet food, the look of the packaging is not first order for consumers, but maybe if we're talking about clothing or furniture, it might be more important.

Hal Weitzman: Okay. And do you also include video? Because you have a lot of video reviews on Amazon or unboxing type reviews. Do you include those?

Giovanni Compiani: Yeah. So in the current paper, we don't. This is something that we don't do, but the machinery that we develop could absolutely use that also.

Hal Weitzman: It would be possible to do that.

Giovanni Compiani: Yes, as long as we are able to sort of reduce the dimensionality of the data, in this case video, and we do have models for that, we're good to go.

Hal Weitzman: Okay. So what you're trying to work out is whether, just explain it to us, whether consumers, you said a classic model of consumer choice, how they ultimately choose to buy the product. Whether they value the consumer reviews higher than the official description of the product. Is that what you're trying to find out?

Giovanni Compiani: So the goal really is to understand what consumers value, how much consumer value different products, and how they would react to changes in the environment. For example, a change in prices. We're in the era of Trump tariff. So we might ask, for example, what would happen if the prices of products from a certain country were to go up?

How would consumers react? What would they substitute to? What would they switch to as a result of this? So these are the kinds of questions that we're thinking about. And they are obviously of interest not just to policymakers, but also to manufacturers and retailers as well who need to decide how to price products at which price point, which products to offer.

And all these questions hinge on, again, how consumers view the products in the choice set in their category, whether they view them as very close substitutes to each other or as very differentiated.

Hal Weitzman: I see. Okay. So presumably it would be also helpful for the manufacturers or the vendors to know how consumers make these choices, right?

Giovanni Compiani: Definitely.

Hal Weitzman: Not just what the substitutes are, but how the whole process works.

Giovanni Compiani: Absolutely. But one example of a question that is sort of important for manufacturers would be not just pricing, but also for new products, right?

Hal Weitzman: Right.

Giovanni Compiani: If you're thinking about developing a new product or if you already have developed a product and you want to predict demand for that new product, which is obviously something you want to do to plan in advance, the machinery can be used for this.

So in this case you would, it's interesting. In this case, you typically wouldn't have review information because the product is not out there yet. So that's not a data source that you could rely on. But as long as you have a prototype, for example, an image that you can look at or a description of the product and its characteristics, that could be data that can still be fed into our machinery and then spit out a prediction for demand for the new product.

And this is crucially going to hinge on, again, how close that product is to the existing ones that consumers have available to them.

Hal Weitzman: Amazing. Okay. Fascinating stuff. So tell us, let's get into what the experiments you actually ran, because I know one was with thousands of people picking books. And you got them to go to think about their first choice and their second choice. Just explain what you did and what you found.

Giovanni Compiani: Sure. So in the experiment, again, we ask our participants, around almost 10,000 of them, to make two choices. So we first ask them, show them 10 books with their covers, with their obviously titles. We randomize the prices of the books and we also provide review and plot information. And then we say, "Okay, pick your favorite book among these 10 books."

And then for each participant, we say, "Okay, now imagine this first book was removed from the choice." And we actually physically remove it from the display, from the interface that consumers see. And we ask them to choose again among the remaining nine books.

And so why do we do this? We do this because the two data sets really allow us to assess the performance of different models at answering these what if questions that we're interested in. So we use the first choice data to train our models.

This is the data that is usually available to us researchers. Data from what consumers picked, what was their favorite option among a set of options that they have available. So this is the data that we use for training our models.

And then the second choice data we use to validate our models, to understand how well they perform at predicting what if sort of choices. And so in this case, the what if question is what would happen if your first choice was removed? What would you switch to? What would your second choice be?

Because we ask this question directly, we have a ground truth that we can use to compare the performance of different models. And when we do this, we realize that incorporating this unstructured data, in particular in the case of books reviews, turns out to be very important to predict the second choices. And the performance really improves relative to the case where we ignore the unstructured data and just rely on the standard numeric attributes.

Hal Weitzman: Fascinating. So if you're selling something, you want to know what the consumers might have bought if your product weren't available.

Giovanni Compiani: Right. Exactly. And maybe there's an out of stock sort of occurrence and then you would want to be able to recommend an alternative product, then that could also be a use case.

Hal Weitzman: Giovanni Compiani, in the first half, we talked about your research on AI and vibe. We haven't really talked about vibe. How do you define vibe?

Giovanni Compiani: I would say anything that sort of affects consumer behavior and how much they like and they're drawn to a product, but it's really just hard to capture using numeric quantitative attributes of a product. So it could be the look of a shirt, of a dress. It could be even the user-friendliness of an app. That is obviously something that affects your experience, but it's just hard to quantify using standard numeric attributes.

Hal Weitzman: Okay. So are you the first economist or first in your group, your core, are you guys the first to measure vibe, do you think?

Giovanni Compiani: I would say people have definitely thought about this as an important driver of consumer behavior, but up until recently, we really didn't have the tools to unpack it a little bit. And so that is the reason why we're able to do this.

Hal Weitzman: I was just going to say a whole new field. Vibonomics. I could see a whole new, could be a podcast on its own.

Anyway, so I want to think about the policy implications of this because you and I have talked about this, not on the podcast, but now we're going to. So for example, imagine there were two huge companies, Apple, Samsung, and they wanted to merge. How could a policymaker use your research to protect consumers and make sure that the choice was preserved?

Giovanni Compiani: Right. So this is a great use case for our method. And let's think about the potential merger of two big companies like Apple and Samsung. One big concern is that prices may rise after the merger. This company is going to have a lot of market power. It's going to be able to increase prices and to the detriment of consumers.

Now, the extent to which this happens really hinges on how consumers view the Apple products and the Samsung products as substitutes to each other. And to clarify, let's consider the extreme case where the products are very different. The consumers effectively don't think of them as substitutes. So there's Apple people and there's Samsung people and they don't really sort of interact much.

In this case, the merger is probably not going to do a lot of harm because effectively there's not a lot of competition among the two companies. There's not a lot of price competition among the two companies because these are effectively separate markets. And so the merger is not going to remove price competition. It's not going to lead to a lot of harm for consumers. In fact, it might lead to benefits if there's efficiencies on the supply side that these companies can, synergies that they can harness.

On the other hand, if these products are close substitutes, meaning consumers do switch from one to the other, if the first one becomes more expensive, let's say Apple products become more expensive, then there is a potential for consumer harm. Because now the merger is going to remove a lot of competition, price competition from the market. And as a result of the merger, there could be a big price spike that is going to harm consumers.

Now, the extent to which consumers view these products, the Samsung products and the Apple products as substitutes or not, is a key question. And the tools that we developed can be used to address this. So standard tools have answered this question just using the physical specifications of the Samsung and the Apple products.

And those for sure still matter quite a lot. But there are other additional features, maybe again, the user-friendliness of the device or how premium a device feels that do matter for consumer choice and that traditionally we just weren't able to capture.

Hal Weitzman: Okay. So there's a whole, I can see the lawyers thinking, oh, this is great. Because there's a whole new area.

Giovanni Compiani: That is the hope.

Hal Weitzman: I mean, a lot of this is about judgment and you're saying, "Well, it's still about judgment." But we can bring in a lot more data that would make that more substantial argument that lawyers presumably could make-

Giovanni Compiani: That is right.

Hal Weitzman: ... arguing in favor or against mergers.

Giovanni Compiani: And it is data that is readily available. And this I also want to stress because it's the data that is on the internet and we can just use it to inform our models and make them richer.

Hal Weitzman: Okay. Fascinating. Okay. So I just want to go back to a methodological question, which is about how the algorithm works. How does it know? If it's just looking at customer reviews, how would an algorithm tell if I'm choosing one book over another?

Giovanni Compiani: That's right. So this is where really I want to go back to what we talk about in the first part, which is these two steps of the approach. So there's the first step which is the ML algorithm or the AI algorithm that turns the review into numbers, like a small set of numbers.

And you're right, these numbers per se need not be related to what consumers care or don't care about. The key is that then we connect these numbers to consumer choices via a model of consumer choice. And so that is the part where we are essentially learning which of these features that we've learned from the first step matter and which ones don't for consumer choice.

And so what we do find is that reviews do matter. So the reviews, even when we sort of shrink them down to these very small lower dimensional objects, they're still informative in predicting consumer choice. But we also find that other sources of data, for example, images in our case, for eBooks, we find that the covers of the books don't matter much for consumer choice, which is perhaps not surprising. Don't judge a book by its cover.

Hal Weitzman: You don't carry around an ebook to impress anyone.

Giovanni Compiani: Right. Exactly.

Hal Weitzman: All right. Okay. Very good. So listening to you speak, I can't help feeling that the applications of this could be huge, right? Way beyond just buying products online because I'm thinking of particularly things like Airbnb rentals or hotel spots or vacation spots. That very often, the official information you get is basically worthless beyond a star rating.

It's a three star hotel, but you never really know what that means. It could be terrible. It could be great. Presumably there's some variation within that.

So I as a consumer will tend to go to reviews. Are the rooms clean? Do they have the equipment? Does it work? That kind of thing. I mean, what are the other applications that you've thought about? Is it things like hotels and vacations?

Giovanni Compiani: Definitely. So that is one that we are definitely thinking about because it's, as you mentioned, one where, for lack of a better word, the vibe really matters a lot. So there's some sort of physical specification, physical features of an Airbnb or a hotel that as researchers, we can collect the number of rooms, the number of bathrooms, whether it has a pool, et cetera, at the location.

But there's also a lot of other information that is conveyed in this case by reviews and I would say images as well are quite important in this category, that by introspection we all know matters a lot for consumer choice. And also it's oftentimes indicative of the experience that you're actually going to have when you do go to that spot, and that is just not captured by standard attributes. So yes, I think accommodation, hotels, short-term rentals, even housing, I think could be an interesting application.

Hal Weitzman: Fascinating. Yeah. So this could really open up a whole load of this. All the photos that we're posting and the reviews that we're doing could turn into important data that will hopefully help improve vacation spots or get Airbnbs to be cleaner and that kind of stuff.

So we talked a little bit about this, but if I were a manufacturer and I were launching a product, how would I use your technology? What kind of are the three things I could do using your approach?

Giovanni Compiani: Right. So you would need some kind of data on the product that you plan to launch. So again, reviews, as we mentioned before, are typically not going to be available for a new product because it's by definition, it's not out there yet. But you might have an image of a prototype. You are going to have some specifications, some of them numeric, but you could also have a description of the product or what the product is intended to do.

And all of this information could be just fed through our algorithm. And what you can get out of this is a couple of things. One could be a prediction of the level of demand that you expect for the product. So let's say you're Apple, you're launching a new iPhone model. If the iPhone model is just a tiny increment, tiny deviations from what is already out there, and maybe it's much more expensive, maybe the model will tell you there's not going to be a lot of demand for this because this is not that much term relative to the price.

But if it is a model that has a lot of new features, it's much more user-friendly, for example, then that could lead to a big demand increase and it would be obviously good for the company or for the retailers to plan accordingly.

You can also use this to decide on a price point, right? So saying given the features of these products, the non-price feature of these products, how should I price it? If it's just a small increment relative to the other options available, maybe I cannot afford to price too high. But if it is something that really consumers view as a big change and an improvement hopefully relative to the competition, then a retailer or a manufacturer might be able to charge a higher price.

Hal Weitzman: Okay. Well, Giovanni, thank you very much for coming on the podcast and telling us about how you measure vibe.

Giovanni Compiani: Thank you very much, Hal.

 

More from Chicago Booth Review
More from Chicago Booth

Your Privacy
We want to demonstrate our commitment to your privacy. Please review Chicago Booth's privacy notice, which provides information explaining how and why we collect particular information when you visit our website.