I Built YouTube's Algorithm - here's What I Got Wrong
About this episode
Jing Conan Wang spent years at Google Brain optimizing YouTube's algorithm for engagement, and discovered it was making people miserable. We dig into why, and why the same problem is repeating with today's AI systems.
The moment he realized optimization metrics ≠ user happiness Why modern LLMs actually lost personalization compared to older recommendation systems How RAG systems are missing the personalization layer entirely The "context modeling" framework that could fix it How high token costs change the entire go-to-market strategy for AI companies
This is essential viewing if you're building AI products, thinking about business models in the era of expensive compute, or wondering why your RAG implementation doesn't feel as personalized as YouTube recommendations did 10 years ago.
Full transcript
Can you describe what was it like working at Google Brain? Yeah, they
have a lot of top tier researchers
and you are one of 'em. Uh,
I always had this imposter syndrome. Should I be there? Not me. And my colleague was The, the people who come up with the standard. Recommendation framework that everyone is using is called two stage recommendation framework.
So you also contributed to the addiction algorithm of YouTube. Then
I think every metric has its own purpose until some point.
How do you measure that?
that actually, the whole experience needs me to think about what is the right relationship between AI and and the human. What is the right reward?
Is the amount of time people spend on the app the right reward or not the right reward?
Hey everyone. Welcome back to To ai. I'm super excited about today's conversation with. Jing Wang, who's building some fascinating AI products while also running a hundred thousand member founder community. So Jing, thanks for [00:01:00] joining us today. For those who don't know you yet, could you give us a quick intro about yourself?
Yeah, absolutely. Thank you very much, Anina, for inviting me to the this podcast. Really great to meet all of you, my imaging. I'm currently a founder of Deep Vista Proactive Agents for busy people, especially for founders and business owners. So the problem we're solving is that we all have a lot of messages and communication we need do every day.
In fact, once available Will got that. Every founder, business owners, we have 127 message every single day. And that's not just message, that actually small decision you need make. Either you commit on a meeting, commit some resource of the company. That all requires thinking and there many of the times we need, we wish that there are some assistant or partner helping us to think through that problem.
Previously. We have to find either really awesome assistance or co-founders or partners helping us deep insights to make [00:02:00] that process easy. That every one of us will have a proactive assistant, which is AI that is available 24 7. To help you thinking through that, to prepare the message dropped for you and to be your psych.
So that's what we're doing. And also a little bit more about my background and how my journey needs to what I'm building here. So throughout my career I've been AI researchers. I actually got a PhD degree of reinforced learning in 2014. That was very early. And after that I joined Google to first work on layout as recommendation systems.
And also I was, because I was very passionate about reinforc learning and I joined Google around the same time as DeepMind. So I closed follow layout after go work. By the time of 2016, DeepMind had this like very great PR like event with AlphaGo beating. Which [00:03:00] basically changed the whole field. Like it was not that popular field before the, the whole reinforced learning field.
But later, everyone want to try reinforc learning and Google also want to try, so they wanna find someone. Who knows Reinforc learning very well and knows how to do recommendation system and how to make money. So I happen to have all those experiences. Then I moved to Google Brain to lead efforts of applying reinforcement learning into recommendation systems, and that's where actually I started to go deeper really into.
How to help to build like a system or chat bot help people. So actually how to improve a lot of the Google product experiences from one round to multiple round in which that, if you think about a Google product, is really, you type some query, you get some results. But for many cases you will not get the results you want and you will go back to type the query again.
That actually will happen multiple rounds and forms a dialogue [00:04:00] and that dialogue has a goal. So we, I was basically using reinforced learning to help align the AI and humans in this goal-driven dialogue. And that leads to me like a lot of thinking about what is the relationship between. AI and human, but actually I started like a very long soul searching journey on that.
The first thing, because when I was working on Google. I realized like a lot of AI was really grab people's attention, like YouTube home recommendations, it actually video you a lot of addictive content, like some content people love to watch. And so I personally felt that I want to make it more useful and people more feel more meaningful.
And so I later left Google to work on startup on that. And so DFI is actually my first step already, but the whole journey I have been through is that I learned that in order to build a really awesome partner, which is [00:05:00] what DFI is, this partner needs to be know your personal context very well. It knows very well and it deeply aligns with your interest.
Yeah. Is for you, is helping you and it's very proactive.
Yeah.
And so it's not like waiting for your instructions and that eventually leads to Di Vista.
I think you pretty much answered all my questions for the podcast today. Yeah, in one shot. Okay. You had amazing experience. You worked at Google and not just Google, but Google Brain, right?
That's like Mount Olympus for AI engineers or engineers at the time. That's the dream job for everybody in tech space. Can you describe, what was it like working there at Google Brain?
Yeah, that was actually such a fantastic experiences. I went there end of 2017. There was time, there was a lot of interesting thing happening in Googler when I was there.
I was thinking, oh, that was the bail naps of my time because bail lab was. [00:06:00] Previously a very famous industry research lab, like at a six Nobel Prize. It was very interesting experience because back then Google has tons of money. They spent I think four or five billions a year in pointing to the research.
That was a lot of research was very fundamental research, not like building the building like application. They can make money and because of that they have a lot of top tier researchers. Mm-hmm. So the people who sit not too far away, me like. The father of carers like, like the people with more than 20,000 citations, there's like hundreds of them like in the same floor, and you are one of them.
I always had this imposter syndrome. Should I be there or not? Should better because I feel like they're all like so smart. So first of all, for me, it feel like a university. Because I had my original like schools when I got my degree of P reinforced learning. And then the first three years at Google I was working very engineering [00:07:00] focused organization as organizations and it move very fast to build a product.
And for them experience actually gave me a lot of P of building a real engineering system and product, which is a little bit different than the research or build a research system. That research system focus on theoretical innovation, whether that thing is interesting, and the engineering system is usually focused on whether it has practical value and it's scalable and that has different mindset.
And when I went to Google Brain, I look at the people around me. I feel like I couldn't offer more than the more atic value than they do. They're just like so smart. But finding my own position in the lab. Is that I can work with smart people, learn from them, but I meticulous take notes. Learning from that, I help them to turn that into real world applications, which I absolutely enjoy doing.
[00:08:00] And I also turns out that a lot of researchers, they also love that they love to talk with someone who can understand them. Because of my research background, I can understand what they're talking about. And they also wanted to talk with someone who can actually turn their ideas into some practical results.
Mm-hmm. And
that's one thing I learned because a lot of the researchers, they were publish papers at the top tier papers, neurons on other top conferences. But when you are hitting like Google Brain research level, you'll have a lot of such papers anymore. So it's no longer interesting. To them, you want to be applied into a real product.
And in my cases, I was actually leading the applications on applying the reinforc learning research into YouTube home recommendations YouTube search as recommendations. So that's different applications. Mm-hmm. That we are trying to apply the technology. And so [00:09:00] that actually needs a very interesting and very.
I would say beneficial, healthy dynamics of research and development. So usually the researchers who, they will do some fundamental research that is driven by curiosity from them,
right.
And they will have a very regular composition with me. And I will also be like talking with them to help them, to guide them the research direction.
To some direction that is useful to the company and to the industry as well. And I'll also come up with a lot of the framework that is useful in the real world. For example, if I find some new interesting ways of training model that actually get great results, I'll actually meticulously take notes of that.
Mm-hmm. And we would have a lot of discussion on that. And then we'll write paper together. Hmm. After that we'll also go to talk with the product team, which is completely [00:10:00] different group of people. Now. They are things like how to drive the business metrics, and I will translate everything I learned from researchers and bridging the gap between these two group of people and.
Then what things practical, what thing is still like I need to make a judgment call about what thing can be implemented now, what is like quarter from now, a year from now? And then creating kind of research roadmap for the product team. And so the product team will looking at me as the face of people from Google Brain, um, for that the product roadmap.
Yeah.
Yeah. I'm wondering if you're the person who they can understand. Because you can, you also speak their language as well. You take the sense to the research team while also take the research and interpret it to the product team. It
is through the whole experience. I realize chat is such an important capability [00:11:00] for any jobs, like for researchers, for engineers.
Like it's a way for humans to convene their ideas from their mind to another person's mind, like chat or conversation or podcasts like we are doing today is the only way and. Different people would have different operating system like researchers. They were operating in very long-term, driven decades and human civilizations, they were talking that language.
Right. And
then the product team will talk like, what is the OKR this quarter? How do we grow metrics by 10%? How do the performance management if, and I have seen many cases in which they are in the same room and they were extremely awkward and basically not in the same spectrum. And gradually I found that my positioning is that everyone loves talk with me.
When researcher talk with me, they feel like I'm one of them. I understand their concepts. And when product team talk with me, they feel that I can give them [00:12:00] clarity rather than it being a LA of land of five years, 10 years, what's going on in the industry.
Yeah. Think the different language. Yeah.
Yeah. So I essentially was a translator for, for the whole research and product teams.
Yeah.
That sounds like a fascinating experience. I'm sure a lot of our audience would love to do that. So for those who wanna work for Google Brain, you know what, it takes at least either research background or you have to be the interpreter, but you also contributed to the addiction. Algorithm of YouTube then could you say you're working under the reinforcement learning for recommendations?
Yeah,
I think every metric has its own purpose until to the point they all optimize it. And so that's whole concept of we wanna optimize user engagement, like this is applicable for. All the apps. And there was one point that every AI resources, like that's field recommendation system that improve engagement, which is make users use the app more, love us more.
So that would be the law style metrics of [00:13:00] every company at some point. And in the beginning it will, there'll be alignment of interest between users and algorithm because if you make the app better, then of course users gonna stay longer time. Over time, if you over optimize something, the interest might not directly aligned from that, and that will lead unexpected behavior, which is, oh, people spend too much time on computers, too much time on this app.
So this relates to what I was working on, like one of the project I was leading to improve the engagement of users, how, how much time people spend on YouTube and we will set an optimization metric, the reward of the. Reinforcement agents to optimize for, and that is how long people would be engaging with the app.
But at one point I, I realized that as we were optimizing it, people's happily actually go down.
How'd you measure that?
So we realized that are a lot of users [00:14:00] who actually will uninstalled app and for a while, and then we reinstall app. So that's basically mass behaviors. People feel guilty, they spend too much time on this app and they're not happy and in some.
Far mode, they will just uninstall the app. But then after a week or so, I never realized I still need this. I couldn't do go with this. So we reinstalled it. I personally also did it for many times, not for YouTube, but for TikTok. I, I uninstalled and reinstalled like at least five times. So that's like indications of people are not less strongly happy about what we do.
But that actually the whole experiences needs to me to think about what is the right relationship between the AI and the human, and basically the what is the right reward? Is the amount of time people spend on the app is the right reward or not the right reward. In fact, there was a triggering event that made me start to think about my last journey.
After Google, we were working on this [00:15:00] YouTube weekly review and we are looking at the law style metrics we're optimizing, which is over time the user spend. And we realized that we have about 10 engineers, researchers in the room that determines the algorithms. A lot of the decision we make, you are changing what people will see on the YouTube home.
That's like what billions of people will see on the YouTube home if we push a button and push it. Like the buildings, people's homepage were changed and that's while they spent like a few hours a day. And I want that experience to be more, less data driven, but more relatable to actually every individual human.
And I was actually looking, I was doing this and I was looking outside the window. I saw daddy. Was walking together with his daughter on the street and it was a very lively family things and so it's everyone's individual life. It what's happening outside [00:16:00] and inside it is like huge table. Like in Google we review because it's very complicated.
We pull out this dashboard and every algorithms, there'll be like, like 300 metrics and looking at this down, up, down, up, there's such a big thrust contrast between the two. Yeah. And, and I was thinking about, I already done a lot of this, these reviews and numbers and all these things. I'm no longer too interesting for me.
But if you look at outside at this family line, how can I create a product that I would feel very proud at the end of my career, right? And how do I create a product that make everyone enjoying their life and being happier? So that actually needs me whole searching process after leaving. There's many reasons I decided to leave Google, but that was one very memorable event.
For me to think about. And then the first thing I actually do after I leave Google, I think [00:17:00] maybe it's because of ads, because I come from Google ads and as I actually did a lot of improvement in ad system over credibility of revenues on that. But then I realized that maybe I wanna work on a different business model because one of the reasons people wanted more engagement is they could have more as time.
For them to show ads. I, I think that's like industry side and that's no good or bad of this thing. Everyone's doing that. But the better way is actually a business model that would align people's interest better. In ads model the advertiser pays you give the product for free to your users. In return they'll get attention.
And this way is created this three party system that is not fully aligned in terms of interest. So I started to think about can we do sas? Which is like subscription based, which is that I provide you the best services and you gave me the money in return. That will be much cleaner business models to do.
So actually after I left Google, I joined the SaaS [00:18:00] startup and be the head of ai. So that's my whole journey and then all startup I doing has been like exploring this, these models is basically the idea is that how to build a useful product. So that people are willing to pay directly themself.
How's that different from the ads model?
In
the ads model that basically like your users will use your product for free and they will give you attention and you will sell this attention to advertisers, which will actually pay for it, and it works very well. And this is what I learned later after I have been doing several startup. It works well for the internet time.
Because the cost of serving a user is very low. If you have a user coming to like google.com or facebook.com, the cost is almost zero. And because the cost is zero, the best way, the best strategy to do is to give the our product for free. And then you [00:19:00] can maximize your reach, which is get a maximum number of users.
And after you get all this attention, you can resell this to the advertisers from that. This model doesn't really work in AI applications because the course of token is very high. It's definitely not free. And so because of that, it is actually not very wise for companies to chase for growth too early on.
That's like different people have different opinions on this, but I have a, one version of my opinion is that the go-to market and how you grow the company should really depends on. The economics dynamics in your systems. The current wave, the Gen I, we are in the token dynamics, which all the, the business strategy of application companies is driven by how expensive the token is.
And then the token is very expensive, then is actually very important to [00:20:00] have a solid business model, actually important to chase high value problems first, rather than chasing more people. Which means that I actually don't quite agree with a lot of the vibe revenue companies in terms of they are chasing growth in the sense of they're actually selling the goods lower than the cost they're paying for cloud or the,
yeah, it's really hard, right?
It's a competitive space and then people are racing to the bottom in order to chase growth. Like in terms of user account, right? D-A-U-M-A-U, those kind of things. You are saying that. Because there is high cost associated with token usage for offering such a product to users. We should be focusing on high value.
So you, we really should add value to the user.
Exactly. Yeah. That's the reason why for our focus on busy people, founders, business owners, and high value use cases. And to your question, I think it is. Everything, every [00:21:00] decision or every framework needs to be considered in the particular context of the technologies constraints or the timing.
Like what I'm think now is like at this moment, it's actually not a good idea to chase it because the talking course is still very high and I have a one yard stick to check whether talking cost is high or not. If you think about the different type of coupling that we are having, there's like hardware company.
Cloud company that's providing computing power. There's a middleware company like long chain providing tools. Yeah, infrastructure tools. And there's application company, which is providing application to any users, deep voice like the application company. And if you think about this is like the, the whole system is the value state in the long-term, steady state, the really the, the layer that provides the computing, including NVIDIA or cloud providers.
They should have very same magic if you, if we electricity, right? [00:22:00] Electricity like we should, like P, does PGLE make any money? They're not. They are a public utility company. They provide this electricity to us so that we can use that electricity to do great things. And this is because they're not making money.
The whole society. Able to create more things. And if you think about currently Invidia get like 75, 80% of margin, gross margin. I think we're still too early. It is more like when the I Electricity was just created in 19th Century. They also have it a 70, 80% margin. And then the only company that is making money is the company who making electricity.
And in these cases. It's actually very important to chase, like high valuable use case. One of the example for this is there is a ly garden in some material. I went there once once and there were one of the very first families had a A light bulb. Light bulb. [00:23:00] Yeah. Back in a hundred years ago, that was a luxury.
Now everyone has it, right? They're very wealthy people and so that's also the high valuable use cases. But when the margin of the electricity company goes low and everyone will get a benefit on that. So the ya stick I mentioned earlier was how much gross margin Nvidia has. I will be ready to change the business strategy whenever I see.
Earlier earning quarter gross margins low.
What are you gonna do when it, when public announced that the margins dropped?
Now we serve wider audiences because I, for D Visa, I mentioned that we're primarily focused on busy people, business owners, founders. We focus on them also because I also run founder community called Founder Coho that has, we have 120,000 followers, mostly founders in Bay and Globe, and they're very high, valuable.
People to serve for. So that's the main strategy. It's like building awesome product for small group of people. [00:24:00] That is based on some of the situation that the talking is very expensive. But whenever that has changed, this Divis tool essentially can be your personal system for everyone. I think that will also change the, our business strategy from that
is your business strategy today.
Also, you know, looking for growth. Your last startup story, tell.ai. I remember reading your blog post and saying that you grew to 550 K in one year users. Right. So that one sounds to me that's the past. Are you looking for the same for this one?
We have a a couple multi stage growth plan. The first stage we actually were laser focused on high valuable.
Customers, which is business owners and founders, and the founder, Coco, essentially already formed a community from them. So in terms of growth, we can just grow within the community initially, but later we'll expand it, but expand [00:25:00] it to other use cases like sales team who need a smart assistant to help them to handle all the communication flow.
So there will be, so based on also there was about 500,000 founders a across the globe there's like about 2 million executives in growth companies that can potentially be our users. And there are also 5 million business peoples in bigger companies can also be benefit from us. So that's, different tiers will grow and eventually we can be an AI system for everyone.
But that is like really turning a vision for us.
How's it different from chat GPT? If you're saying that's, that's everyone,
right? Yeah, that's a very great question. In fact that when we look at the chat, GPT, a lot of people use that to handle messages. But we find that like through the surveys of the many users we talk to, the right now, the experience of child GBT is [00:26:00] problematic for several reasons.
The first reason is that. It's still very transactional. Although chat beauty has a global memory, the memory is very simple. It's flattened a couple like list of texts. It doesn't really understand you deeply, and so every time you open a chat, write messages, it asks you like three questions before you can even get the initial responses and you answer those questions and in closing the open new conversation it'll ask you again.
So that's one is transactional. The second thing is that a lot of the chat GP responses. Still not personal yet. It is like you still feel the AI is not you because it doesn't know a lot of your personal context. And the service problem with chat DT is because it's chat app, it's not a task tracking app.
And for business people. The reason why they send messages to move forward for the business side, they need someone to follow up on the items. So they really need these task tracking capabilities. Someone [00:27:00] remind them when they need to do sentencing, so chat to me doesn't really do that. And so we really differentiate us.
With chat DT. In fact, we actually use a lot of the LM API form open, I form cloud where application company, were not building foundation models, but what we do that those big company wouldn't do is we build your personalized context model, which doesn't need to be a large model, but it is a small model that restores all your contexts and are able to generate.
The right prompt for large language models to execute. And the framework we are having here is that in, in order to build a personalized AI system, we actually need to have two layers. The bottom layer, which everyone talking about is executionally. If you really think about our large language models currently, it works very similar like a interpreter, pricing interpreter.
We have this context. Context are a sequence of instructions. Which is [00:28:00] large language. You send that instructions into our, and you get things done, you get responses. And now you can also get a tool calling and get the things done. And this work exact same way, we'll send a Python program into Python interpreter and get it done
right.
Makes sense. So that
execution layer, and we'll leverage that capability as much as we could. But on top of that, we still need a layer, which is who is writing this Python program? Right now everyone's saying, oh, they're doing either prompt engineering, which we expect that everyone's just being able to type that program by themself.
Yeah. Louy prompts. Yeah.
It's not easy. Yeah. Or another concept, people I are very popular right now is called Context Engineering. It's basically creating a set of rules to create that program. So that program essentially is a couple of rules. What we believe is that problem engineering context engineer has own.
But the main problem there is they're not deeply [00:29:00] personalized. If it is like you require people to write prompt, most people couldn't. And if you rely on rules to generate context, engineering context, it's the same for everyone. So what we are building is this personalized context model, but deeply understand the past interaction.
Between the users and agents, just like our part like recommendation, understand your interaction between you and the, like the YouTube it understand every chat conversation you had with this agent, including other data sources. And it translate those contexts based on what your type now into a set of instruction is.
Writing this like on the fly small program and that program is essentially the LM context. That will be sent to the LM to execute it. And it also gave us a very strong positioning is now every time opine release new model or cloud release, new model, we're very happy. I'm like so happy for that [00:30:00] because we are working on top of them.
We can easily adapt our model to write better instruction and which can be sending lm. So the analogy will be we'll always have this capability writing a good pricing program. And open eye and cloud are building the best pricing integrator and if they upgrade, like we're absolutely super happy.
So you're saying your personalization is through on the fly context engineering for the user based on their past interactions.
Because you still let the user write whatever prompts they're gonna write. You're not using a one size fits all context engineered context for everybody. But you're building that based on real time conversations so that you, I think it's very human as well. 'cause we're friends. The more we talk, the more I know about you, and I always retain that context.
I always know I'll learn about how you work and how we can work together. I have that context and I distill that. [00:31:00] So basically it's like a dynamic context engineering based on personal interactions.
Actually, I wrote an article about that. So this new paradigm, I coined a new name for it. For context modeling compared with context engineering.
It is really a dynamic context engineering that everyone will have different context engineering tactics based their own situation. So compelled with context engineering is really the engineers who are doing the thing. Yeah. So what is the difference of context modeling component compelled with context engineering?
The main difference here is that instead of have engineers come up with rules. To retrieve the relevant context and the user template, putting it together, there will be a model that will dynamically retrieve the contexts for the users that is relevant to current combination, and then not using a template, growing them together, basically digest it [00:32:00] and generated the instruction.
The RM instruction, we talked earlier. Directly, which will be sent to the front end to be executed. And that model is what we call a personalized context model to doing that. One of the main difference that model will do compared with privacy is this dynamic retrieval processes or retrieval models.
We're talking about building a personalized. LLM or personalized AI system. And that ties to interesting historical involvement of ai. Like we used to be able to do personalization very well, and that capability get lost. So before LLM, the model was much smaller. And when we have this algorithm in the YouTube or it is highly personalized, everyone's homepage of YouTube is different.
Everyone's is completely different. They're trained in the real time. But now fast forward when we have LM, every release of the [00:33:00] model is one and a half a year or two years. Like we got a G PT four in April of 2023 and we just got G PT five a month back, which is like August of 2025. That's almost two years, more than two years in terms of the model development.
And this whole model is exactly same for every single person. And why that happens is because the model gets bigger. And so it's getting harder and harder to train. To
train.
And in fact that the note on that is like when the G GPT three world, everyone's talking about with fine tuned model. In the G PT five, nobody's fine tuned.
GP five.
Yeah. Nobody has
the resource to fine tune G PT five. And because of this, as we are having scaling the model, we lost personalization.
I wanna double click on that.
Yeah, absolutely. Yeah.
So pre GPT we're able to personalize like let's for YouTube. And is that model a non lm, like just a deep [00:34:00] learning?
It's non
lm,
it's machine learning models.
Machine learning model, and I can go a little deeper into how that model looks like. And in fact, me and my colleague was one, the people who come up with the standard recommendation framework that everyone is using is called two stage recommendation framework.
It basically have two steps. The first step, and let me use a YouTube as an example. The first step is that let's not focus on. Accuracy. Let's focus on finding any possible thing that you might like. Even if I suggest you, something that you might not like, doesn't hurt. So I will get as many candidate as possible and give it to the next stages.
This is called retrieval. And the second stage is called re-ranking. And now the re-ranking need to handles instead of billions of candidates, they only need to handle the candidate generated by the retrieval model, which usually is the hundreds or like a thousand. And they are focused on accuracy among these thousand candidates.
What are the [00:35:00] top five that will be relevant to you in the current situation and that recommendation system? Is using most of this AI company like TikTok, like YouTube, they all use this framework and why? And the reasons why they use this is because they need to scale the system, maintain the personalizations, and with the model, when you don't have that many users, you don't actually need two system.
But when you have a lot of users, you realize that is actually very important. To have candidate model with retrieval models to be able to handle billions of potential candidates. So you can focus on that, and you can have the re-ranking models. You focus on the personalizations, do the most accurate re-ranking.
When you separate out these problem, your system can scale,
and
then that's how we are able to, for example, like in the TikTok or the YouTube, billions of [00:36:00] users train continuously. Everyone have their deeply personalized homepage. This is through this architecture.
And you're saying this is no longer true.
They're not using this model anymore.
That's a great question. This whole industry is making this true again, but initially it's not true for a reasons. So when the large language models comes out, it solves a much generic problem than recommendations. Like this personalized recommendation system. They can do depersonalization, but they can only do one thing, which is ranking.
You have a lot of YouTube video, which five relevant to you. They can solve this problem very well, but LM is solving a different problem, is solving generic q and a is. I can type any inquiry and give any responses and people love that. And when the LM gets creative first and people were directly chat with LOM, and [00:37:00] in these cases, nobody need to worry too much about scaling or other things.
And so they were directly doing that and the performance will be best because this is the end to end solutions for that. And over time people realize that they need to have some additional capability on topic to. Process much more information in real time fashion. If you are only chat with lm, if the LM knows it, it's very powerful.
Great. But if you chat about a topic with LLM, most of the time it will hallucinate,
right?
It doesn't know this information, but it will pretend it knows this information, and we call the hallucination from this. So hallucination has been. Really news for the LM. It hasn't been solved even now today because LM only know a subset of information and whenever you go out of the boundary, it'll start to rambling.
And later everyone is trying to build this AI G Rail saying [00:38:00] Don't say things you don't know, but refuse to answer your questions. And so people are building this additional system on top of it. To be able to handle data in real time, which is now rack retrieval generated augmented generation. And if you think about a rack, it is actually work very similar to the recommendation system.
Talk about,
it's the retriever,
it's very retrieval. They even carry the same name. In fact, they actually use the same technology.
Okay. Like most
of this like pine corn or this like vector databases? Yeah. For the rack based applications. So people started adding this back. But the current situation we are in is that when we talk about, like when we look at this really the rack or the Lina retrieval, augmenting generations, we have this LLM, which is the second stage actually works similar.
Like Rka now is LM, but it's G five. Nobody can do anything for it.
Right.
And you have the first part. [00:39:00] Which most people are doing this called Science Similarity based retrieval. And so they will run this like dog product and get a top K and then get a result as in the lm. And that rules is the same for everyone.
And if you think about this two system, there's no personalization at all. The first stage is the same for everyone. The second stage is the same for everyone. Then how do you have any personalization? And that also cause another problem, which is like. Everyone saying application company don't have mode.
Of course, you build it like a system without any personalization. So like no matter how much data you have, you don't actually improve user experience. Now there's no data flywheel in this whole current rack based applications. So we are in an interesting situations in which that everyone is chasing more data, trying to grab it and hope that they can use it.
But no one is actually using it. Yeah.
Very interesting insight. So are the [00:40:00] recommendation, the origin recommendation engines, going back to the earlier algorithms like you two? There is
actually, it'll never go back to the same algorithm as before, but it will involve just like the, in the later, there was many like species involved with similar forms.
Because the environment needs that type of things. I think it involves in similar fashions, which you will have the retrieval system. Is focused on being able to handle things in real time, handle a lot of more documents, and you will have lm, which is very powerful and focus on accuracy that is still coming.
And then we, in the process of doing that and what is, what I think is next, which is also related to context modeling, is that it used to be when you do build a recommendation system, the retrieval staging, the recommendations is a model is actually called retrieval model. Deep retrieval and it is highly [00:41:00] personalized model that learns from user feedback.
So give back to the example. When you have this retrieval modeling, like YouTube home recommendations or TikTok recommendations, it actually learns from every single action you have. Like you like it, how you watch on this, and based on this feedback, it will give you different candidates. And which will be sent to the re ranker, which is also personalized or learning on this.
So basically that's two layer of personalization happening for, and then fast forward to our rank based applications. Neither of them have any personalization. Nope.
So the original version of this is. Because you are training based on my action, it's already personalized towards my action, right? That will guide the retriever.
'cause the retriever is modeling against my action. So the retrieve the results will be personalized versus the rag approach is just co-sign similarity. So based on my query, which one is similar to my query, right? So there is no feedback in terms of my action. And the same for [00:42:00] the re ranker as well. If you don't.
Put these parameters in it explicitly, then it's not working.
Yeah, exactly.
Yeah. Thanks for explaining this. You are having this context modeling for Deep Vista. Can you speak about the first use case, since you're not rolling out to everyone yet, what does that actually look like in practice? If I were to sign up today and what's gonna happen.
Great question. It actually tied back to the conversation we had earlier.
Yeah. In
order to build a great company, you need a great research, which we talk about context modeling and as deeper research happening, company for that. But you also need a great product and great engineer, like excuse experiences.
Now let's switch gears to the second part. Go back
to that. Yeah,
so the first initial use cases that we're solving is for busy people, especially for founders and non for busy people. For basic people.
Okay. And we
wanna help them to reach inbox zero all the time after. How is that possible?
Okay. Tell me exactly.
That's a great [00:43:00] question that you need, like AI agent that really first deeply understand you. The second are very proactive and in fact that a lot of founders and business owners, they're already doing this. They're doing this by hiring another people doing this like consecutive assistant or consecutive admin.
A lot of founders I know, they hired this and they were trained their executive assistant to help them to manage their inbox for them, and it's very costly. If you do this in the US it's easily more than a hundred thousand a year and even you do offshore. It is like easily 30,000, 40,000 a year. And so if you use our product, you basically will be replacing this executive system.
So we also AI executive system for busy people and the user experience. When you log in the our product, the first thing is that you will like after you authorize with your emails and will later focus on email first, and then we'll start with Slack and Discord and it will start to [00:44:00] proactive processing all the past unreal messages for you.
Based on the context you provided and could also give guidance, just like how you were trained, your executive assistant guide, oh, I like this email. I, I think this is less important. You can have conversation on that, and then your agent will use the instructions you gave to basically classify all the emails, process them for you.
Our agent will basically ask a few questions initially, just connect information. Who are the important persons? I care about what topics I care the most, what topics the agents can handle automatically. The Vista agents will proactively handle this and 80% of the messages, and they will archive it for me or will do a quick reply automatically, and there might be like 10% or 20% that needs my attention.
Even in that cases, either we're pulling all the contexts and writing a draft for me so that I can just like build on top of it. [00:45:00] And as I'm doing more iterations on that, I'll send the feedback. And so like in the future for these cases, you should handle like this. And this is how I would write emails for this situation.
You can actually give instructions to our defi agents to do that and they will learn that. And in the future, we are just pretend. They handle the email the way you would be handled, and eventually after you finish wholesale, you will get inbox zero and then it goes to the next second stages. When you reach the inbox zero, it actually proactive monitoring for you.
We'll keep using inbox zero state, so every morning you log in, it'll give incept summary. Everything the Vista agent has done for you. If you, you got a hundred messages in the, the agent will handle 80 of them and give you like a draft for 10 of them for you to process in the morning and also will give you a proactive suggestion what is focus for that day.
Is it, is it [00:46:00] integrated with like my Gmail so that I can just one click to send the response over?
Yes. Well also take action for you when you do the inbox. Zero. Your agent actually go to the Gmail and actually archive the email for you. And also in the future, when you need to send the email, you don't have to go to the email, you just say, please send it.
Okay. And the agent
will send it for you. Yeah.
I can tell you some of my S one is some people wanna schedule a meeting with me. Can they actually go to my Google calendar and actually make it happen for me? Is that action
exactly be
taken as well? Okay, awesome. Exactly.
Yeah. And in the future. It's not just in the Google Calendar email.
We also have the plan for Slack and, and like discord LinkedIn message. So it'll be all in one inbox that you need handle, but you don't have the login to LinkedIn to reply your messages or, or Discord or Slack it. Just talk to the agent, send this message for me and it will go as send for you.
I [00:47:00] think that's exactly what I'm looking for.
So I get maybe 50 per day on emails, requests, asking for sponsorship. I just wanna respond to them. Those agencies like Head or AHA Labs, I don't know if you know about them. This is part of the PLG growth hack, right? I have tens of thousands of different emails from the same company and I cannot filter the emails.
90% of my inbox is like those kind of spammy emails. It's just a pain. This is why my inbox is not zero. Right. What they're optimizing. They're optimizing to hit my inbox and I'm trying to like, I don't wanna see you, but I don't know how to spam you. You know what I mean?
Yeah, exactly. Yeah. Yeah, yeah. Yeah.
Thanks for great, great help that you, you are my perfect ICP ideal customer.
Are you open for beta testing?
So yeah. We are open for the sign for waiting list, the domain, deep Vista ai. And we are rolling out the [00:48:00] testing, like the offer testing later this month.
Okay. Awesome. I feel for the stuff that you wrote before, your blogs and your LinkedIn posts and things like that, you mentioned this.
This term appears multiple times, so signal from noise, a lot of your philosophy is around trying to increase the signal level from a lot of noise. Is that still true?
Yeah, that's true. That's true. Yeah. I'm really a productivity person. I'm very fascinated about all the tricks and ways to help people to gain their time back and get more things done, get their dream done, and I think one of the main problem we are facing.
In the modern world is that there's too many informations for us to processing and like you mentioned, all the startup are doing is really focused on that signal from noise. And because this is a problem that we modern people face all the time. In ancient time, people don't need to solve this. There's not many [00:49:00] communication tools for them to share the information.
I really think the signal noise ratio problem is a big problem for us, our generations, because we just finished one revolution, which is internet revolution. We are going to this AI revolution. What I see, this is AI revolution, is anti-D dose of internet revolution. The reason is that. If we think about these tools, the technology we have in the past century as a continuous spectrum, we have first like computing devices.
We have computer now. We have personal computer, now personal computer connected because they wanna have more communication with each other. As we're doing this, there's more and more information generated and more and more information communicated and where in a world everyone knows everything. Like we have the YouTube influencers, like millions of followers or tens of millions of followers, and this trend is still going on.
What [00:50:00] that means that one person message does a million people receive it, and we're all subscribed to thousands of influencers. And that's just for this online influence channel. And there's also email channels we talk about. There's every a thousand people trying to reach me as well. And then there's DM channels as well.
So. It's a very common problem because of how convenient it is to communicate with other people. We have this social media, we have these emails, we have Im, it is so convenient, communicate with other people thanks to the internet revolution.
Right
now we have a problem of steroid ratios to no. Which I think the AI revolutions help to solve that problem.
We will have the best AI assistant to really help us to make sense of the world.
You're saying that it can be a filter, like it's a lens we can use to build the world [00:51:00] so we get our sanity back in a sense.
Exactly. So that's what I say is this anti-D dose and maybe one day. We have such a powerful, I think that the technology involves in the cycle.
So because we have so many information, we need the future or personalized assistance helping with that. But later we might wanna connect more directly with each other, not with the first. Then there are maybe another like revolution to help people to connect more in the future. So there's like this more signal noise ratio and solve that problem.
That's create another one. It's just like. Cycles going on.
It's too much noise and people wanna reduce it. And if it's not enough noise and people want to hear more noise, it's like
we all, we also do the same, right? Sometimes we have go to too many events, rest for a few days, but when we actually take a rest for a few days, I wanna go to more events.
Okay. Yeah, exactly. I love your community events. I go to all of them. So strong supporter of all your events. Oh,
thanks. Yeah.
But for all my audience, [00:52:00] you're local at Bay Area. I think Kojo has really good events, so please join. It's very fun.
Thank you. Thank you. I
think you hit on a really deep topic today.
I think you mentioned about two different mindset about building products. One for optimizing for metrics and for engagement for money. For revenue, right? Yeah. We cannot build a company or business without that. So that's still important. And then the second one is adding real value. That's good for the human being users.
So I feel like you have a very empathetic approach of building product. And when you mentioned about the contrast when seeing a family living their daily life versus tons of. Metrics on the screen. I feel like I almost wanna cry. I can visualize that. I can see that. So. Do you think this is the reality?
Do you think building products are [00:53:00] trending towards rethinking to add true value for humans and we're rethinking this machine human relationship?
Yeah, that's a Thanks for this such. This is such an awesome question and thanks for asking this. My whole philosophy building product is evolving as I'm doing this whole journey and when, talk back to the story that like I shared early.
In the YouTube. That was 2018, actually, 2019 when I saw that thing. And then I don't know how to make that happen. And so there's a whole soul searching journey for me. How can I build a problem people truly love, not just have good metrics? It took me some zigs out journey to do that, but at the end of the day, I realize the only thing that will help you to guide through.
Whether you build a good product or not is actually real world in action with your users. See them every single day. It's not through any metrics. I have to change my philosophy a little bit because I was like, had my PhD [00:54:00] training, I was doing the algorithm. I saw, oh, you can always come up with a set up, a dashboard, and I'm seeking here on the dashboard, okay, this measure is going great and I will have a great product and gradually a job benefit philosophy.
That also leads me to make more friends, to build a community of the, because I realized that the purpose of building the product, the true value is building a product that people love. This is fundamentally about people, both for the builders and for the users, and you have to have this like conversation and the actions.
To them and there was different processes that the guiding through the whole thing. Like the first thing is that you have to talk with enough people, say now, facial expression, not just metrics. So tied to back the story is like I had is that on the one side is some numbers on the other side is people,
right?
I would rather than I build a product and talk to a hundred people, they are smiling rather than I build a product. The metrics are [00:55:00] fantastic, but no one's really smiling. Form that. So the philosophy I build is I find those group of people first, and I'll make them smile.
I love it and
how I make that happen.
Yeah.
I love it. Thank you, Jane. I really enjoyed the conversation with you today. For anybody who's looking for a product like a AI assistant who can take actions on your behalf and can be personalized to your needs, I would consider use in the future. I would love it to, to come from. Very empathetic founder who cares about deeply if I smile or not.
So yeah, really a pleasure to have you here today, zing.
Thank you. Thank you Angelina for having me, and really great pleasure.
Sounds good. I'll see you next
time. I'll see you. See you next time.