Why Most AI Agents Fail in Production (and How to Fix It)

2025-08-27 47:34 Guest: Yann Bilien Watch on YouTube

About this episode

In this episode, I sat down with Yann Bilien, co-founder and Chief Scientist of Rippletide, to explore the cutting edge of AI agents and autonomous systems. With a background that bridges advanced AI research and practical deployment, Yann shares how Rippletide is building scalable agent architectures designed to handle complex workflows with reliability and precision.

From the challenges of orchestrating multi-agent systems and designing robust evaluation frameworks to the importance of grounding, memory, and feedback loops, Yann offers deep technical insights into what it really takes to move from demos to production-ready autonomy. He also shares Rippletide’s vision for how agent-based systems can transform industries from knowledge work to enterprise automation.

Whether you’re an AI engineer, researcher, or founder curious about the real-world potential of agentic AI, this episode delivers a thoughtful look into the future of autonomous systems and the scientific principles driving Rippletide’s approach.

YB
Yann Bilien Rippletide Chief Scientist

Full transcript

Yann Bilien00:00

My name is Yan and I'm the co-founder of Reconcile. So I was doing AI research and now we are working on, uh, autonomous agents.

Angelina00:08

Let's say I'm your client, right? How do I get started?

Yann Bilien00:10

I would ask you the, this, this question is 80% enough for your agents. The biggest LMS now just have seen the same amount of information of a 4-year-old.

Yann Bilien00:20

So we will soon be missing some data to keep the improvement rate of language models. Regarding this compute is not the limiting factor that is this evaluation scheme So far has been a challenge in AI and it's not solved yet.

Angelina00:37

What do you think of a very smooth transition from a demo? To a more dependable production use case look like

Yann Bilien00:45

when building agents.

Yann Bilien00:46

You have now, typically two steps

Angelina00:49

should companies build or buy for these AI agents needs?

Yann Bilien00:54

That's a tough question. Um, maybe two elements to, to answer that. The first one is an observation, which is

Angelina01:02

Do you think AI agents will ever become truly autonomous?

Yann Bilien01:06

So autonomy is being more and more characterized. When you are in the valley with, uh, way more self-driving cars and so on, but definitely yes, autonomous systems will be next.

Yann Bilien01:17

Next big thing on

Angelina01:18

your website also mentioned that AI agents, you're bringing like from 80% accuracy level to 99% accuracy level. I think it's pretty bold statement.

Angelina01:30

Hi y. Hi Angelina. Good to see you and thanks for the invitation.

Yann Bilien01:36

Good to see you again as well. I met Y at the Gen AI conference and you gave a fascinating talk about AI agents in production. So this is our topic today. Would you mind, introduce yourself and tell our audience what you do and why you working on it? sure. So my name is Yan and I'm the co-founder of Rippleside Previously, so I was doing AI research. I was. At Imperial College in London and now we are working on, autonomous. Triple title. So helping people build them, helping people build AI agents. That's it. We, I started two years ago basically, and so we, we started with the intuition that we could reach autonomy and that to reach autonomy, there is some tech stack that is missing today.

Yann Bilien02:25

And so we started building agents, selling agents and so on, and we built the technology to what was missing. And I'm sure we talk about that. And so today we are selling this technology for people to build agents. Awesome. I wanna know what's missing. I wanted to ask you like what is, how old is Ripple tide, but I guess you told me the two years old, and on your website it says it mentions a lot about AI sales agent. So can you tell us a little bit more about what is an AI sales agent? Yeah, sure. So we, when we started Ripple Tide, we wanted to really go forward regarding autonomy. And so to start doing an agent, we started building a sales agent. Why a sales agent is because people are involved in sales processes.

Yann Bilien03:11

And so that means that's where we can express autonomy in a very good manner because we cannot know in advance what will happen in the situations. So we started building this sales agent. we made the r and d we, we developed the tech components to do such agent and. One day people started, with the hype of agents to every day building themselves agents.

Yann Bilien03:34

And so they told us. Can we use the technology you developed to solve the, to solve it and to do autonomous agents? And so now we are also serving this technology to other people. Very interesting story. So you're saying a situation like you cannot predict what's happening. let's say we are on a call together right now. you cannot predict what I'm gonna ask you. I cannot predict what you're gonna say. So is that a perfect scenario for using the AI agents? Yeah. that could be. And I think, maybe we can start from the definition of what is an autonomous agent compared to what's existing today.

Angelina04:08

Because what's interesting is, I guess there are two types of agents, basically one that is more focused on automation, RPA. So you take a business workflow, for example, and you want to do it using technology. And on the other hand, there are some situations where. You cannot say in advance, what is the workflow, what is the process to do that?

Angelina04:31

And especially when humans are involved in it, as you mentioned, I cannot predict your next question or so on. So I will have, based on your question, based on the surrounding, on what's around, I will have to make some decisions to formulate an answer. And so that's where autonomy comes into play. And my, my, my guess is that. It'll enable a lot of new use cases to have autonomy and not just automating some workflows and so on. And so that's where we are working on it. That's a great definition. I think you cleared my mind a lot 'cause I was working on NAN automation, so I guess I'm working on automation, not agents. So that's a great delineation of the two concepts.

Angelina05:11

Can you also tell me what's the difference between sales agent and the AI SDR. I think there are also a lot of confusion about that as well. Yeah. I'd say that a safe agent is just a category of subagents basically. For example, if you rely on the digital worker notion that a lot of people are talking about, the idea is to do a job using an agent.

Yann Bilien05:34

And you have two ways to deploy agents. Either you're trying to give like a coworker to, to someone. So an agent that would work along each person. Or you want to handle entire use cases for a company, for example, you want to relay to do all the inbound leads that enter into the sales motion or things like that. So down those two ways, and I think that the sales agent rely on the category of a lot of different use cases. That,we can discuss. you will have in the sales motion, you have the SDR that is trying to get and book some meetings. You have the AE that is trying to move forward the deals you have.

Yann Bilien06:14

You can have an agent that will close the deals and so on. And so the AI SDRs is the first one, in the line. And I think a lot of people try to build AI SDRs because that's can make probably the easiest one to start with. And so that's why there are so many companies doing that today. It is a very crowded space.

Yann Bilien06:32

And are you guys, the first thing that comes to my mind, it's, I have to tell you, it's very confusing. There's so many AI SDRs or AI agents, companies, and sometimes they all say that they work for the sales team, but it's very unclear to me what exactly they do for, in. 11 x, right? It's a famous company. Are you guys similar? Are you similar to 11 x? so we are not building SDR agents, so people can build a lot of different agents using our technology. And as I mentioned, SDR is probably the. One of the easiest one in a way because you can automate a lot of things that rather workflow you want to do.

Angelina07:09

You want to get some information on the internet, then you want to draft an email and you want to send that and to do that, you don't really need autonomy actually. and the interesting points regarding AI is a good observation. What is the bottleneck to put on agents in such situations?

Yann Bilien07:29

And sometimes it's not only the technology, it can be external factors. And regarding ai, SDRs and all the stuff you, you probably heard about, the, those companies is you can send tens of thousands of emails a day or millions per month that's not a technological problem. With that. the limitation of this is rather. The only variable you have to improve your outcome is volume. So you will need to send more emails every day. And everyone is receiving thousands of emails a day and you can't read them all. And in the end, the addressable market you have is the limitation of such solutions because if you send millions of emails a month after a few months, probably, you have no one to reach out after.

Yann Bilien08:13

So that's not a technological problem. It's rather you want. Variables than volume to improve. And that's where autonomy can come into play. In a way you can improve your outcome by, I don't know, being more personalized, being more respect the company, how you talk to your prospects and so on. And not just improving volume as much as possible.

Yann Bilien08:36

Is that what a ripple tide is working on? other than Yeah, we are working on this autonomy, so having systems that can definitely think reason and make decisions and not just doing just the volume. Yeah. Sequence of actions. Predefine, which is more related to automation and we are working through autonomy on our side. Okay. Maybe that's the reason why these AI, SDR companies are, or products are failing. we see a lot of social media posts about people are trying it out and then abandoning them. exactly like what you said, and I wanna know how's the autonomous agent different? what more parameters are?

Angelina09:14

Coming into play. You mentioned some, can you say more?

Angelina09:16

Yeah, probably. So let me give you an example of an autonomous system than interact. So if you take way more self-driving course, when you get into the car, the car doesn't know, it'll have to accelerate, I don't know, in three minutes and 47 seconds, or then break or something.

Yann Bilien09:32

So based on what's happening around the pedestrians and so on, it'll make a decision to act in a way or another way. That's the same in agents. If at the very beginning, What you will have to do, the exact sequence of actions with the timing and so on. You can just follow this workflow and probably it'll be more reliable that if you try to introduce such autonomy. So I guess that it makes sense to have autonomous systems in all situations. You cannot predict in advance what will happen. And in many situations, that's where humans are involved into it, that you, as you, you cannot predict and anticipate what will the human do. You'll have to find new solutions based on the on, on what's happening.

Yann Bilien10:14

the key in autonomy is being able to make some decisions and if you have to make decisions. Sometimes in real time, so sometimes not. You need autonomy to do that. And I don't know if prospection or AI is deals is the best case throughout autonomy because it's pretty based on the workflow. So I guess there are tons of other use cases where autonomy is probably best suited to do that. We are autonomous agents, right? We're autonomous. Yeah. In a way we, the, that's probably that, my guess is AI systems. Do not necessarily need to mimic how humans think, how humans reason, and so on. But definitely there are some patterns in, in the way we,we make decisions and sometimes based on we, we see, a signal, a noise, so we have a lot of sensors and in the end we make decisions.

Angelina11:07

So on the high level, I think you, we can take some inspiration for AI systems. But at the low level, we are pretty far to be thinking the same way as humans do. That makes sense. We're just guessing how we work as well. In your product description, I mean you also mentioned a lot of the word of deterministic, but if you are, autonomous, right?

Angelina11:29

The agents are autonomous. The autonomous is based on reasoning models, which is intrinsically probabilistic, right? How is deterministic possible here?

Yann Bilien11:39

Yeah, I get that. And I think you are referring to the famous trade off in AI between control on one side and flexibility on the other side. So flexibility means you, you'll be able to adapt to new situations you didn't see before, even if you did not anticipate them. And control is rather, you want your system to do, you want to expect something from the system and the system will do it the same way as you expected. And so that, that's a very famous trade off. And I guess the answer to that, it shouldn't be a trade off because you need control at, in some parts of the systems, and you need flexibility in other parts.

Yann Bilien12:15

And it's not at the same place that you would, that, that you want to have this trade off. And so when you're referring to LMS in agentic frameworks, and it's the way agents are built today. So if you take the basic framework, you have an LLM at the center that takes a question or something like this.

Angelina12:32

That can call some external tools to perform some actions, but in the end it's the LLM that is making decisions. And as you mentioned, this is probabilistic. And so I guess in your agents, and especially in sensitive context, high stakes agents and so on. What you want to be deterministic is not necessarily the behavior, it's rather the quadras that you're providing, the system. So for example, you don't want your system to talk about pricing. If you have a customer success agent, for example, if you put that the traditional wage build agent, you will just prompt on LLM, don't talk about pricing, and in the end. Just the highest probability tokens that will be selected by the LLM.

Angelina13:16

So it'll be probabilistic, and sometimes if the signal is higher somewhere else, it won't follow your Quadra. So I guess on this side you need deterministic quadras, but at some point you definitely need some flexibility. Otherwise, you can just do workflows, if you can't, build logic and new decisions based on what's happening.

Angelina13:36

not at the same place between control and flexibility, different areas of the agency framework, how do you put in those controls? what you mentioned, example. so we worked a lot on removing the decision from the LLM. So in the agentic framework that is not on LLM, that is making the decisions, that is choosing to call some external tools and so on.

Yann Bilien14:00

That is other type of technologies. And for example, not a lot of research works regarding neuro symbolic ai, regarding graphs and so on. you can then leverage all those technologies to, to become deterministic. You can be binary at some point, you can forbid some graph areas. you can leverage those technology and especially there, there is one that is quite old, but that is coming back to the front, which is a neuro symbolic ai.

Yann Bilien14:27

And so neuro symbol AI means you have this neural part regarding flexibility. So reinforcement learning, traditional AI systems. And then the other part, you have the old way of doing things.

Yann Bilien14:39

So symbolism, so being able to put some rules, some logic around this. And so if you leverage this, definitely you can, for example, have deterministic quads if you do that. You're saying you're putting, you're combining the rule based engine, like the same old together with that's way of doing. So clearly, I guess you, you can, probably split your problems into some parts where it'll be AI with,probabilistic computations and so on and at other areas. You want something that is more related to symbolism, to, to roles and things like that.

Angelina15:12

Yeah. Yeah, that's a really good insight. Not putting it in a prompt but you put it in within the rule engine is very interesting and I don't hear often like people talking about it, so thanks for sharing.

Angelina15:24

Yeah, no, no problem. And I think that today everyone is building agent in the valley and during the same way because using an LLM is naturally the first thing to do because that's easy to use in a way, right?

Yann Bilien15:38

you can just do some API calls and so on from something, and. The demo will be very great, but we are hitting a wall toward this and regarding agents that are not going into production regarding autonomy that you cannot achieve with such elements and so on. So I guess in the next coming months people will hit this wall and will go back to maybe new or previous architectures to do that and we'll try to switch the way they are usually building agents, I guess so.

Yann Bilien16:07

I like it when you say that it's in the next few months. It's not the next years, next few years, it's next few months. That's how it's really quick nowadays. So if I wanna be realistic, I think it's rather in the next month. Oh my goodness. Yeah.

Yann Bilien16:20

I think that's also why also a lot of the AI agents, even my clients, they have a really nice demo and as you said, they don't deploy them into productions.

Angelina16:30

Can you share, like your observation, like

Yann Bilien16:32

what can go wrong? Yeah. Yeah. So basically I think it's related to this LLM making decision, basically because when you build your demo, so using LM, it's really quick. So in a week you can have a very amazing demo and as we mentioned before, in the production situations, you cannot anticipate what someone will do and. That's the way humans are built. in a way there are so many questions If you have a customer facing agent, for example, that a human can ask with a very little probability, for it.

Yann Bilien17:07

So that's called long tail distribution. That means very little probability, but so many possibilities. So you can't anticipate everything in advance.

Yann Bilien17:15

So on your prototype, it's working because basically you just took, uh, I give you the perfect data. Yeah, exactly. So the same question as always. You iterated a lot of your prompt and so on, but the day you release, it's in production. You have things that you did not anticipate before and that day, usually it breaks because the DLLM as not truly understanding what's happening around.

Yann Bilien17:37

It's not making correct decisions. It's rather trying to generate coherent text based on the training set that is, has seen before. And so in production on new question outside of the training set, it's usually breaking and maybe Alex, sometimes the chess metaphor to explain this, which is. If you are a very experienced chess player and you are playing against the beginner, you don't need to think or to reason at all.

Angelina18:01

You will just recognize some patterns of what the beginner is doing, and you'll find naturally the way to corner the beginner without, thinking. But if you're playing against someone that is also very experienced. How humans would do that is they would try to simulate some moves. Say, okay, if I do this sequence of moves, will I win?

Angelina18:21

Okay, so maybe that wasn't the right decision, I'll try something else. So simulate different possibilities or actions and in the end, pick the right move. And that's what happens in the agents. When it's a prototype, you are playing against the beginner. So just generating, without thinking, without needing to think.

Angelina18:39

So just generating what is naturally coming and so on. But the day one production with new questions that you did not anticipate before, you would need to think to reason to say, what should I do, even if I don't know.

Angelina18:53

And that's where if you have the DLLM at the center of your agent. It's usually break. that's a really good analogy. So the demo AI agents are maybe experienced enough, but then they're fighting against the truly autonomous agents, which are humans, right? Yeah. So for sure we're more experienced. and you mentioned about in this pipeline building an AI agents like this, right?

Angelina19:14

and we're removing the decision points from EL and

Angelina19:18

who's making the decision? How do we bake decisions in this pipeline? That's a great question. I, and I think today it really depends on the use case you are addressing and so on. Since there is not a foundational technology such as language models to do that and so on, you cannot just take something that is doing the general case and you can apply it to reason and make decisions.

Yann Bilien19:40

That would be a GI basically. And so today we are not on this, so some people are trying to scale compute and so on to, to reach an unstructured data to reach this level of autonomy. Another path to do that is to be more verticalized, and you see all those companies that are doing vertical things so that you can develop reasoning engine decision making engines on such verticals because you need less data, less compute and so on.

Angelina20:05

And so the rules and the business logic. Just works on a small subset, just on the vertical, but at least you can try to mimic the reasonings and the behaviors of what people would do working in this industry. I remember reading one of your blog posts. Yeah. And I read some of your blog posts. It's very well written, and you mentioned about.

Angelina20:25

verticalized reasoning engine requiring much less compute and in comparison with the typical chain of salt model development, which might get maxed out at scaling point by 2026. You have more comment about that.

Yann Bilien20:43

I think the main limiting factor now is not compute. it's rather data in a way because compute is create extensive, develop, are building new data centers, new compute things to, to do those computations. But what's we are really missing is data in a way, because such element we are referring to language model. So you need. Text to feed those models and text is, has a low information density, so that means you need a lot of text to be able to do some logic on it because, for example, compared to a video or a sensor or something like that, you have much less information, logic, things that are not explicit in text.

Yann Bilien21:25

It's difficult to do reasonings that are not explicit on text. And there is this, conference at Neuralips. It was, last year from Ilia, the co-founder of Open AI and subsequent intelligence around the typical scaling low. So until now, there was a low where when you improve the number of,of tokens, so the data you feed to your model, the performance of your model will also improve. and that's a logarithmic, uh, load now. We, the, this low is plateauing in a way and will soon plateau because we are missing data and,I can quote this French scientist regarding the amount of information l LMS have seen and that

Yann Bilien22:08

the biggest LMS now just have seen the same amount of information of a 4-year-old child.

Yann Bilien22:13

And that we don't get much more text to use to, train LMS accepting synthetic data and things like this and that. So we will soon be missing some data to keep the improvement rate of language models regarding this. So I think. Compute to solve. Compute is not the limiting factor, but data is definitely with such architectures

Yann Bilien22:37

and synthetic data is not gonna be enough. Yeah. De depending on so different applications, but for a lot of applications generating data, if you can generate the data, basically it won't bring any new signals into your model. So you will loop on a closed loop. And the your models will slowly derive and at the end lower their performance. So there are some cases where it works, but the majority of cases it doesn't work.

Yann Bilien23:01

So far, I'm shocked that you know the largest oms only see what 4-year-old has seen.

Yann Bilien23:07

And that's very, in a way, that's only because a child has sensors so he can see a lot of things he can hear. So there is a lot of data that the child can see, but the LM has just this text, so you know, you have to find some pattern, some information, some way to reason, just based on semantic.

Yann Bilien23:27

So some text, and that's something in the future probably that will improve. That's the key to unlock robotic, especially because we can leverage sensors and so the density of information is much higher. Hmm. I see. the pure volume of data is actually not that much that it's being trained on. Yeah. It's the density of information at the end. Got it. Yeah. I guess AI agents or the OMS need some life experience like living a life like a, that would be definitely, and there are some ways now that people are exploring to train agents, like having kind of a buddy. you bring your agents with you and then you will learn alongside what you're doing.

Yann Bilien24:05

So seeing the actions you are doing and so on to train it. Instead of just training foundational models and then fine tuning it on a use case. It's just bringing a system with you every day and in the end it'll learn the same way as you're doing.

Angelina24:19

Yeah, A little bit scary, a little bit creepy, but yeah, I get it.

Angelina24:23

You also on, on your website also mentioned that AI agents, you're bringing like from 80% accuracy level to 99% accuracy level with. Hypergraph decisions and so on. How do you get to 99%? that's pretty bold statement. Yeah. Yeah. So based on this, maybe we can just come back to what we mentioned before on agents where you want to do automation and so on, and the ones you need autonomy. And I think there is another distinction is when someone tries to build an agent, I usually ask, is 80% performance enough or not for your agents? And in a lot of use cases, 80% is enough. For example, if you are just solving customer success tickets, alright? If you are already solving 80%, that's fine. You, that, that's already a big gain for your company, so you can stop here.

Angelina25:14

On the other hand, there are some use cases where you need 99.9%, performance, and especially when it's customer facing, when it's related to regulatory industries or things like this. For example, take a personal assistant agent that is booking meetings for you with your prospects or whatever people you have to meet.

Angelina25:34

80% is clearly not enough. If it sends the wrong time slot one time out of five, that's not okay. the people you will meet will be very angry about that. So there are some situations where 80% is enough and other situations where 99.9 is required. And so we are very focused on ourselves to bring agents to 99.9 for those high stakes agents.

Angelina26:00

So in a way, how you can do that is clearly removing and some probabilistic behaviors at some points. And why this 80%? Where does it come from? This 80% is the compounding of errors. And because when you are building agents, usually you take a process that has mult multiple steps. And I dunno, if you take an LLM for example with a market standard 5% hallucination rate, and you have 10 steps in a row.

Angelina26:28

At the end, the success rate will be just 60%. So close to one out of two you will have some hallucinations, anything like this. So that's working for 80%, but that's not working for other 99.9% agents needed. So in those cases, you need something else that the LLM, and so that's why we've been working on you.

Yann Bilien26:47

You cited this hypergraph to bring some higher safety and reliability for the agents in production. What is Hypergraph? Yeah, sorry for that. So a hypergraph is just a graph with multiple dimensions. So when you build an agents, you have two main layers. At some point you have a memory layer and you have a reasoning layer.

Angelina27:07

And so that's the way you can represent the data there. You can have some text, for example, just from LLM, otherwise you can bring some, what people call knowledge graphs. So that means bringing some logic. So saying, what's the relation, for example, between Yian and Angelina, so that when your system is talking about you, it'll know that you are speaking with me, for example, that relationship is deterministic, right?

Angelina27:30

So that's not Yeah. On campus in the,in the memory layer at some point, yeah. Yeah. 40% error rate is pretty high. Yeah, it is. It is. But that's fine for you for some use cases. So that's why I'm always asking, is 80% enough or not? If it's enough, just work with that. And even maybe you should be able to do that with a workflow, so it'll be a hundred percent deterministic, which is even better.

Angelina27:54

But if you can do a workflow, just do a workflow in the end. That's when we put something, when we put a, a prototype agent in production and how do we validate when there may be data drift? Oh, So data drift. So you means systems,the performance will decrease a long time product.

Angelina28:09

Yeah. Yeah. That's interesting subject. And at the same time, there are actually very few real autonomous systems in production. So this problem of drift. Is maybe less important that other issues, and especially today, many issues are hallucinations and guardrails are not followed, but drift can come maybe as a third issue.

Yann Bilien28:33

So that's tough to evaluate actually, because you need a way first to evaluate your agents, which is more broad, but. In the end, you need to do that. that's the same thing as you have a prototype. You want to release it in production. How do I know it's safe enough to release it? And usually people are scared of releasing agents in production because they don't know what will happen in different situations. How, what people will ask and how the agents will react.

Yann Bilien29:00

And so this evaluation scheme so far has been a challenge in ai and as you the main solutions. Show that it's a tough challenge. For example, you think the, maybe the most common techniques which are LLM as a judge. So you use an LLM to evaluate the output of another LLM or you put human in the loop.

Yann Bilien29:21

So that means you have someone that is reviewing what the agent is doing. Such solutions show that it's a tough challenge and that it's not solved yet. If you need a na, a human to evaluate your agent in the end. That there is, it's not autonomous at all. Same for evaluating. So that means probably based on your business logic to have something. To evaluate your agent. I don't know if you have this customer success agent, probably it's related to the percentage of tickets it can solve alone. If you have a sales agent, it's maybe the conversion rate, you have and so on. So, I would advise definitely to have really business metrics, business evaluations related to the agent, and not relying only on foundational evaluations because it won't be very suited to your agent.

Angelina30:11

That makes sense. I'm a machine learning engineer, so I know, for structured data, for measuring data drifts, and we look at the distribution of the data in production. So it's more straightforward compared with using agents in production.

Angelina30:23

Yeah. And. Probably that there, there are also two ways to evaluate.

Angelina30:27

As you mentioned, you can have some statistic approaches and especially that's something used when you want to do feedback loops. So for example, you are doing reinforcement learning. So you will do the actions and then you will evaluate them and you will try to retrain a bit your model. But on the other hand, you can also do that in advance.

Yann Bilien30:45

So trying to predict what will happen if you do something. And so at least you don't need to, I don't know if you have a sales agent to burn a lead to, to learn something new. you can predict in advance. Are you saying that, just stress test the agent? Kind like, yes, but even at the foundational level, there are some research works regarding this, which is predicting things and not just evaluating at the end with a feedback loop.

Angelina31:10

So prediction. Interesting. I have to read more into that.

Angelina31:14

I also wanna ask about what you think of a context windows, Because we are feeding a lot of information into context window as prompt engineering goals, right? Yeah. What do you think of what if we, stuff it as so full that we cannot stuff more into it, and will we have infinite memory?

Angelina31:28

Is that even desirable?

Yann Bilien31:30

So regarding that, there are two things to consider. the first one is. If you want to build autonomous agents, as we mentioned, you need enough data to describe the actual situation that is happening. So in a way, you need to put a lot of data and to process a lot of data. On the other hand, today, techniques with large context window. I was discussing with someone working on Gemini yesterday, and you know this. 1 million context window, right? It has still a lot of issues regarding pits of attention in the middle. For example, a lot of tokens have less intention attention and so on.

Yann Bilien32:05

On the, and the third point is today in gen AI projects, usually if you separate the memory and the reasoning layers, you have a pipeline between both. Usually it's either a rag graph rag or whatever. So a retrieval pipeline, which is you have a query and you want to get some related information so that you will be able to answer to the question.

Yann Bilien32:27

And usually those pipelines are bottlenecks in the projects because RAG is not exhaustive. Sometimes you have chunks of context that are not related to what you're talking about, so that in the end you will hallucinate and so on. So if you should probably try in your projects. To remove those bottleneck pipelines and try to bridge the gap between the knowledge, the memory, and where the action is made, the decision is made. And if you can bridge the gap between both. So one way to extend this context window is a way to do that today. Still has lot of imperfections and so on. But there are other way to, to do that, and especially if you structure your data another way, because context window is very useful if you have large text, for example.

Yann Bilien33:15

But if you are able to structure the text, so to keep the same amounts of information, but with less tokens, probably at the end you will get a more accurate results. Is this what people are, talking about that's called context engineering these days? Like your, oh yeah, that, that's interesting term context engineering, because people probably.

Yann Bilien33:36

So that what they called prompt engineering and so on is not the step to spend some time. And probably it's not just regarding a prompt, because if you're just doing a prompt, you are just trying to fine tune an l, LM and so on, and that there is something else related to how you bring data to your models, what type of architectures you are using. Do you need the whole context to feed to your system? Every time, or you just need to give it a subset of it. So this context engineering term is very interesting to me. and I think it's going to toward the right direction to what models need to do that, how to give that information so that models are more accurate in the end.

Angelina34:18

Yeah. We just made a video about it, so we'll. Oh, nice. Yeah. My co-founder. Yeah. what do you think,

Angelina34:24

what do you think? a very smooth transition from a demo. To a more dependable production use case look like. Yeah. I think today the, let's say the reliability and the safety of agents is clearly key points to release in production and especially when you have high stakes agents and you cannot allow yourself to have hallucinations GU race not followed, anything like this. So I think it's really regarding this, trying to have. Agent, maybe that can do less things, but at the end, it protects the brand you are representing. For example, if you are a luxury brand, you want your agent to be a premium experience. You don't want to have, to be spammy, to be noisy, and you really have to represent the brand you are working with you.

Angelina35:16

You cannot invent things. You not follow roles, for example. And so being very close to the business logic and being sure that. The model has all the elements to really reproduce what the business is doing is key. And so brand safety, so representing the right manner, the brand is, definitely key.

Yann Bilien35:36

And to do it's safety, reliability. The difference between both the reliability is if I do the usual things people would do with these agents. What's the percentage rate? Will it work every time you do that and safety is rather focused on not di divulging information, not hallucinating and things like that. You know that there is this example with a company aligned where the agent, the customer success agent invented a discount for customer. And so then it went into a lawsuit and so on. So cost a lot to people. So releasing in production is really regarding safety and reliability.

Angelina36:10

Yeah, that makes total sense.

Angelina36:12

Yeah. Let's go back to how ripple tide work. Let's say I'm your client, how do I get started? What will happen?

Angelina36:18

Alright, so I would ask you the, this question is 80% enough for your agents? Let's say, yeah, let's say 99 9. let's say we're a customer service bot, but for healthcare, right?

Angelina36:32

For healthcare devices. so it's almost zero tolerance. You cannot go wrong on healthcare.

Yann Bilien36:37

Yeah. sure. So when building agents, you have now typically two steps. the first one is regarding structuring your knowledge, having pipelines, and having things that. Agents will be able to use and describing your business logic and so on. This, what are those data? Depending on what you, the agent should do, I don't know if your healthcare agent should answer questions regarding your health insurance. For example, you'll need to learn what is mandatory in the health insurance. You'll have to bring some logic between all the concepts of those health insurances and so on.

Angelina37:12

Then you'll need to have a system that is able to make some decisions based on, what the user can do with the agent. So define what is the behavior of the agent. In the end, having a system that can at the same time break this, trade off between control and flexibility. And you mentioned, and especially so being able to pick some right decisions based on the current information you have.

Yann Bilien37:33

Maybe it shouldn't answer if you don't have enough information. maybe,it should follow some. So you have to define them. So it's really related to, to business and at the same time, having the right technology to do that. Who are your ideal customer profile? Who are your, on our, yeah. on our side.

Angelina37:50

We are really working with people with high stake agents,

Angelina38:08

should work very reliably in any situation.

Angelina38:12

I think that sounds to me a very good differentiator among all these agent companies, right?

Angelina38:18

Yeah. so that's the first point. And the second one is do you need autonomous systems or is the workflow enough to do what you're doing? And I think it makes sense, these.

Angelina38:29

Things and so on. If definitely you need autonomy and so by saying autonomy is, you can describe in advance what the agent should do so that you need to bring some intelligence to the project.

Yann Bilien38:40

Should companies build or buy for these AI agents needs? That's a tough question. Maybe two elements to, to answer that. The first one is an observation, which is. Who are the best people to evaluate the results of an agent, and usually that's business people that are using the agent. If you give to an IT team that is building the agent, they won't be able to say the agent is working properly and so on. Yeah, because you, we don't have these metrics and so on, on text or thing like that.

Yann Bilien39:08

In the end, you probably need someone from the business to be involved in that. So I guess, and the second point is if you need reliable system that are autonomous, probably that you'll need to overcome current state of the art limitations. So that's require, let's say, strong machine learning, data scientists, teams and so on.

Yann Bilien39:27

So definitely I think building inhouse. Is a good point as you can be close to the business logic, but at the same time, you need the right technology to push forward current limitations. So either you can buy this technology and then integrate by yourself in the company because you'll have the business proximity and so on. And if the, let's say you, the teams are not really working on the subjects and then. Buying Makes sense so that you have the integrated technology and it'll be rather working with this company that will integrate regarding the knowledge, regarding the business logic and so on. you can buy just the technology or you can also buy the integration with it.

Yann Bilien40:08

And the sub path is definitely to, to build in-house. But today. I'd say one limitation of building such, systems inhouse is going very quickly. And so you cannot work on old subjects, in ai. And so definitely you have to pick some battles at the end. So if you pick the right ones, you can definitely do that.

Yann Bilien40:25

But maybe the next hot subject you will need, the technology you will need or something you did not consider before. So from a timing perspective, having external technology is, can be helpful.

Angelina40:37

What about data? what about internal data? Like companies might have their proprietary data, right?

Angelina40:41

That may be a driver for why they wanna build their own stuff. Yeah, definitely. So there are usual data lakes that you may have and agents also need some something else. So like playbooks thing like this. So something. Maybe it doesn't exist in current systems so far. So it show proximity with the data is key.

Angelina41:01

Then it's to see how it is structured, data, sanity and so on, because agents require some stuff on it. Or on the other hand, there are now some systems that can work with, let's say, data, very incomplete data, very badly annotated, anything like that. And there are some agents now that it's possible to build some that are working on such data.

Angelina41:22

And so I guess depending on what you have today in your company, is that clean data or is that messy data? Probably there are solutions for both, but you should consider technological solution based on what you have today.

Yann Bilien41:34

Okay, awesome. Thanks for your guidance. and I have two more questions, a lot of the, public perception about OMS and where these models are going is these models are getting smarter and smarter, right? And. Can we expect the production AI agents will be easier and easier due to the fact that model's just being smarter or it's not the same thing? Yeah, I guess that's a right point. And I think there is still a huge gap between foundational models and so on, and the applications you can build, using them.

Yann Bilien42:08

And so this gap has always existed and on, on one way, you have so very capable models.

Yann Bilien42:16

Even today. for example, language models are very good at two things. The first one is compressing vast amount of data. You can give trillions of tokens, to the system, and you will be able to retrieve the signal using language models.

Yann Bilien42:30

And second thing is generating natural language. And so we have very capable models regarding language understanding, also language generation and things like this. And, but there, there is this gap in applying them and that's why you see so many companies that are building vertical applications. in every industry you have tens, and dozens of startups that are building vertical agents on it and so on. And there is this gap between research foundational models and so on, and the ones that are used in the real world. And I think, so there is plenty of space in the middle of this to bridge this gap. And I think so for sure, the models will get smarter and smarter. The thing is how people can use them.

Yann Bilien43:14

Probably at the very beginning, there are some big tech labs with very capable models and so on, but no one are using them in production because that's too heavy. You need too much compute for insurance of those models and so on. So you cannot serve those models easily. And I think when you say that models are getting smarter and.

Yann Bilien43:33

Probably also they will be smaller and smaller. I believe that you will need less computer run each model.

Yann Bilien43:38

But on the other hand, the fact is you will be able instead of doing one LLM request, for example, to do 10 or 20 for the same thing so that you will get more accurate answers. So you just, you don't just need a model that is smarter. If you can do 10 or 20 requests, probably you will get a more accurate result. People, I think the compute needed for people to use such models in the end will remain quite the same because each model will require less compute, but you will call the same model multiple times in the end. So I think bridging the gap between, so those foundational models and the vertical application is still probably a challenge.

Yann Bilien44:15

And, the distribution between both and how you can serve and build real system user leveraging them is still a challenge.

Angelina44:23

Yeah, the challenge will probably exist still, for the foreseeable future. I don't know.

Yann Bilien44:33

We still hire people have experience in something. So people, we are verticalized as well. A data scientist. A data scientist, you won't hire data scientist as a doctor to treat your, symptoms. Right? So now, but you can address that by designing agents that have, I don't know if you're building a sales agent, when you're hiring a sales person That the person doesn't know your company, but. If he already sold product before, he has a sales sense, and something like that. So when you're hiring someone, he knows how to sell, but he doesn't know your company. And with agents that can be quite the same. You have reasoning techniques and so on. So an agent with a sales sense, for example, and then you just have to give some information about your company.

Yann Bilien45:12

You have a very capable system from the beginning. I feel we are getting closer and closer to, human beings being potentially replaced by these easy AI agents.

Yann Bilien45:21

Where do you think that is? I don't think so. I think we can do a lot of things with autonomous systems and so on, and the idea is clearly that you can do more with less people probably.

Yann Bilien45:33

And if there is a company that there are all of those digital workers, narrative and so on. If you take a company that is. Hiring digital workers and then is firing people and the other company is hiring the same amount of digital workers, but keeping the same people. But they can do 10 x more.

Yann Bilien45:51

Because it's compounding between digital workers and humans. I'll let you guess which one will be out of the competition. And so I think that definitely combining P positions will change a bit, but combining humans with autonomous systems and so on will definitely be better than replacing people. And if you want, I guess that so far, if you have such digital workers and so on, it's a competitive advantage compared to your competitors.

Yann Bilien46:18

But in a few months to years. If you don't, you will just lag behind. It will be the one who doesn't have agents and so on will just lag behind. So it's now time to, I think to dig into the subjects and to see how we can compound the effects of having digital workers working alongside with humans because they are doing different things, but combining them is definitely a better solution to deliver more than replacing people.

Yann Bilien46:45

Amazing. Do you think AI agents will ever become truly autonomous? Yeah, I guess so. So autonomy is being more and more culturalized and when you are in the valley with, Waymo self-driving cars and so on, we can do a lot of things, the surroundings, the research is accelerating regarding autonomous systems and so on, and I'm pretty proud of what we are doing, triple tie toward this and I think we're on a lot of situations.

Yann Bilien47:10

We are clo close to it and at it in some situations, so definitely yes. Autonomous systems will be next. big thing.

Angelina47:18

Awesome. I'm looking forward to that. That's all my questions. Thank you Ian. I really, yeah, thank you. It was a pleasure to have this conversation. Yeah. yeah. I learned a lot and thank you for sharing all these information I'll see you next time.

Angelina47:31

Yeah, sure. I'll see you. Thank you very much. Bye-Bye.

More episodes