Why Your Model is Probably Using Bad Data
About this episode
Ahmed Rashad went from drilling for oil to running growth at Scale AI to raising $17.5M for Perle.ai. In this conversation, we dive deep into why 95% of AI systems fail in production (spoiler: it's the data), why medical AI requires zero error tolerance, and why "vibe coding" creates beautiful demos but terrible products.
We explore the uncomfortable reality that most companies face: your model is probably using bad data, and synthetic data won't save you. Ahmed shares real stories from the trenches - including those 9pm panic calls from companies whose AI is about to explode before Monday's launch.
Key topics covered: 🛢️ From offshore oil rigs to MIT to founding an AI company 📊 Why we're running out of quality data (and why it depends on your use case) 🏥 Medical AI labeling: When hallucinations literally cost lives 💰 The TCO approach: Why $??/document beats $3/document 🧠 Extracting human wisdom vs. just collecting labels 🔬 The difference between demo success and production reality ⚖️ Legal AI and the "it depends" problem 🌍 Why the world is too complex for synthetic data 👥 When to worry about data quality in your AI system
Key moments
- 00:00 Introduction: Ahmed Rashad’s career transition from offshore oil drilling to software and the founding of Perle AI.
- 02:00 The AI Data Problem: A look at why AI will always have a data problem as applications become more sophisticated and specialized.
- 09:00 Live Demo of Data Annotation: A walkthrough of the Perle AI platform showing automated transcription, PII removal, and expert validation of medical data.
- 14:00 Expert Context & Human Wisdom: Why "consensus" isn't always the right answer and how to extract high-level reasoning from specialists.
- 29:00 Cost vs. Quality (TCO): A case study explaining why a higher upfront cost for quality labeling ($18 vs. $3) leads to lower long-term expenses.
- 41:00 Capturing Experience & Future Vision: Exploring whether AI can eventually capture human "wisdom" and the possibility of eliminating the human-in-the-loop.
Full transcript
Do you think your company and your vision is ahead of time?
I think it's gonna be the next step of how data pipelines are managed. Hallucinations cost a lot. You cannot make a mistake, so it's like, oh, but it does well enough. But whatever Chad GPT, or Quad does well enough, well, well enough is not good enough.
Can we eliminate human in the loop one day?
I think it's possible sometime in the reasonably distant ture we can capture the human input without the human having to go into a specific place to give that input.
Some company wanna use your data for training. Isn't like labeling with experts really expensive?
The short answer is yes. And
let's say I wanna build an AI agent system. When do I need to worry about data labeling?
The answer is, it varies.
What does data annotation mean in your definition?
For me, data annotation means, let me actually show you how about that?
Okay. Awesome.
Hello everyone, this is TSET ai, and [00:01:00] today I'm sitting down with Ahmed Rashad from pearl.ai.
Yeah, absolutely. I had a very interesting, career trajectory. I, started my career as an offshore oil driller and, taught myself how to code and built a product and became a, uh, worked since, been in software for the last 20 years since then.
That's really cool. how did you transition from the oil rigs to like data labeling
complete, coincide? complete coincidence? Absolute coincidence.
yeah, absolutely. So you're solving a training data problem for ai. I'm curious, does AI actually have a data problem?
AI will always, I think, have a data problem. I think, what I think, so I think,as the applications get more and more sophisticated as we progress, even the older applications, as we start refining them, the data needs evolve and get more sophisticated as we move forward. you look at the use cases we used to solve a few years ago, 10 years ago, they were, they used to be a lot simpler.
you're the expert in this area. for me,for the ai era of, data tagging, I'm thinking of let's give you two images. Is this a dog, is this sort of cat and Right. it's with a person with common sense, like myself.
it varies depending on. How advanced [00:04:00] or how much progression has been made in the specific area. for example, in, in some languages where the AI hasn't developed as much.
Mm-hmm.
there are still some things like, oh, we need to actually transcribe this properly. We need to actually identify what does this word actually mean within this context?
Oh, sure.
Yeah. it, it matters like dramatically changes the context. It dramatically changes the meaning. Now in other areas where we've overcome a lot of the language and the language isn't as problematic, like English for example. We've gone into other fields where we're looking at, much more sophisticated applications that aren't necessarily dedicated to the language like English is pretty much in very good space.
It's not a common sense person can, tell what is right, what is wrong.
Absolutely. if I show you a medical scan now of someone's brain, can you tell me if this is okay or not? It's,
can you show me, do you have one?
I mean, I [00:06:00] can, I just dunno if it's gonna be HIPAA compliant.
Okay. Yeah. It's not a simple dog recognition. All right. you mentioned about the reality of the real world, Today of AI use cases and the complexity comes in with translation, underserved languages, and then serving the right context. you also mentioned about the, the image part, right?
Is that a problem you are solving too?
Mm-hmm.
But if you're looking for. Let's say hybrid speech between someone who switches back and forth between a specific dialect of this language and French, then that might be a little bit more problematic.
Right?
so the answer is, again, it depends on the specific area in a lot of areas, but they're broken down into three categories. One, we have a lot, two, we have a lot, but we don't have access to.
And all three categories are needed for training, language models or multimodality [00:08:00] models.
yes. It's correct. Yeah. 'cause you think about language models and language models are evolving so fast and they are learning to do new skills every couple of weeks, like e every couple of weeks you hear about a new breakthrough, right? Yeah. And these breakthroughs aren't just happening, they aren't being developed in these two weeks.
I'm very tempted to ask you about the trend in ai, but let's take one step back. I think you were talking about a different type of, data annotation or data labeling from my definition,
For me, data annotation means, actually, you know what, let me actually show you. How about that?
Awesome. Yay. Love live demos.
this one's slightly old, but it is. It is cleaned and it's HIPAA cleared.
Ooh.
Yep.
Oh, cool. And they can
Yep. At comment comments. They can add comments and at the end of the day, you could imagine that this actually, the way it happens is it gets triaged first. The case gets triaged from a general practitioner's oh, actually this is a head scan.
is the, brain [00:12:00] image also labeled as well? Is that one also labeled? Yeah,
yeah. Yeah. If you look at the brain image, like you see these dots, these yellow and blue dots,that's labeling.
what can we do with these labeled results
Yeah. So this teaches the customer's models. How to detect these anomalies and how to detect these contrasts and how to do the triage and how to extract the key values and the named entities, like the clinical signs, the pathologies, et cetera, et cetera.
Right?
And sometimes the speech is just, it's an ongoing conversation and it's not very [00:13:00] clear. the problem with medicine is that there's no room for error.
Zero tolerance.
Zero tolerance. Zero error tolerance.
this data can be used for, medical training and the workflows in the medical practices.
Yeah.
So I think you answered a lot of my question already. Visual is better than a thousand words, right?
Yeah. so expert context, let's take a step back. So when you think about the expert context,when you learn pyres theorum, when you first learn it, it's A squared plus BAA squared plus B squared equals C squared, right? And then later on,when you advance a little bit, you learn that there's more to it than that actually it's not just that there's actually a little bit more.
Uh, they give you less instruction.
so it's a very tricky balance because in most use cases it's okay, but then on the edges you actually need the smaller population of people who are more creative, who are more the edges themselves to actually stand out. \
That is very different from, if I say we just prompt engineer the workflow
it doesn't work the same way they, you are. Don't get me wrong, we still use a lot of AI inputs and we still use a lot of prompt engineering.
that's why you put that on your website, extracting this wisdom or giving your data wisdom. So it's actually from experts,
from humans. yes. and how do you create the right conditions for people to thrive and actually give you their wisdom.
that sounds like the wrong flywheel to get on.
It's a very wrong flywheel to get on. It's a very bad flywheel to get on.
you wanna build a premium product that actually works, right? Not the offshore teams.
Yes. Yes. And a premium product with low total cost of ownership. like you're genuinely delivering low total cost of ownership for the customer.
I'll say best things about you on Reddit.
Yeah. Sometimes. Yeah.
what's the commonly expected error rates in handling these, uh, data quality in your use cases you've seen? Is there a tolerance level that you've seen in the industry or verticals, or It's purely zero,
It is a very tricky question because in, in a lot of the industries that we handle, there are three main types of error rates.
right?
The problem is in these use cases that, and we've seen this over and over again, customers are used to handling high error rates and they're just frustrated, and they're just used to, frankly not very good quality.
And
by the time we're done and we finish the POC and they see the quality, it's oh, okay.
you're saying actually the, in the industry, the quality is actually not very high?
it's very variable. Like sometimes it's high, sometimes it's very variable. The [00:23:00] variability is the problem.
is that benchmarked against the human error rate because, if a doctor makes a mistake. The worst thing is they get sued.
So yes and no. I think in medicine the objective isn't to replace the doctors, per se. I think the objective is to help the doctors be more productive.
Which
is the value add is them actually seeing patients. Them doing paperwork, them walking from building to building, them doing work that a nurse or a physician assistant or an AI could do. That's [00:24:00] all incidental or known value add, right? Absolutely.
is accuracy the most important metric?
Yes. In a lot of ways, yes. Because if you are, see the answer is always, it depends. You're becoming a
lawyer.
I'm a lawyer so
much,
I spend so much time with lawyers.
Exactly. So to
drive adoption, they need very high confidence that this AI thing is not gonna make a mistake that's gonna get them in trouble.
That makes sense. Well, the, ai adoption process is a whole different story. Yes.
The short answer is yes. And the way we do it is so when we build this whole thing, we built it,with the intention of having it at the end so that the customer can do it fully self-serve, even without our intervention.
What does that mean? So,
so it means that the customer can go in and you, they can use our tooling and they can set up a data pipeline. They can collect and label data through a combination of experts, agents, ai, other software linkers, et cetera, et cetera, and design their pipeline the way they want it [00:27:00] without any intervention from a third party like us or anyone else.
right?
So once you figure out what the optimal flow is
It sounds like if I'm a medical center and I walk with you and you will set up my pipeline for me during the POC, and then you'll help me get to the self-service, state, and then the cost will be lower.
the cost will be significantly lower. the longer term the project, the more time I have to standardize it and to make it much more efficient.
do you have any ROI stories that you tell about does this make financial, sense? I feel like if I'm on a fence as a medical center or law firm, I really don't know what to expect. Maybe a lot of your customer have this objection or question as well. How do you answer that? ROI question?
Yeah, absolutely. [00:29:00] so when we talk to customers, we never talk about, it always comes up the question of cost.
Okay. 6,
6, 6. We did the POC and we ran a little bit of volume through, [00:30:00] the customer came to the conclusion that with us they needed to basically, statistically speaking, they didn't have to QC or QA anything.
Mm-hmm.
So that $18 didn't sound so bad after all.
What were they trying to build?
something In medical.
Medical using ai? Yeah, something. Okay. You need data? Yeah. Yeah. Okay.
So the answer is, it varies. so is the question is your model doing or your agent doing well enough or not?
Right, exactly.
Right.
Let's fine tune,
let's fine tune, yeah, there's a cap on how much you can fine tune.
I would do prompt engineering first and then maybe consider fine tune last. But all of these tactics are more related to the algorithm part or the architecture part.
Less about the data. Yes. Yes. I think the, in a lot of [00:33:00] cases, so I would say if I were doing, I'm not saying this is the right approach. I'm not saying this is the perfect approach, but if I were doing it, I would very quickly test.
right,
so my, my, [00:34:00] my subset is quite limited, so I'm biased, obviously. Right?
I maybe I represent more top of the funnel that you haven't talked too much with. If I know that I have a data issue, I will call you. please.
a lot of people who have serious data quality issues, they come to us and it's usually they come to us in a panic.
Ah, safety.
yeah. Yeah. a lot of our, a lot of our customers, like the first call is, Hey, I got your number from a [00:35:00] friend, and this is usually like a eight, 9:00 PM call on a Friday or a Saturday evening.
right?
So people got used to that.
I remember I saw a research from MIT saying that, 95% of the AI systems in the past two years fail in production.
So a lot actually That's, that's quite significant. I'll tell you, like I like playing ground. I like breaking things. I like building things and testing things out.
Right.
They don't have the, they don't have the infrastructure, they don't have the depth. you can demo whatever you want. Like it's not a problem at all. And I think in the middle of the AI frenzy, we have a lot of companies that are able to demo things that raise a lot of money and based on very [00:38:00] beautiful demos.
vibe coding. Yeah. Yeah. I'm using test data, so it works on test data, but a lot of people building these systems are not anticipating what can happen in real world.
Yeah. Yeah. and the other thing that I hear a lot is, oh, we're gonna use synthetic data. It's okay, synthetic data, great, but like synthetic data doesn't deal with the real world, right? Like synthetic data is not gonna [00:39:00] generate exceptions like you see in the real world.
Exactly.
there's there's a chance that we are actually having this conversation and there's a non-zero chance that I am talking to myself against a wall. Like it's a non-zero chance. Like literally this is a non-zero chance.
I know what you mean.how the real world is functioning it's not something we can [00:40:00] prompt engineer. There are subtle rules or principles that we all abide by, like whether it's physical or temporal, whatever it is, we don't think about it.
Absolutely. Absolutely. Absolutely. and even our understanding of those rules is similar to the Pythagoras theorem going back, like our understanding could be the basic understanding, but there's a lot more to it.
Right.
It's Newtonian physics all over again. Like we thought that this was the world, actually the world was a lot more sophisticated than that.
And there are, there's a whole lot more that we don't know yet. So we can't simulate that.
Yeah, exactly. Exactly. I've heard this over and over again.
Expert opinions are dead. It was like, whatcha
talking about?
Okay.
Yeah.
Great.
Yeah.As you imagine, [00:41:00] I have very strong opinions about this.
help me to, think about,a top up the funnel potential customer. Let's say. me, I'm a company, I'm building AI agents
That experience is possible to capture in quite a few different ways. So first of all, we can, you can capture exactly how you do it. You can capture how different people do it. You can capture and then how you label it. You can describe it, you can abstract it further [00:42:00] into definable buckets.
Right,
right. Maybe the morning.
that's the wisdom.
that's the wisdom. And then, then there's another layer on top of it. So you abstract it, you abstract into a bucket. Let's say the bucket. Oh, this person is,let's take some, something very simple. Let's say, you're booking travel, right? or, I don't know, gimme an example of something that you would Hmm.
maybe, let's say I'm drafting patents out of my, product. I have to draft patents and and I have all the documents, conversations with the team, building this whole thing, how it works, evaluation all the data points, the end result is some sort of patent application, let's say.
Yep. so you can approach this in multiple different ways. You are approaching it in a specific way. you have a bunch of inputs, you have a bunch of outputs. And the outputs is a document that's structured, that has very specific clauses that need to be filled in a specific way. And the inputs are whatever the inputs are, and the inputs are very specific, finite number of things that could vary.
I imagine this is needed for a patent office, for instance. a patent office might be working with many , of these products , and then they have a way of doing things, [00:45:00] right? Yes. Do you think they could use your service,
I think they could. I think they could.
So, in your process, the difference will be, you'll have experts in personal injury law to be part of the team. At least review 5% of it this way. your labeling process includes both AI and human expert so that your judgment on the insurance, payout amount will be more accurate than And to your [00:47:00]
ai
Cheaper. Okay. It's very different. 'cause I was trying to compare with how I'm gonna solve that problem was through prompting and workflow. 'cause I can talk with a lawyer and I can get your experience and prompt it through.
Yeah, we can actually, we can combine the two and that would be actually, that'd be awesome. That'll be a fantastic.
Can we eliminate human in the loop one day?
I think it's possible sometime. In the reasonably distant, in somewhat distant future, we can capture the human input without the human having to go into a specific place to give that input. Especially with a lot of devices and with wearables and a lot of devices around us, we can capture that input significantly higher without the human having to go and actually describe it.
Hmm.
I think it's, yeah, I think it's gonna take a long time, but I think it's not out of the realm of possibility.
Let's, keep watching.
Yeah.
Do you think your company and your vision is ahead of time? What is your vision?
The concept? the concept itself isn't necessarily new. The concept of having human in the loop existed for a very long time. there were companies that you could hire 50 years ago that would provide expert service to augment your existing processes, right?
Thank you Ahmed. Really enjoyed the conversation with you. I feel like I learned a lot
about Thank you, Angelina.
I've [00:50:00] talked with founders who are building like second brain building. Pure AI employees. I feel you're on the opposite side of that. You know,it's a spectrum of philosophy and how AI adoption is gonna change our lives and improve productivity.
we need both. We need both.
Yeah.
Yeah, no, of course.
Yeah. sounds good. I'll see you next time.