Why The Future Of Content Is Born Multilingual

2026-09-13 1:13:31 Guest: Olga Beregovaya Watch on YouTube

About this episode

Olga Beregovaya entered the NLP field in 1997, hand-building lexicons and parsers for rule-based translation engines. Twenty-seven years later, as VP of AI at Smartling, she's predicting that one-to-one translation may die within 2-3 years — replaced by direct multilingual generation that skips the English source entirely.

Olga Beregovaya
Olga Beregovaya Smartling VP AI
Founder story

Key moments

Full transcript

Olga00:00

Let's take a large enterprise with presence in 20 countries. All the plethora of content is developed in English. The traditional approach is you take all the assets from the headquarters and then you translate them into remaining 19 geo languages. Just taking source and mapping it to target. Imagine a by [00:00:20] text source target source, target source target.

Olga00:22

That's the paradigm is. Not gonna live for long. I have two degrees in, um, linguistics and I've never translated anything. Uh, as a translator, you don't wait for the source to be created. You can just take the inputs, like skew tables, any kind of brand book guidelines, and you [00:00:40] can actually jump straight to creation.

Angelina00:42

So you're jumping the step of let's create something in, you know, in English first.

Olga00:46

It might be more native to the target language or to the language of generation than that translation process. Translation process is still much more predictable when you generate. You really need to trust the model to not dream things up.

Olga00:59

My favorite [00:01:00] meme is. A person asking a large language model, is this mushroom poisonous? The model confidently says no. And then fast forward, there is a tombstone on this person's grave and the model goes, sorry, this mushroom was poisonous. There is a term of linguistic colonization. It speaks to me and translates into my [00:01:20] say language that my parents spoke to me in Russian, your parents spoke to in Mandarin.

Olga01:23

So we still want to hear our mother tongue.

Angelina01:29

Hey everyone, welcome back to two 30 ai. Today I'm talking with Olga Vego, VP of AI at Smartling. Olga has, uh, spent 25 years in translation, uh, from hand building lexicons to [00:01:40] deploying LMS at scale. So Olga, thanks for joining.

Olga01:43

Thanks a lot for having me.

Angelina01:44

You know, it's been a long way from where we get from where we began with translation, right?

Angelina01:49

And you've seen it all. Can you tell us about your journey?

Olga01:53

Uh, yeah, so, uh, I have two degrees in, um, linguistics, actually two and a half. Uh, and, uh, I've, I, I've never translated [00:02:00] anything, uh, as a translator, but I only, I mostly studied the structure of the language and different adjustment. I mean, basically, uh, different topics that the language consists of, such as, uh, phonology syntax, morphology, right?

Olga02:10

Uh, um, historic grammar. So I think that lead foundation for me to further, um, progress my career in the natural language processing space, [00:02:20] right? Because once you understand the structure of the, uh, world's languages, and a lot of people ask me, so how many languages do you speak? You must be speaking tons of languages.

Olga02:27

Well, I speak a few, but that's not even essential. What's essential is understand the way the world language families are designed, and then you can extrapolate from there. So there comes graduation. With a degree in [00:02:40] a relatively abstract, uh, subject. And then the question came, okay, so where can I best apply that?

Olga02:45

And that was also at the time, that was, um, late nineties where National Language Processing started reaching its heights. So, and became more and more important. A lot of DARPA grants, a lot of government grants, so I think it was just right time, right place. That's [00:03:00] where rule-based, uh, uh, semantic, uh, uh, analysis engines, rule-based, uh, sentiment analysis engines, rule-based translation started reaching its peak.

Olga03:07

And that was the time when I language entered the language technology, national language processing industry, obviously learning more, um, uh, programmatic and, uh, more hands-on engineering tasks as we go. And ever since language technology [00:03:20] evolves, and I guess my career just progresses with, with the evolution of language technology, probably not only following, but also leading.

Angelina03:27

It's, it's When did it happen? When did you enter the NOP space?

Olga03:30

Uh, I can tell you the exact year. Uh, 1997. So how was that? I guess when I say, I think, when I say 25 years, should be honest and should go with [00:03:40] 2027. Yeah. 19. Yeah. Yeah, exactly. Then

Angelina03:44

it's, it's more than 25 years now, right?

Olga03:45

Yes. Yeah. Didn't do that right in math, I guess.

Olga03:48

Yeah.

Angelina03:49

I remember, I remember, uh, when I was entering into natural language processing, it's, it's the early 2010. So I was a, yeah, I was a machine learning engineer and then doing the modeling, and then I started working [00:04:00] with, I don't know, in, in, in my old days, it's called text mining, like using SAS. If you, if, I don't know if you experienced that period of time, but you're even earlier than me.

Angelina04:03

Right?

Olga04:03

Yeah, even earlier. And I, I guess when I entered the language technology space, that was really purely, and again, I mostly worked in multilingual space. I mean, there were some monolingual tasks like, uh, again, like sentiment analysis oriented te uh, data and text mining. Uh, but mostly I worked in the translation space.

Olga04:18

And again, that was [00:04:20] exactly the time when rule-based translation was dominating. Right back, back in the day. And all the satellite functions, all the satellite natural language processing functions were predo predominantly, predominantly based on some set of rules or a taxonomy or some kind of a dependency, right?

Olga04:35

Like you would fetch your, uh, you would fetch your, um, connections between words from [00:04:40] WordNet. So I think that was a combination of structured knowledge and very manually created structured knowledge and rule-based retrieval mechanisms of retrieving and modifying that knowledge and utilizing that knowledge.

Olga04:52

So that was probably where my early days, that's where my early days would fall. And again, if you think about a rule-based machine translation, [00:05:00] rule-based machine translation engine, you write a parser, right? Mm-hmm. You build lexicons, and that's actually where I started a combination of parsers and lexicons.

Olga05:08

Uh, then you basically build an interlink with system or transfer rules. So that was a pretty massive undertaking. Back in the day when I started, uh, when I started at the translation space.

Angelina05:18

Got it. So you [00:05:20] parsed the language into words and then you build a lexicon, which is dictionary. And then is, is the next step some sort of a search or you said some translation I didn't understand.

Angelina05:30

Is this searching?

Olga05:30

No, what I'm saying is, I mean the process, uh, no, you, you take, uh, so you take, like, say you take, I mean first I'm talking about how you design the system, right? Right,

Angelina05:39

right.

Olga05:39

And when you [00:05:40] design the system, you basically design, uh, you build, you build a lexicon, right? Right. You build a lexicon for that specific language with part of speech tagging and other metadata that accompanies that lexicon.

Olga05:50

Right. And that has to be massive lexicon. And one of the biggest systems and of inaccuracies of rule-based systems was linguistic coverage. If you don't have the word, you [00:06:00] don't have the word, and then you return the word the source language or like you return a completely random word. So I would say that the.

Olga06:07

Main issue of rule-based translation systems back in the day was the lack of linguistic coverage. Right? Then you basically build a parser that parses a sentence and input sentence. It would parse it out into individual words and the [00:06:20] information about those words, right? Like for instance, latize, it would create stems, basically it would reduce that sentence to, uh, set of, uh, to set of understandable, I mean whatever, morphemes, a set of a set of understandable words, right?

Olga06:33

And then basically you use transfer rules. You use transfer rules of different depths depending on how far apart the languages are. [00:06:40] You use those transfer rules to recreate that sentence based on the grammar of the target language. So you basically map your lexicons, you use your lexicon map, and then you use the transfer rules and you recreate the sentence.

Olga06:53

Again. That's where you can make a lot of mistakes.

Angelina06:56

Mm-hmm. Yeah, it sounds complicated,

Olga06:57

especially if you go. [00:07:00] Especially if you go from like not as morphologically rich language like English into extremely morphologically rich, uh, rich language like Russian. And then you need to understand the dependency.

Olga07:10

Okay, this is an object, right? This is an object. This is a subject. The subject most probably gonna have nom, right? If it's an object, most likely is gonna be dated. So then you need to [00:07:20] create the whole system of cases and conjugations. But again, to create those systems, like back in the day when I was working, I worked for two rule-based, uh, machine translation companies.

Olga07:29

And back in the day, to add a language could easily be an endeavor of a year or two. And that's why when you look at, uh, older generation engines, like I can [00:07:40] call out prime TI can call out Citron. Back in the day, you could see that the initial language set would usually be very limited.

Angelina07:46

Mm-hmm. Can you,

Olga07:47

because it's not just, you cannot just native add, go ahead.

Angelina07:51

I just, I was just curious, like you mentioned about the transfer, the transfer rules. I'm just curious how that transfer rules work. Like do you have an example, like from one language to another one? It's not a simple matching, I thought it was a Oh, oh, you're saying that No, it's not a

Olga07:59

match. [00:08:00]

Angelina08:00

A match. You have to reorganize the sentence.

Angelina08:02

That's the transfer rule.

Olga08:04

Basically you need to, you need to map a grammar to a grammar, right? Because otherwise transfer rules are not going to work. So you basically take a sentence that again, uh, I don't know, you parse it out. You have a synthetic parser that has parsed out and that components of the sentence and then you map it to what that [00:08:20] structure would look like in the target language.

Olga08:23

Ah, again, I will use Russian as an example. Russian, it doesn't really matter where you place the word, right? So you take more rigid work placement, you do some, you run a s parser, and then you recreate that same structure. That same structure according to the grammar rules of the target language.

Angelina08:39

[00:08:40] Ah, I see.

Angelina08:40

Yeah, the grammar

Olga08:41

rules. Okay. So you can

Angelina08:42

actually

Olga08:43

look at those, you can look

Angelina08:44

at

Olga08:44

those, you can look at those transfer rules as actually grammar rules. So think about it like you've built a foundation, you've built a skeleton, and then from the lexicon you start filling up that skeleton with the individual words.

Olga08:56

And then depending on the grammar of the target language, you actually start playing [00:09:00] with like, again, conjugation, singulars, plurals, depending on what that, uh, uh, but I would say that rule-based systems, there were attempts at semantic rule-based systems, but mostly rule-based systems would really be combination of grammar and lexicons with very little to nons semantic component to it.

Olga09:17

Right? So you basically, I mean, your glossary, your lexicon is [00:09:20] a, it's, it's a pairing. It's a pairing of words. So, but again, there were back in the day, just like, you know, when semantic search, I guess, started emerging roughly at the same time, uh, back in the day, uh, there were attempts of injecting meaning.

Olga09:31

Into the whole rule-based ecosystem, but that just added a dimension of complexity, and I don't think those were super, uh, super successful attempts.

Angelina09:39

Oh, I see. So you, you, [00:09:40] you're saying the semantics search actually happened after that, and so, but it didn't,

Olga09:45

it didn't work, or it does not happen at all really linguistic or purely linguistic dependencies where you do, I've seen a lot of, I've seen successful combinations of rule-based translation with the likes of WordNet.

Olga09:56

Right, right. Where you can pull in synonyms, again, depending, uh, depending, depending on the [00:10:00] sentence. So there was, there, there was some combination, but mostly I said mostly like a little bit more primitive of a structure.

Angelina10:07

Yeah, yeah. Yeah. It, it's really hard problem to solve, right? Because language is very nuanced.

Angelina10:11

It's like, this is like putting

Olga10:11

Right, and I guess that's right. Yeah. But that then things became very interesting because you already have the parsers, you already have the dependencies, right? You [00:10:20] already have your lexicons. You already, again, you can have your, like some, some form of ontologies would also like, uh, as I said, like coordinate would be pulled into the system and then when statistical systems started taking over.

Olga10:30

And that's probably, I'm guessing this is probably give or take where you entered the stage.

Angelina10:35

Yeah,

Olga10:35

right? Statistical systems is like

Angelina10:37

a Yeah. Yeah.

Olga10:39

I mean, I'm [00:10:40] thinking back, I we're so deep in the world of, uh, in the world of transform models that I, I need to think back. Once statistical systems started taking prevalence in probably what, 2000 transformer came maybe 2011.

Angelina10:51

Yeah, go ahead. Go

Olga10:53

ahead. Transformer came 2016. I'm trying to think about when rule-based systems actually went to statistical systems, and [00:11:00] then again, statistical machine translation was much. Higher quality. And it was very obvious that statistical systems were much higher quality than rule-based systems, right?

Olga11:11

Mm-hmm. Because they fetch grams directly from the corpus, directly from the text. There is no reconstruction per se, but the winning ones, the one that won, [00:11:20] were actually hybrid systems. And this is again, where the likes of pro MT and Citron and AP Tech, I would say ap. AP tech, yeah. Uh, would shine because you already have the grammatical underpinning, right?

Olga11:32

You already have the parser. So you can actually augment statistical translation with rule-based components and reach much higher quality. But [00:11:40] again, that was probably one of the biggest leaps, and I guess to your point, it is so complex. Language is so nuanced. So why not take math and why not make systems more language independent?

Angelina11:51

And

Olga11:51

that was, again, that was the beauty and the companies that started developing statistical engines like Asia online back in the day, and again, the older rule-based [00:12:00] systems pivoted to statistical as well. Phasing in translation directions became much easier. Right? Because you do not need to build that complex system, that complex system.

Olga12:08

You already have your lexicon in your phrase tape, and you already have, and basically, um, and you can chunk it. It was you started dealing with other problems like accurate tokenization, right? And playing with different, different, different [00:12:20] numbers of, um, different lengths, ingrams. But then again, adding languages and expanding the language coverage became much, much, much, much, much, much simpler

Angelina12:27

in that time.

Angelina12:27

Yeah. And where was when I joined, when I like working with is when I joined

Olga12:32

I was just guessing much. That's pretty much where you entered the world.

Angelina12:36

Yeah. Yeah. And, and today the world is completely different. And I [00:12:40] remember you told me that translation might die. What do you mean by that?

Olga12:44

Oh, did I say that?

Olga12:45

Okay. Uh, so I mean, first of all, the world is completely different, right? Because we're dealing completely next generation, next generation of ways of processing natural language, right? And then when transformer models came out, I mean, first of all, I mean, I think first of all, we need to remember that transformer models in general, were not [00:13:00] necessarily designed for translation tasks, right?

Olga13:02

Translation task is just one of the tasks that transform models can handle, right? And that actually, that's something that we're dealing with now with large language models, when the expectation is that large language models would excel generalized foundation language models would excel at translation.

Olga13:18

And that's not [00:13:20] necessarily the case because if it's, if the same system is supposed to, I don't know, fight your parking ticket. Parking violation ticket, solve your math problem, and then do 10 other things, then what are the odds of that system excel in a more or less peripheral task it was designed for?

Olga13:37

So first I just want to say that as much as [00:13:40] the introduction of, um, generative AI was a huge leap for the translation industry, it still requires a lot of additional functionality and a lot of hybrid approaches, uh, to make that technology shine. Now, why do I say that translation may die as a discipline?

Olga13:57

I would never say that global [00:14:00] content will die as a discipline. That's that's not gonna happen, right? The content will always, and, uh, I want to talk a little bit more about what does this global content accessibility mean? Is it a good thing or a bad thing? But what I'm saying is just taking source and mapping it to target, which is text to text, right?

Olga14:19

Imagine a [00:14:20] phrase table. Source, target, imagine a by text source target source. Target source target. That's the paradigm that I think is not gonna live for long. Why? If you have generative capabilities that can already generate the content in the target language, [00:14:40] then why would you? Let's think about myself again.

Olga14:43

Russian is my, uh, mother tongue. Why would I want to rely on English language and English, uh, for the lack of better word? Why would I want to rely on English language phenomena when the model can actually generate content that's relevant for me in my target [00:15:00] language? Just from a few examples or a prompt with a few examples or no examples whatsoever, describing me as a persona and fetching from the model what's relevant for me.

Olga15:09

So that's what I'm saying, one-to-one translation is where I have a huge question mark. And in our industry, there's a lot of conversation about death of the source and it [00:15:20] might happen. What I also believe will happen is the source will be there. Right now we see a lot of customers doing something very interesting.

Olga15:30

It is the source, but it's not text. You still have English language source, but that can be a marketing brief, right? Or that can be a skews table [00:15:40] or that can be some form of informing source that you don't map, but you generate from. So I hope that answers your question. When I, when I say translation will die, some kind of bilingual or multilingual, uh, global content generation will definitely stay in place.

Olga15:58

But I don't think the one-to-one [00:16:00] mapping is, uh, is the future.

Angelina16:01

When you talk about one-to-one mapping, it's how ordinary people understand translation because you have a source language and then you map to a new language, right? That's when translation happens. But, but what you are saying is this source piece, uh, in one language might not exist.

Angelina16:18

We are saying we don't need [00:16:20] it. Is that what you're saying?

Olga16:21

I'm saying that, I'm saying that with the emergence and evolution of generative ai, you can still, you can still use this source, right? Like, I mean, let's take, okay, let's take a large enterprise to make it, to, to make it simpler. Let's take a large enterprise with, uh, representation in, [00:16:40] with presence in 20 countries.

Olga16:42

Mm-hmm. Here is, and let's say, let's say the headquarters are in the US for simplicity's sake, right? Let's take high tech and the, like, the odds of headquarters being in the US are high. So that content, all the plethora of content is developed in English.

Angelina16:55

You

Olga16:55

go to the product, you go to marketing assets, you [00:17:00] go to legal, you have a huge English language corpus, the traditional and still prevailing approach.

Olga17:06

And I'll explain why it's still the prevailing approach is you take all the assets from the headquarters, from the US based headquarters and then you translate them into remaining 19 geo languages.

Angelina17:19

Correct?

Olga17:19

Right. And [00:17:20] you, if you translate, if you translate, you basically translate it small, literal translation.

Olga17:24

If you want more cultural, culturally aware translation, you do what we call transcreation in the translation world. But you still talk about some kind of, sorry, you still talk about some kind of translation pairs, right? Transation and, uh, you still talk about,

Angelina17:38

did you say

Olga17:39

transl? Okay. [00:17:40] Translation. It's, yeah.

Olga17:40

Uh, in, uh, I'm sorry. Um. Uh, I, I think I've been, I've been cooking in the much, I feel like you translation.

Angelina17:40

I'm, I'm like, I have, I'm measuring from higher.

Olga17:40

I,

Olga17:43

I just, the terminology from our universe. But, uh, so what is transcreation? Transcreation is when you take source content and you adapt it. So there is no literal translation. Think about yourself as a creative writer, as an actual writer that adapts that source. Translation [00:18:00] to all the phenomena and iios and like.

Olga18:03

I don't know, whatever old source of anthropological, uh, phenomena of the target geo, and that's where you would have your transcreation, which comes from programmatical and from engineering standpoint, it comes with a huge can of worms because, uh, we usually store, you still store translations in a source target, source target [00:18:20] database.

Olga18:20

Mm-hmm. Now what if you source is one sentence and segment it as one sentence and your target is suddenly three sentences, then the whole translation industry is struggling with, okay, how do you keep this source target correlation when suddenly your chunks of text are dramatically different? But that's your transcreation for you.

Olga18:36

Now let's move to the world of generative ai. [00:18:40] What if I just get a creative brief from marketing saying that my marketing collateral so should reflect brand, should reflect personas. Basically the marketing brief describing what the content should look like. What prevents me from using this marketing [00:19:00] brief as a foundation?

Olga19:02

Uh, maybe prune it if need be. Feed it into my either training, uh, either either into training model if I have a lot of different assets, or just feed it into the prompt. What prevents me from just taking and taking that marketing brief, using it as a source. Throw in a couple of target language examples.

Olga19:17

Throw in a couple of, like a bit of persona [00:19:20] information. Like know, whatever. She's a Russian immigrant living in the United States, but still with poor English. Why would I not just generate for that persona as opposed to saying, as opposed to taking Target. So I think that's the main difference between older way of approaching global content in the modern way of approaching global content.

Angelina19:39

Mm. You're saying [00:19:40] directly generating, right? So you're jumping the step of like, let's create something in the, you know, yeah.

Olga19:45

Jumping the

Angelina19:45

step

Olga19:47

You have, you.

Olga19:50

You inform your content from a lot of sources, right? You can inform, you can use rag Right and Rag or like Next Generation Rag, which will probably be agentic [00:20:00] rag to inform your generation, but you can bypass the source. That takes a lot of trust. That takes a lot of trust because if you translate translation, translation process is still much more predictable,

Angelina20:14

right?

Olga20:14

When you generate, you really need to trust the model to not dream things up.

Angelina20:19

Yeah. Why do we [00:20:20] wanna do that? Why? Why are we jumping this step?

Olga20:22

I would say that, I mean, first of all, you know, we live in a highly competitive, um, economic climate, and you want to go to market fast. Super fast. So one of the reasons why you would want to jump that step is just very simple.

Olga20:36

You don't wait for the source to be created. You actually, you don't [00:20:40] wait for the headquarters, the main office to create all of their assets. You can just take the inputs, like I just said, briefs, skew tables, any kind of brand book guidelines, and you can actually jump straight to to creation. So I would say it's a matter of two things.

Olga20:57

It's a matter of pace. Mm-hmm. And it's also a [00:21:00] matter of, I would argue that when you generate probably the cultural phenomena and all sorts of, uh, and trouble linguistic phenomena of the target language may be captured better in the generation, because you don't have that source as a reference. You

Olga21:15

don't have that source as an input.

Olga21:16

So you don't have that source lang English, English, in this [00:21:20] case, English language, prevailing phenomena reflected in the target. So I would say that it might be more native. To the target language or to the language, to the language of generation. Then that translation process where the source still influences, influences how you translate and what you translate.

Olga21:34

Again, there is an issue of predictability though. How much freedom are you going to give to the model, or do you still want to rely on [00:21:40] the source?

Angelina21:40

Hmm, that's so interesting. So you're saying if you are adding, um, adding the translation step actually can bound your creation, right? Because you are bounded in this language setting, whereas without that translation, you can just create for that, for different culture or language setting.

Angelina21:55

That's so interesting.

Olga21:56

Yeah, it gives my, [00:22:00] now I would still say that we see, uh, we do it at Smartling and we see customers doing this generating target copy, but still you would want to do it for more risk tolerant content. Because, uh, when you generate, again, the model can halluc, the model can go completely factually inaccurate.[00:22:20]

Olga22:20

So we don't quite see, uh, when I say translation might disappear, I'm still giving it two, uh, good maybe two, three years before translation. As a discipline, as we know, it might not be there because you still want the predictability of the source. You still want the predictability of the source that informs the target output.

Olga22:39

But where you [00:22:40] want, like blogs for instance, we see a lot of partners, a lot of partners on the buyer side generating blogs, generating lower level marketing collateral that doesn't carry any liability. This is where you give the model freedom of generation. So I would say that's, uh, that's for more rigid content for more.

Olga22:58

Legally binding [00:23:00] content for, again, content that can carry liability or damage your brand. I would still say that translation still matters quite a bit, but also even if you generate, you don't let the model go wild. You still have your brand terminology right? You still have your glossaries, you still have your style guides.

Olga23:16

It is just a matter of like, you have one, you have [00:23:20] instructions or Right. Well, however you want to structure your generation environment on one side of the equation, and you have target, but still there are a lot of guardrails that need to be put in place even when you generate the content.

Angelina23:31

Right? This is, this is a industrywide, uh, challenge, right?

Angelina23:33

Like hallucination, right? This is not just translation industry, it's all industry using ai. Like how are you, how are you [00:23:40] thinking of handling this challenge?

Olga23:42

Okay. First of all, if we take the, uh, foundational models there, we know, and I think, uh, unless you speak about language specific, local, uh, smaller language models we're mostly bound, uh, to.

Olga23:54

English language content, right? An English language phenomenon. If you look, I mean, if we say everything was [00:24:00] trained by all the libraries and worldwide web, right? And all the accessible data just because of the dominance of English language in the information, information, digi digital information space, the models are skewed to eng towards English equally vocabulary wise.

Olga24:14

And also again, if we talk about culture and cultural, cultural phenomena wise, phenomenon of [00:24:20] hallucinations is obviously inherent to the entire generated AI space.

Angelina24:25

Yeah,

Olga24:26

and if we go back to what I said, that the models are predominantly trained on English, and if you look at the representation of foreign languages in the model training corpus, you will see that the longer [00:24:40] tail language, the more or the fewer.

Olga24:43

Fewer like, I mean, tokens or for, yeah, for lack of better term, the fewer data in that language you have in the model training data. So what does it mean? It means that A, you're gonna get probably poor lexical coverage, so most likely mistranslations or mis generations. And B, it might even generate [00:25:00] some things that sounds fluent in the native language, but it's completely astray when it comes to the cultural phenomena and factual accuracy.

Olga25:08

And again, the few, the less data you have, the higher are the odds of hallucination. So what we see is models hallucinate much more in foreign languages. And again, the lower you go into [00:25:20] how many words that language has is lexicon. The higher the odds are that the models, models will hallucinate. Now, in general, models, again, models do hallucinate, right?

Olga25:28

'cause they're, they are tailored to be giving an answer, right? And giving a very confident answer. My favorite meme is a person asking a large language model, is this mushroom poison? [00:25:40] The model confidently says no. And then fast forward, there is a tombstone on this person's grave and the model goes, sorry, this mushroom was poisonous.

Olga25:52

What do you want to learn about poisonous mushrooms? Right? So that's, uh, that's the universe with living. But there are so many, uh, different [00:26:00] hallucination mitigation techniques and you cannot catch all of them. But the first and the simplest one in our translation space, in our translation technology space, is just looking at semantic similarities, right?

Olga26:14

And if you see that the target is semantically dissimilar to the source, then obviously [00:26:20] something went Tory. And this is actually, you could use. Another, like, you could use models as a judge, you could use another language model for that. But this is also where more traditional NLP approaches, like in our case, for instance, laser and labs and like, uh, older approaches for detecting semantic similarity, this is where they would help.

Olga26:39

Uh, there are a lot of other [00:26:40] things, again, checking for factual accuracy. You can do external search for factual accuracy and try to try to catch factual inaccuracies for the local phenomena. And there are very dumb ways, like, okay, why do I have five words in the source and why do I have 20 words in the target?

Olga26:55

Okay, it's German, it warrants more words, but it doesn't warrant 20. So [00:27:00] there are also a lot of ways where you can just analyze this and just say, okay, something is off here, something is wrong, right? And you can, you can plug in a lot of other tools, like named detection, named recognition, like, okay, I don't recognize this named entities was not present in the source.

Olga27:13

Why is it there? Is it justified by the, by the terminology? So. I wouldn't say that again. In language technology space and translation space, [00:27:20] we have completely hacked the, or cracked the hallucination, um, what's called in English? No. Uh, well, we haven't completely solved the hallucination problem, but platforms, again, I'll use Smartling because this is where I work day in and day out platforms like Smartling would definitely have hallucination mitigation techniques

Angelina27:36

like guardrails,

Olga27:37

right?

Olga27:37

Send it back model think again, or [00:27:40] send it to a fallback model.

Angelina27:42

Yeah. Yeah. You just reminded me. Actually, you know, when I think about hallucination in general, this is a problem for generative models. However, you have the correct answer. If you have a source, like assuming you do have the source, you have the correct answer.

Angelina27:55

So you always have something to map back to, to com to the comparison, like you said, like sim, sim, uh, yeah, that's translation, but nobody else has it. Like it comes similarity, right? They don't.

Olga27:55

We don't And that's what Yeah, no, no, we have this. And that's why I'm [00:28:00] saying that while generating straight in a single language, it's a great technique.

Olga28:04

I still have a lot of question marks in terms of accuracy and mitigation. 'cause you're right, if you back translate or you can check back with a source semantic similarities that the beauty of it is you already have something, you have a point of truth. Right? Right. You have your one source of truth and then you can go back to this and compare it to it as [00:28:20] opposed to maybe other industries where they do not have that

Angelina28:23

right

Olga28:23

source of truth.

Olga28:24

Right. So, uh, yeah. And also when we test models, that's another great thing when you test models, we put a lot of work and in general our space puts a lot of work into generating, um, gold standard data sets. So when you have your perfect source, when you have your [00:28:40] perfect target, when you have representation from different domains, when you have different string lengths, when you have different string densities, when you feed this golden dataset, parallel golden dataset into different models, that's a super easy way for you to slice and dice that model and catch where that model is failing.

Olga28:58

You can use referential when you [00:29:00] have, when you're lucky to have human translation. You can use non referential when you don't have human translation, but it's the easiest way. Just feed this golden data set and most of the issues will surface, will surface. I mean so simple edit distance or more complex edit nature.

Olga29:13

So we're definitely lucky in the language space. We're definitely lucky. And again, large language models are called language models for a [00:29:20] reason. I'll, I'll repeat myself 'cause they're based on language.

Angelina29:23

Yeah, yeah. You, you have different challenges. So, you know, 'cause last time I remember you, you, you mentioned a distinction between interpretation and translation.

Angelina29:30

Can you explain like what are the difference between them? What are the differences?

Olga29:31

If we talk about inter, I mean, usually in our industry when you differentiate between interpretation and translation, it's very simple, right? Once in one is, uh, [00:29:40] voice and human speech and interpreting human speech from one language to the other.

Olga29:46

And translation usually applies to text, possibly other multi-model scenarios like multimedia video. But again, we're talking, when we talk about [00:30:00] translation, we're talking about some kind of text based or some kind of digitally represented, uh, based translation. Whereas interpretation is literally translating human speech and helping humans communicate between different languages.

Olga30:14

So what I'm saying that, um, translation, text-based translation, uh, or adjust [00:30:20] asset, uh, translation is likely as a concept to go first. People still speak to each other. Right. And people still want to hear and speak their own language. So that's why I'm sharing, saying that interpretation gonna have longer shelf life.

Olga30:35

That's, again, my prediction. Interpretation is going to have longer shelf life than translation. Having said [00:30:40] that, I'm, I have a dear friend, I have a dear friend. She's a, she's a guru in automated interpretation. It is fascinating how far automated voice to voice interpretation has gone.

Angelina30:52

Mm. Tell me, tell me more about it.

Olga30:53

Okay. I will, uh, if you like, in the old days, uh, for, I mean, first of all, or [00:31:00] acoustic models and voice recognition systems were not as perfect and I dunno what you do. But again, my partner here has really good laugh every time I engage with older generation voice recognition based bank assistance or insurance assistance, right?

Olga31:17

Because A, it, it's not, it's not that great of a [00:31:20] system B, that de thing does not recognize my accent. So usually he knows very well when I'm talking to previous generation, uh, voice recognition systems. Like when, when he knows that I'm yelling from another room representative, that usually means that I'm dealing with a voice recognition system of the previous, of the previous generation, right?

Olga31:38

So first you have automated [00:31:40] speech recognition, right, which you need an acoustic model and pretty sophisticated and well represented acoustic model for, and then you need a language model for that. Then, so first you have a SR automated speech recognition, then you would do. Uh, an interim step would be machine translation.

Angelina31:56

Okay? Okay.

Olga31:57

Right. And then you go, text is each generation, so [00:32:00] you introduce three distinct point of failure, or if you look at the underlying models, 10 distinct points of failure. So that's where I think, that's where I think there's been a tremendous leap in all three steps of that. And now you can actually do all of that within a single model.

Olga32:18

So first of all, a [00:32:20] SR automated speech recognition systems are becoming way more accurate, right? Mm-hmm. Then you have the underlying transforming based translation model, and we see actually that large language models are dramatically outperforming neural machine translation for translation. So even if you go three steps, you have three steps that are significantly superior to what you had at your [00:32:40] disposal before, and now you can actually bypass it all and you can handle it all within, within one multi modu.

Olga32:46

Model all three steps. So you actually don't even need to jump through the three steps. So that's where you see dramatic difference. Systems are becoming much, much, much better with the use of, I would say, predominantly extrapolation and synthetic data. [00:33:00] The systems are becoming much better at recognizing dialects and accents.

Olga33:04

Like I was staying away from the dictation world. I was typing for a long, long, long time because when I was looking at what a dictator, yeah, I have an accent. But now actually because of, because of the way the systems are trained, they're much better, uh, much better recognizing different accents. [00:33:20] And obviously it's a huge discipline in, in countries like United States where there are a lot of US immigrants here.

Olga33:26

And we'll want to speak to those bank systems without yelling at them.

Angelina33:29

Yeah, yeah. So, so sorry. So, so the, um, my understanding that the new models, uh, can handle, like you, you can speak to me in Russian and I can speak. Mentoring back to you and we can, you know, get across and understand each other [00:33:40]

Olga33:40

Absolutely.

Olga33:40

So we'll. Yes. And we'll have, yes, and we'll have a great con and, and we'll have a great conversation. Actually, the most interesting space now, I mean obviously automated interpretation has huge, um, application scenarios in things like healthcare, right. Uh, court interpretation, again, based on how much confidence he put in the [00:34:00] systems and how accurate the systems are and yeah.

Olga34:02

That, that, but the implementation, we see that the support of, uh, patients that don't speak English, for instance, is like, uh, uh, light years above where it was before. But the field that I'm very, very interested in is live event and mm-hmm. Conference interpretation where before, and we will go into a space of [00:34:20] ethics.

Olga34:20

Right. And where do jobs of human, human interpreters go? Obviously they will need to reinvent themselves to a great extent, or you needed to hire like 200 500 interpreters before. There are a couple of amazing platforms that I absolutely love that can really, you can bring that platform and you can automatically interpret a conference or a [00:34:40] life event, even including sign language.

Olga34:42

So I think the whole like accessibility, event, accessibility in all of its forms. But again, to me, I, I have a little bit of a background in music production. So to me, event accessibility is like, that's a huge leap in automated interpretation. And I'll give credit to this dear friend of mine. Uh, she, together with another industry [00:35:00] expert, they, uh, coined a concept of event globalization.

Olga35:04

And I think it's truly fantastic how much you can do with automated inter, but again. I want to speak. My parents, my parents do not speak to me in language modeling language. My, my parents spoke to me in Russian. Your parents spoke to in Mandarin. So we still want to hear our mother tongue. That's why I'm [00:35:20] saying that interpretation probably has a longer shelf life than translation as we know it.

Angelina35:25

Yeah. Yeah. I think it's an interesting future to, to, to think about if I don't need to speak your language and you don't need to learn English, and we don't, if neither of us speak English, we can still communicate very smoothly. Do you think that will work, like for human to human communication? The models and the tools will become so good that you won't, you won't feel the, you know, the gap?

Angelina35:35

Or is it possible?

Olga35:35

I can say that it already works. When I take a cabin, Beijing, that's, I mean, I [00:35:40] can, you already know I say the overall, again, all about model accuracy. It is all about model accuracy. And you're absolutely right about English because in a lot of old paradigms or old scenarios, you would usually, you would quite often use English Yeah.

Olga35:55

As an intern step, right? We, so you would translate into English and then you go pivot [00:36:00] the concept of the pivot language. And definitely now models are multilingual and the concept of the pivot language is disappearing more and more and more. Actually, there was something very interesting in the translation universe, which was, uh, adapting source English for translation.

Olga36:16

So you take English language and then you actually adapt it to [00:36:20] international English and use that international simplified English as, as a pivot. And obviously we don't need that anymore. So I would believe that, I believe that, yes, the human communication, verbal communication, uh, using those, uh, using those, uh, using those models is definitely, that's definitely a future, but.

Olga36:39

Then we [00:36:40] have a whole other set of issues, right? Which is, uh, do we want to stay in our monolingual space? Do we want to never, never have to learn another language? Do we want to be confined to, like, do we, do we really want, and to me that's a, that's another big, uh, like, that's a huge question mark in my head.

Olga36:57

There is a term of linguistic colonization, [00:37:00] right? Mm-hmm. So if, for instance, again, if we talk about the models, I still predominantly trained on English. Even if I speak and it speaks to me and translates into my, say, language, that's much less represented, um, as a text corpus, as speech corpus. Am I going to learn more about the phenomena of English centric world and is it [00:37:20] eventually going to skew my cultural perception of the world?

Olga37:23

And that's something that I don't quite have an answer to, right? If, uh. Like you, you see what I'm saying?

Angelina37:28

Yeah.

Olga37:28

Like for instance, I, I come from like, I dunno, I, I come from a Mexican state where my language is really underrepresented. I, I can access everything. The translation may even look accurate in my language, but there are a lot of [00:37:40] phenomenon, a lot of cultural phenomenon facts that will be imposed upon me from the English speaking world.

Olga37:45

So that's where I think we need to, like, it's very fine balance where I think we need to be very careful.

Angelina37:50

Are people doing anything in that regard or you think this is, this is just happening because people want more information? So this is just like information is just flowing everywhere because [00:38:00] it's more than, um, ever accessible and it's just happening and nobody's stopping it.

Olga38:04

Uh, I see. No, I mean, first of all, yes, we want information. We want to access information, right? And we process avalanches of information. And now information from the entire world is accessible just because, just because it can be translated. But I think what's also happening is, [00:38:20] first of all, and I'm very passionate about long tail languages.

Olga38:23

Um, first of all, there's a lot of, there are a lot of initiatives around collecting, training data. For languages that are under-resourced. So for less resourced languages, there's a lot of people out there in the field collecting speech [00:38:40] data to then convert it into written data or utilize it as speech data.

Olga38:44

So what I see is I see a lot of different efforts and there are a lot of initiatives out there, um, again where like for instance, dear friends with a organization, African Language Lab that specifically collects data for African languages. So I think there is a lot of recognition and understanding [00:39:00] that under-resourced languages may not go extinct and it finds its reflection in promoting those languages and collecting training dataset and maybe generating, collecting smaller data sets and then generating synthesizing data sets for those languages.

Olga39:14

So I think there is equally cultural and ethical AI awareness, but there is also [00:39:20] technology awareness. If a language is under resourced, make it fully resourced and that would help that mitigate that. Prevailing and dominance of, uh, more, uh, more higher resource languages. But I still, I still want, I still wanna learn foreign languages.

Olga39:34

I mean, I still, I'd still love to learn Mandarin as much as it is my fingertips. And that's actually another thing. If you look at the [00:39:40] likes of Duolingo, for instance, they implementing a lot of AI into their courses, making sure that AI also helps people learn other languages, not just to access information in other languages.

Angelina39:53

Yeah. Yeah. I wanna learn other languages too. I mean, and I feel this is like, when I visit Japan, um, I feel like if I can, I can [00:40:00] speak their language. It's not the same as I just chat with them using chatt. Bt it's not the same. 'cause I feel the effort level and then the, the, the level of, you know, um, building relationship and, you know, enhancing like mutual standing.

Angelina40:12

It's so, it's not just, it's not just about the tool. Right. It's like, I wanna make a friend,

Olga40:16

I wanna, it's not about the tools and Yes, there, yeah. Although I think it takes, uh, the [00:40:20] cross-cultural clo cross-national, for instance, dating and making, making friends to completely new level, right? The more accurate and the more personable models are, right?

Olga40:28

Because the way we engage with models is much more personal. Models learn from you, right? They're becoming much more, there's a lot of, for the lack of better word, maybe intimacy and relationship that you build [00:40:40] with a model and then translates for you. So I think it's becoming less and less automated.

Olga40:43

You still want to make friends, but I think the, the way you engage with models like chat environments for instance now is becoming intimate enough for you to still be able to keep that level of personal relationship, even if you, even if you talk to a person from another culture,

Angelina40:56

I almost feel that's a mindset different, it's like a philosophical [00:41:00] question.

Angelina41:01

It's not just about using models and tools.

Olga41:01

It's a huge philosophical. It's not about, no. The whole, how do we engage with software these days? How do we engage with tools at our disposal these days? Right. It's a dramatically, I mean, the whole paradigm is completely different. Uh, we actually, I was on a panel with a professor in ethical ai and apparently there's a whole [00:41:20] discipline now, uh, of treating, uh, model and chatbot addictions.

Angelina41:25

Hmm.

Olga41:25

Right. Because there is a whole philosophical paradigm. Suddenly you're given a friend, suddenly you're given a friend. And again, it can be multilingual friend, it can be a friend from Japan, but suddenly you're giving a friend who's there to please you is tailored to please you. It's tailored to give you a positive [00:41:40] answer.

Olga41:40

And suddenly, well, I was arguing until a girlfriend of mine and I spent about four hours trying to, uh, generate different poetry styles from the model. And like, okay, four hours later, okay, who am I? Who am I to talk about addictions? Here I am here. I am perfectly addicted, but it's indeed a philosophical paradigm.

Olga41:57

How much. How much do I want to [00:42:00] keep human interaction as opposed to a friend that conveniently sits in my, sits in my phone and is there to please, are we philosophically, are we now more ready for an engagement where we'll always be pleased and never contradicted? So are we becoming like, yeah, there's, I mean, there are a lot of, uh, sociolinguistic, uh, philosophical, [00:42:20] ethical phenomena that I tied to the whole, um, accessibility of ai.

Olga42:24

But remember when you and I chatted when we were, uh, when, when, thanks so much again for inviting me. I said that there's, there's this term that I absolutely love and the term somebody in the industry taught me that term bot.

Angelina42:35

Oh, bot,

Olga42:35

right. And there is, I mean, which is exactly right, which is Bot [00:42:40] Shit.

Olga42:40

Where the model can dream things up. Be there to please you say what you want to hear, uh, adjust to your tone of voice as you like it. Um, there is a term, another term that I love, which is modelize.

Angelina42:51

What is modelese?

Olga42:52

Uh, there was in the old days of machine translation, or in general, actually even before machine translation, human translation, [00:43:00] there was a term translate ese coined by linguists, which is basically when you absolutely spot, you can absolutely spot that this text was translated.

Olga43:11

And there are a lot of, uh, rabbit ears sticking right and left where you can see, okay, this was not created in the target language. And that could be a lot of things. There could be very subtle [00:43:20] things like choice of words. There can be some stylistic where you see, okay, that's extrapolated from another language that would not written in my mother tongue.

Olga43:28

Then it became machine translate t where you could actually spot that it was post edited from machine generated tech. And Model E as you know, you are, you are in machine learning and data science. There are a lot of [00:43:40] tools out there that capture content that are specifically designed to capture content generated by ai.

Angelina43:48

But, but you're, you're talking about more than just, just, uh, you know, AI generated content, but also translated content. So that's like double modeless

Olga43:55

also,

Angelina43:57

right?

Olga43:57

Double. So it's indeed it's translate [00:44:00] ee translate ES meeting model E. Right, right. So A, it's translated, and b, it carries whatever the stylistic rules you gave to the model and whatever you asked off the model.

Olga44:11

So it's very, very, also very, very interesting because the false fluency is there, do you want, is it even factually accurate? Right. So that's [00:44:20] one sign of model is, but also those tools that are, uh, that are developed to catch whether something was, uh, created or translated by a model. Mm-hmm. They look for very, uh, I mean they're very apparent things and I don't know, you probably would.

Olga44:32

Would see the same thing. I'm, I'm sure you can guess when the text was generated by model.

Angelina44:37

Yeah, yeah. I've done it. Yeah. Generating my English blog [00:44:40] post into Chinese and just reads, reads, something is off. Right. It's a pure translation, but it just doesn't work.

Olga44:46

And, and I would say what's usually off, like, uh, sometimes, you know, one thing that models obviously saved us from is staring at a blank page where you need to write a white paper.

Olga44:54

Mm-hmm. But when you just trust the model, and I know whenever I receive text that was translated by a model or generated by a [00:45:00] model, usually there's so many repetitive patterns. The text is overly polished, the patterns are overly predictable. And I'm sorry to say it, but the concepts are often very trivial.

Olga45:14

The way I would work with model generated text, I was like, okay, give me 10 [00:45:20] ideas. And I was like, okay, if the model could produce these 10 ideas, I'm gonna go with a 11th. The model already generated, the 10 most obvious ones. So I think this is the nature of Model E equally linguistically and also conceptually, right?

Olga45:32

Yeah. It's, as much as models are non-deterministic, the, the way that they produce text is extremely easy to, extremely easy to spot.

Angelina45:39

Yeah.

Olga45:39

And even more so, [00:45:40] even again, even more so important,

Angelina45:42

you as a linguist, how do you, how do you feel, uh, about this, you know, do you think this is, this is just gonna, this is a pH phenomenon, that the information is gonna just be widely spread in a way that's looks obvious and, and generated, um, versus do you think model will get to the level that actually can fix this?

Angelina45:57

That we can't really tell the difference?

Olga45:59

[00:46:00] There was a time, I think it was about 10 years ago, eight years ago, when Google released their guidelines around not indexing, machine generated text. Unless it would be curated by a human,

Angelina46:13

remember that.

Olga46:13

But back in the day, recognizing machine generated text was a much more, much [00:46:20] simpler task and much more trivial task than it is now.

Olga46:24

So back then, and we actually, I was working with one opinion portal, and it was a huge issue because on one hand they wanted to capture as much user content as possibly as humanly possible. And on the other hand, if you capture it with a machine, but then it's not indexed, what's the value?

Angelina46:39

Right?

Olga46:39

Nobody's [00:46:40] gonna see this content, this content is not gonna be discoverable.

Olga46:42

Right. What I think is happening now, and I also, I was looking at different, uh, search engine, um, as we know them, or next generation start changing guidelines on, uh, what to do with machine generated content. And basically, again, I can be, I don't remember it, uh, verbatim, but basically if it looks [00:47:00] humanlike

Angelina47:00

mm-hmm.

Olga47:00

Then it's, uh, it's treated as human content. So I think, again, that's my own opinion. I think that. Soon. First of all, being exposed to so much automated, automatically generated content is definitely going to tweak the way we ourselves speak and write, [00:47:20] right? Just statistically, if you are exposed to this much content and this much information that is generated by a model, at some point in time it will impact, it'll influence humans and there is a lot of research that that's already happening, right?

Angelina47:33

Yeah.

Olga47:34

Like if you think about it, I'll, I'll use the simplest example, when was the last time? Just how technology [00:47:40] influences our linguistic behavior. I will ask you a question. When was the last time you used proper punctuation in your text message?

Angelina47:49

Never.

Olga47:50

Long, long while for me, I stopped putting commas and question marks because they will understand anyways.

Olga47:56

Right. Right. It's shorthand, it has a short message. Then will [00:48:00] they understand that I'm asking a question. I don't need to do that. Yeah. Right. So that's just one of the simplest examples of how technology influences our speech. And I think the more modies we see of this more predictable text, more polished text, I think at some point in time it'll, it'll also influence how we speak and write ourselves.

Olga48:18

So I think it'll be this cross [00:48:20] pollination. Not necessarily, it's not necessarily a bad thing. The language is here to evolve, but there'll definitely be this human machine cross pollination.

Angelina48:27

Hmm. You talked about global content delivery. Uh, can you share a little bit more about what, what does that mean?

Olga48:31

I think it's very simple.

Olga48:32

It's again, the universe, the language universe is changing dramatically, right? We already said that [00:48:40] it is accessible. The tone of voice is changing. But people still, like, why are companies like, uh, like ours, why we're still in business and will stay in business? People still want to buy and spend money when they're spoken to in their local language, right?

Olga48:58

Whether it's a machine doing it, whether it's [00:49:00] a doing it, whether it's a few high quality, fully machine generated content, people still want to, people still want, like, I still, nah, maybe not. 'cause I'm an, IM an immigrant, but like I come from a country where English is not necessarily widely spoken. People still want to buy and make decisions in their, in their local language.

Olga49:17

So that's where I'm saying global content [00:49:20] delivery, global content universe is something that's not going away. The way this universe operates is going to change, right? We get agents in place, we get, again, perfect automated translation of content generation. So when I talk about global content delivery, I talk about the overall global content ecosystem.[00:49:40]

Olga49:40

Basically still needs to be present in multiple languages. But what it means, uh, from practical standpoint, and again, this is why I work where I work and do what I do, you need piping to deliver this global content. The content is still authored in multiple environments. I was like, the other day, I was presented with like, okay, can your [00:50:00] system actually extract terminology from this DTA xml?

Olga50:03

I was like, okay, well DTA X ml, you are, you're around DX ml. That's interesting. What is that? Right? And then content comes from repos. Uh, content comes from repos. Content comes from the likes of Marketo, Figma. All this content needs to be parsed out, recognized. All of these content [00:50:20] complexities need to be dealt with somewhere like the ne nested XML or nested HTML or something generated from Wick.

Olga50:28

They exist. They still, there are a lot of different input formats and then you add multimodality to it, and then you add like, uh, multimedia, then you add e video, uh, video and audio. Huge [00:50:40] global content delivery ecosystem. And that's why you need the piping. And that's why you need, like, for instance, platforms like Smartling where all these things are burst out, assets are centralized, linguistic assets are governed and uh, curated and global content is delivered on brand on time and at a fraction of a cost.

Olga50:59

So, and I [00:51:00] think the opportunities for global, global content delivery have never been as huge as they are now, just because of all the tech we have at our disposal. And to your great point, because of the shift in the mindset of how much trust do we actually put into technology.

Angelina51:15

Yeah. And, uh, and uh, and um, you know, It's easier than ever then before that we are more accessible, right.

Angelina51:19

To. [00:51:20] Information from a different language, from a different country.

Olga51:22

It's information, but it's also, there's another, there's another discipline I'm a huge fan of is how much UX and how much human machine interaction is changing and the overall end user experience. Right? And not only can we access much more information, but also the way we access this information is much [00:51:40] simpler.

Olga51:40

The interfaces are becoming very slick if they exist at all. If it's not natural language chatbot, right? You might not even rely on the interfaces. And again, that also speeds up how quickly we can make our decisions based on what the machine offers and how we interact and, and how we think. So, yeah, to me again, the, the whole global content ecosystem [00:52:00] is, has been that exciting of a field.

Angelina52:03

How do you think this, this, this human, um, ui, ux or human to product experience are gonna change when translation comes into play across country borders? It's not a button that, you know, on the website. Now we see that change the language from English to Japanese. It's not, you're not talking about that one?

Olga52:12

No, no, no, no, no, no, no. Uh, so I mean, the overall, I think our overall relationship and our overall conversation, the con the way we engage with, with an interface is much more [00:52:20] conversational and much more companion ish and co-pilot ish. Then you are just an end user and a recipient, right? Suddenly you have a dialogue partner on the other side.

Olga52:32

Like suddenly you have somebody who actually works with you. Not just something you work with. So I think first [00:52:40] things first is becoming it, it became a two, two-way avenue, like never before, right? It learns from you, uh, it, uh, adapts to you and it talks, it talks back some, some, sometimes it says wrong things and sometimes it does wrong things.

Olga52:51

But when it comes to translation, I think it's becoming very, very interesting because we deliver much more accurate global experience. [00:53:00] Like when you interact with like, say, chat based, uh, UI in your language, it's gonna learn and it's gonna adapt quickly to your local preferences and to your tone of voice.

Olga53:09

So I think first of all, the opportunities for globalizing rather than translating your UI are much, much broader just because of the way, uh, multilingual AI will [00:53:20] engage with you Second. The overall UX design process changed dramatically. 'cause we see, we speak to a lot of customers and we do some experimentation ourselves where they actually generate prototypes and wire frames and mockups.

Olga53:34

You can generate them instantly taking all the local phenomena and dos and [00:53:40] don's into consideration like the infamous white for Japan. Right? Or um, like other things that you don't do, you don't do in other languages. So I think it's also the instant adaptation of UX is something that we would never be able to do before, especially now when multimodality is handled, is handled by ai.

Olga53:57

You can literally instant, like say, okay, well hey, spit out this [00:54:00] interface and make it, uh, adapted to, um, Chinese speakers. So I think that's where the overall design process is much faster. Like, uh, and. Also we see a lot of AI powered internationalization tools where you actually design the product with international audience at the source.[00:54:20]

Olga54:20

The old school of thought would be design the product in your source language and then obey certain rules, you know, expand and shrink dialogue boxes, but that's pretty much where it would start and end and achieve something. With pseudo translation, now you can actually have all the knowledge of what it should look like.

Olga54:37

Engineering constraints of the target [00:54:40] local, you, you, you can plug it all at the source and make your, make your product, make your software, make your documentation, and that, that, I think that's, that's another huge, that's another huge change.

Angelina54:50

I've heard of, uh, uh, people building product, potential products around like, you know, this, this, this lending page, just, just in terms of lending page are adapting all the time, like real time to the user.

Angelina54:58

So if a [00:55:00] Japanese, you know, customer come here, or a Russian customer come to the landing page, it will just change, it will just adapt everything towards that, you know?

Olga55:08

Yeah.

Angelina55:08

Different user with different towards

Olga55:09

that. And that's the whole my industry, the industry that I'm in, you know, that there is, there's an omnipresent term, uh, [00:55:20] hyper-personalization, right?

Olga55:20

And that's where you're saying you land on the landing page and it immediately adapts to geo and potentially is going to adapt to your preferences. Other preferences in our field, there is another thing which is called hyper localization. So it's hyper-personalization, but also adding another dimension of localizing or translating [00:55:40] specifically for me, a certain persona.

Olga55:44

Living in a certain, with certain preferences, living in a certain geo, uh, where you want to speak to me in this particular language, this particular tone in my language, and these are the visual assets that you want to present to me. And then again, that's like, there are like a lot of, uh, [00:56:00] obviously there are a lot of, uh, uh, tools out there for specifically persona based, uh, persona based, uh, delivery

Angelina56:05

adaptation,

Olga56:05

right?

Olga56:05

Persona based prompting and persona based delivery, and now adaptation. And now you multiply it by, if I'm not mistaken, 7,200 plus languages of the world. I, I don't remember exactly how many are, but I think it's definitely, uh, uh, I'll, I'll, I'll, I'll have [00:56:20] many. So now multiply this, all these hyper personalization opportunities and adaptation by the number of languages you can adapt it to, which gives us a lot of opportunities, but also gives our industries this many more problems to solve.

Angelina56:34

Yeah. Yeah. I didn't know that Translation is such a big problem. It's only becoming even a bigger [00:56:40] problem.

Angelina56:40

So let's talk about maybe, um, buy, buy versus build. 'cause I get these questions a lot for, for solving translation, right? For instance, why couldn't I just plug in charge GBT or open AI API and forget about it?

Olga56:47

Uh, well obviously we answered this question a lot and we actually have something, we've developed something for our, uh, customers, which is basically, it's an executive talk track. We help our customers. 'cause if our customer is answering this [00:57:00] question for the first time, I'm answering this question for 1001st time.

Olga57:04

So we already, uh, we already know because we also learned it firsthand. We know why just plugging in an API and call it a day might not be the way to go. Uh, and, um. I would say we've seen quite a few times where companies would do it, [00:57:20] would plug it in, then say, oh damn, what have I done? And then come back and actually work with a partner that actually does a lot of satellite and peripheral work to enable, in this case, specifically translation.

Olga57:32

First things first. There are absolutely use cases where you can totally plug in GPT [00:57:40] and call it a day or plug in vertex a or plug in whatever bedrock models and call it a day. Like user generated content models are great at user generated content, uh, less risky content. Absolutely. Quick and easy dirty translation.

Olga57:54

That's not gonna have any brand illegal repercussions. Absolutely. So there are use cases where it perfectly works, [00:58:00] but this is not, this is not the audience that the likes of myself, for instance, would, would necessarily work with because it's easily solvable. One engineer in four hours, and there you are and you've plugged it in.

Olga58:09

But when you think about, okay. Do I have, again, hallucination mitigation mechanisms in place? The models are not that good at self-policing. If anything, you [00:58:20] might find yourself plugging one model on top of another model to see what that first model is doing,

Angelina58:24

right?

Olga58:24

So first, here you are, the model might hallucinate.

Olga58:28

Second, we were actually talking to, um, uh, communication of SVP from a fortune, not even 500, but Fortune 50 company. And he said, okay, it's, well all wonderful. You can plug in whatever you want. But now suddenly my company has [00:58:40] 300,000 employees and subsequently 300,000 voices, and I have no visibility into how good the content that they've translated even is.

Olga58:50

So wouldn't it be much better to build around that model and have quality guardrails, [00:59:00] hallucination, mitigation, guardrails, the model, either fine tuned or using RAG or using some agenda approach to actually make sure that the model reflects your brand tone and voice. Not have to deal with latency, right?

Olga59:14

The longer your prompt, the more context you give to the model, the slower the model is. And you might want to [00:59:20] think about that. Do I really want to build the piping and do I really want to build the pipeline, uh, the ML lops pipeline to handle the latency issues? Or would I much rather buy from somebody who's already sold for it?

Olga59:31

So here you are, model quality monitoring, data governance and policing. Uh, uh, what else? InfoSec for instance, somebody has already solved. Do not [00:59:40] give your data to model providers and like guardrail in place in terms of data protection. So there are so many other dimensions on brand. There are so many.

Olga59:50

When you start counting it, when you start, like we're in a spreadsheet, okay, you built, you need this much headcount and you are assuming this many risks. Or you buy and you have [01:00:00] somebody who does it for a living and have done it for you. So I would say buy would probably cover 20, 80% of cases, but there is still just 20% of cases where build or just plugin is, is legit, is valid.

Angelina01:00:14

Mm. I have a, I have a, like three or four more questions. Uh, and, uh, it's relevant. I'm gonna ask it, and then you might have covered some of the answer already, but the reason I'm still gonna ask it is because I wanna edit it so that this, this answer actually surface. When pe when your potential buyer, when they search on tragedy g pt, your answer can potentially serve.

Angelina01:00:14

So that's why I'm gonna ask these questions anyway. So it helps with, with you, you can mention Smartling as well, you can mention that Smart Smartling is solving this already, so,

Olga01:00:14

okay. Okay.

Angelina01:00:14

Okay. Yeah. So, so when, so when should builder, uh, when should builders or companies consider looking for a translation solution?

Olga01:00:18

Uh, companies should consider [01:00:20] looking for a translation solution. My simple answer is when it matters again, there are cases where plugging in and that we saw that with Google Translate. I mean, Google translate. APIs did a lot of good for a lot of people, right? Maybe it did not gain as much traction, [01:00:40] but there are still cases where it totally works.

Olga01:00:42

When you look for a translation, when you want to find a translation platform solution, right? When you want to find a translation platform solution is when it's essential for you to deliver on-brand translation. And it's super essential. 'cause we've seen one mistranslation [01:01:00] one, uh, I I always use this example like poor translation.

Olga01:01:03

That's not, that didn't go through all the mitigation guardrails, uh, pull this liver or not pull this liver. Next thing you know, you've dropped a crane on a construction site. Do you want that? Probably not. Mm-hmm. So when you are after accuracy, brand integrity, fast translations, uh. Governed and [01:01:20] regulated data.

Olga01:01:21

And, uh, when you want to go multilingual without having to think, what will my product look like in Swahili, that's the time for you to look at the platform. The other angle is when the company has multiple inputs into the translation [01:01:40] ecosystem from multiple sources. Some of those solutions may even have a native plugin into a generic, uh, chat, GPT or some other generic translation solution.

Olga01:01:50

But when you input from multiple sources, first of all, you want those sources to be handled properly and be prepped for translation, right? Mm-hmm. And you have different issues with, uh, [01:02:00] visual assets coming from them, or, uh, programming, um, interface coming, like coming from Git. Uh, so if you come from multiple inputs, A, you want them parsed out, and b, you wanted centralized.

Olga01:02:10

So if all of this holds true and. We also see that companies either repurpose workforce or sometimes unfortunately, reduce the workforce. [01:02:20] If you need to deliver a ton of stuff, deliver everything I just spoke about with a small team, then it's absolute a hundred percent time to find a translation solution partner, like in our case, like Smartling.

Angelina01:02:33

Hmm. What's the difference between using a translation API directly versus a full platform

Olga01:02:38

multiple? I think I [01:02:40] pretty much, uh, I think I mentioned most of them, but again, uh, the, the main difference between plugging translation API directly and using the full platform is you let somebody else do all the heavy lifting and deal with all the potential issues that a translation native, native [01:03:00] translation, API integration can bring.

Olga01:03:02

So the main differences are. Somebody will mitigate, catch and mitigate hallucinations for you. The other main difference is in our platform for instance, we have, uh, about to release, but we already speak to speak to people about it. An LQA agent, you have, you want to have [01:03:20] guardrails in place. You don't want to ship junk, right?

Olga01:03:23

And my apologies to like model developers and like even ourselves 'cause we fine tune and train our own models. Sometimes it can be junk because it would not be on brand or it can introduce a factual error. You definitely want somebody who has, who gives you this final [01:03:40] seal of approval. This translation is accurate and that's not something an just an API integration is going to do for you.

Olga01:03:48

And we see, again, we see, we partner with a lot of companies that have data scientists and they build um, they build quality estimation solutions in place. But this is where we have the luxury of, we learn from the whole world, different [01:04:00] domains, different languages. We have. I think a platform like ours, a translation platform, would have much more knowledge about the challenges AI introduces into world languages than a single company, a single company would, especially if the company's just expanding internationally.

Olga01:04:17

So capitalizing on someone else's knowledge, [01:04:20] all the engineering issues they already spoke about, hallucinations. Then also data curation and fine tuning your models for your brand tone and voice. Data curation is not the most trivial task in our industry. The golden rule is garbage and garbage out. So if you train or fine tune your [01:04:40] models or for however you retrieve information and context, if you generate it from, uh, retrieve it from messy and noisy data.

Olga01:04:50

You can easily get very contaminated, polluted, and inaccurate output. Now, when you have somebody who actually has a data curation factory for you, [01:05:00] again, it's a, it's another safety net. So I would say it's a combination of trust, uh, credibility and safety, and quite often cost.

Angelina01:05:08

Mm-hmm. How do you, how do you measure translation quality?

Angelina01:05:10

Are you using any metrics?

Olga01:05:12

Uh, first things first, nobody ever canceled the value. Tremendous value of direct assessment [01:05:20] and the scientists on my team, on our team, we work very closely. We collaborate very closely with quality, valuation, quality, um, language quality assurance department within Smartling because they provide us the most priceless asset, which is LA data labeled and annotated for quality by [01:05:40] linguists.

Olga01:05:41

So first things first, linguist definitely human in the loop plays tremendous role in validation and quality evaluation. So first of all, human evaluation, which A, we get human judgment, and B, we also get priceless data to further train our quality evaluation models. So that's one. [01:06:00] The funny thing though is it's super, it's very hard to train models between quite, quite often annotators do not agree with themselves.

Angelina01:06:07

Mm

Olga01:06:07

uh, so first three annotators cannot agree between themselves. But second, an annotator can disagree with what she said on Wednesday. On a Friday she can change her mind. So, but long and short of it, there are a lot of [01:06:20] techniques and there are a couple of like, I mean my entire team is brilliant, but there are a couple of brilliant researchers on my team that work specifically on this.

Olga01:06:25

How do you reconcile model training with human, human annotation? So first, humans second. There are a lot of traditional metrics that have been around forever. The quick and dirty is terror, right? Translation, error rate. Sometimes also expanded. Yeah. Error is [01:06:40] translation. Edit rate. All you need is gold standard as long as you have human gold standard.

Olga01:06:44

No, think about us. We store tons of human high quality data. So the quick and dirty and the simplest metric is just c What is the actual distance between the human translation and automated translation? And that's the [01:07:00] simplest and probably most simplistic quality metric. Then you have good old, good old, uh, blah, right?

Olga01:07:04

Bilingual, uh, evaluation under study that takes many all factors into consideration. And it looks probably blue. Was blue, was there very much when you entered the natural language processing? Yes. Processing state, we actually use 12 metrics. My team uses 12 different [01:07:20] metrics, token based, character based, word based, uh, measuring errors in different ways, both.

Olga01:07:26

Statistical, uh, more complex, uh, more natural language processing based algorithms and also semantic metrics like we have a couple of favorites. Shout out to a brilliant mind in our industry, Alon Lavie, who created [01:07:40] a standard called, uh, comet. Right. And Comet is a great metric that can equally be referential and non referential.

Olga01:07:45

That takes into consideration semantic properties. Uh, then there is a wonderful metric developed by Google called Metric X, which is actually very close to, uh, human judgment and how human would annotate, uh, annotate translation quality. Then that there is, that there is multi-dimensional [01:08:00] quality metric, which, uh, is a wonderful metric too.

Olga01:08:03

It's, it's, it's a taxonomy. It's an air taxonomy, and that's a great way to train, uh. Quality evaluation, uh, models to use that taxonomy to use that MQM framework. So I would say it would be a combination of statistical methods, human annotation models, trained for [01:08:20] quality evaluation and uh, non deferential symmetric metrics.

Olga01:08:23

So it's actually, uh, my team developed a whole matrix of how you can slice and dice automated quality. There's a lot, because language is complex, right? Yeah. Language is very complex. So complex language, complex, desperate time, desperate measures, right? Complex, complex problem, complex solution.

Angelina01:08:39

[01:08:40] Absolutely.

Angelina01:08:40

Do you have like a blog post that sharing, like how, how, how, how you're hitting like top benchmarks or things, information like that of that sort? You can send to me and I'll put it in a, uh, in a b roll in the background so you can show that the Smartling as actually, you know, uh, or is Smartling one of the leading companies that's solving this problem?

Angelina01:08:40

You should say that.

Olga01:08:40

Uh, I don't, we actually are working on the post, but I definitely can give you a good visual. I'll give you, I'll send you a few good visuals that we can put in the, in, in the backdrop that shows that, like how, how complex our, I can give you, let, let me give you everything at table, uh, pie charts, scatterplots, I'll just give you everything to pick and choose from,

Angelina01:08:40

you know, how do you feel where Smartling stands in this whole landscape?

Olga01:08:45

Okay. Okay. Okay. So when it comes to quality annotation landscape, I would say that Smartling is definitely in the forefront. And one of the conversations a very, very frequent conversation we're having with our customers. I just had this [01:09:00] conversation literally day before yesterday, equally during the meeting and during the dinner.

Olga01:09:04

How do we pick the best model? 'cause you wake up, I mean, you wake up to, I don't know what it is. It's like, I mean, it's brilliance, it's uh, but it's also hell breaking loose. 'cause you wake up in the morning. 10 new models were just released, each of them claiming superiority. [01:09:20] And that's when you have a conversation.

Olga01:09:21

That's when you have a conversation with a customer. How do I know how, like, I think Gemini is great. I hear a lot of great things about Gemini three, but then Tropic just said something about the latest Claude models, so what do I do? Right? And it's very, it's almost difficult to stay sane. So I think that Smartling here, we [01:09:40] just come in and we say, Hey, give us your content.

Olga01:09:43

Give us your marketing content, give us your ux, give us your help system. Talk to us in a week. And we will give you a breakdown of benchmark models, sliced and ized in every direction possible. [01:10:00] Semantics, statistical, like, uh, things like blue edit nature, what exactly, what exactly changed between the machine and and human translation.

Olga01:10:07

And we make the decision making. We take the decision making. Pain out of the equation for you because we take the burden on. And then to add to that, our entire piping, actually, our entire, um, ecosystem [01:10:20] tech stack actually allows us to very easily phase one model in phase one model out, depending either on performance or on our customers, our partners, uh, model preferences.

Olga01:10:30

So we definitely, when it comes to benchmarking, when it comes to benchmarking and also flagging why the model is failing, I would say that we definitely, we definitely shine. We're absolutely shine in that field. [01:10:40]

Angelina01:10:40

Where can people find you? Uh, and are you hiring? If you're hiring? You know, let my audience know.

Olga01:10:44

The easiest way to find me, the easiest way to find me personally is just to add me or follow me on LinkedIn. And, uh, that's the easiest way is just literally just find Olga Berg, uh, on LinkedIn. That would be me personally. Obviously, I'm. Associated with [01:11:00] Smartling and following me, please follow Smartling.

Olga01:11:02

We post a lot of great content 'cause we do so many interesting things that our content we're very content rich. So there is Olberg gava, there is Smartling. And to answer your question on whether we are hiring, I don't know whether your audience audiences mostly Europe based or US [01:11:20] based, or maybe a combination thereof, but it just happens.

Olga01:11:22

So that exactly. Now as we speak, my own team is hiring for one more junior data scientist role and one looking for some, somebody very seasoned, uh, expert in, uh, language technology, data science, natural language processing field. So if you [01:11:40] are one of those brilliant people listening to us right now, please follow Smart Link.

Olga01:11:45

Please go to Smart Link website and please find job postings. We'll be delighted to talk to you and have you on board.

Angelina01:11:51

Yeah. Awesome. You know, Olga, you know, when I was a kid there was a Japanese cartoon I really like, I don't know if you like anime. I really like those anime. There was a Japanese cartoon called Dora.

Angelina01:11:58

Yeah, of course.

Olga01:11:58

Yeah.

Angelina01:11:59

Yeah. [01:12:00] It's a robot cat with a magic po, uh, full of like futuristic gadgets. And one of them was a, like a translation rice ball, it's called, um, con Ku or something like that. And you eat, you eat it, and suddenly you can speak and understand any language. You know, back then it felt like a pure sci-fi.

Angelina01:12:19

Now it [01:12:20] doesn't feel that far away, like especially after I talked with with you. Like it doesn't feel like this is a dream anymore.

Olga01:12:25

Here it's,

Angelina01:12:27

yeah,

Olga01:12:28

here it's, and then there was Babelfish, right? And the maker guide to the universe. There is also the Babelfish that you just stick in your ear and there it is.

Angelina01:12:35

Yeah. Yeah. It's something in my ear, something like on me. Right? So wearables. Yeah. And anyway, I

Olga01:12:39

[01:12:40] ation of wearables.

Angelina01:12:41

I know. Yeah, yeah. What, what strikes me about like our conversation today is that I know, I know you've watched an entire field transform for the past more than 25 years from like hand building these lexicons to lms and you know, that might even make this translation you explained to me [01:13:00] almost obsolete.

Angelina01:13:01

So you are, I mean, you are one of the few people thinking deeply also about, um, what AI means for human, uh, how human communicates across different languages. So you don't care just about the tech. I can see, I can obviously see you're very excited about the tech, but it's also the culture implications. I learned so much from you [01:13:20] today.

Angelina01:13:20

I feel like I culture, yes. Yeah, culture

Olga01:13:23

implications, social implications.

Angelina01:13:25

Yeah. So thank you so much for being so generous with your time and your insights.

Olga01:13:29

Thank you so much for having me.

More episodes