Mistral AI has just come out with a brand new AI model, and it is called VoxTroll.
Now there's a bunch of interesting things about this model in particular, one that is an open speech model, the benchmarks and the data essentially of how its word error rate and its price I think are particularly interesting.
I'm going to break all of this down, especially on the heels of a potential $1 billion round of funding that Mistral is looking at doing and all of the rumors swirling about Apple acquiring the company.
What the what the founder has said about that, what the plans are for the future of this company, this is an interesting time for Mistral, no doubt, we're going to get into all of that.
But before we do, I wanted to mention if you want to try out all of the latest models from Mistral, including all the latest models from a lot of different companies, companies, the top 40 AI models, you can go check out my own startup, which is AI box .ai.
Over there we have code stroll, mistral 3b compact, pixtrel, large vision, mistral small, a ton of these interesting, well, all of the mistral models, but a ton of interesting models from Anthropic, Cohere, deep seek, Google, meta, Microsoft, NVIDIA, open AI, Quinn, grok from xai, all of the all of the interesting companies.
So you can check out all of those models, and a bunch of other their image speech and audio models at AI box .ai for one subscription, 20 bucks a month, you don't have to add subscriptions to all of these different platforms, you could try them all out.
And there's a link in the description.
All right, let's get into what mistral is doing.
So with vox troll, the most interesting thing that they've essentially announced is that vox troll is this open model, it's going to do transcription.
But basically, what that means is like it can take an audio files understand what they're what the audio files are saying and respond so it has its own voice it's a competitor to 11 labs and a lot of these other players and what they're saying about it is they've actually built three models specifically with three different use cases but what they're saying is this is a much more efficient this is way cheaper than using something like 11 labs or other players so this is pretty interesting they have this kind kind of a diagram that they have its price USD per minute.
And they on the other column, it's the word error rate.
So basically, how often it messes up the words that it's saying.
And they have like a couple competitors, they put on this diagram, one of them is scribe, which is like, super expensive.
And then on the other field, they have they have some others, which are interesting, essentially, the fact that it is an open model.
So they're allowing you to take the model, run it locally on your own devices or server.
And I think for a lot of companies, this is, you know, this is quite exciting to have that capability.
When you look at, you know, right now, if you want to use these text models, you got to use something like 11 labs, open AI has some options, but it's all it's all going to be things that you have to, you know, pay for API usage.
So when you have the open models, it's it's pretty interesting, being able to try and run them locally for a lot of companies.
And they even have a super stripped down version that essentially allows you to run them locally on a device.
And so this is kind of what a lot of people were saying that Mistrol was going to be doing with Apple.
This is why Apple wanted to acquire them is because they have a bunch of these tools that are stripped down and able to run on like on device.
You could imagine a tool like this would be incredibly useful for something like Siri, where you could run essentially an edge model.
So they have this one in particular called vox troll, mini and vox troll mini is the error rate is not the best. It's better still than whisper large v3 from open AI.
And it's still a little bit better than Gemini 2 .5 flash.
But it's, it's not as good as GPT 40 mini transcribe, but but it's, it's way cheaper.
And it's it can run on your device.
They also have one called vox troll mini transcribe, which is also super cheap and has a much better word error rate.
So in any case, they have all of these different models specifically that they have, and they're able to run locally on devices.
So for Apple, for iPhone, they could essentially grab one of these models if they acquired their company or maybe make a partnership with them, put it on your iPhone, use it to power Siri, and even without the internet, Siri would still be able to understand what you're saying.
They probably have another model to back it up, maybe something from Mistral or from another player, maybe an open source model from Mistral, But using this in conjunction with that, they could essentially run Siri with no internet, which would be really, really crazy.
And I think that'd be something that Apple would be interested in doing.
So people have essentially been talking about these rumors that Apple is interested in acquiring Mistral.
The CEO of Mistral said they have no interest. I mean, they weren't specifically talking about Apple, but there's like, we have no interest in being acquired.
They said that they would like to IPO the company, essentially.
And Mistral really is kind of like the crown jewel of Europe.
It's the number one AI company coming out of Europe.
It's raised the most money.
Europe as a country has backed it and given a lot of resources, whether that's compute or special deals, essentially.
And so I think they've been like they've largely benefited a lot from a lot of programs in Europe.
And so I think people want to see it stay owned and operated inside of Europe.
it. But overall, definitely, it's building a lot of really interesting tools that would be very useful for a lot of people.
What's interesting, Mistral says that VoxTroll can transcribe up to 30 minutes of audio, because it has the LLM's backbone, that Mistral Small 3 .1 can understand up to 40 minutes of audio, which is honestly fantastic.
I mean, for a majority of all conversations I ever have, it's going to be less than that.
So essentially, you can ask questions about audio content, you can generate summaries, you can turn voice commands into real time actions, like calling API's or running functions.
It's also multilingual.
So you know, I mean, you can imagine a lot of these cases, it's like you upload an audio file to it, probably less so the live talking is not what this is used for as much but you upload an audio file to it, and it can understand what's in the audio file, you can imagine something like a big use case of this technology would be like YouTube, where you have the transcription of every single YouTube video on the side, YouTube is using their own transcription models for this.
Obviously, Google has their own tools.
But you can imagine like other players that aren't Google that don't have that massive tool would need to use something.
I mean, maybe even something like Vimeo or any another like video kind of platform out there or companies that just want to have transcriptions for or transcribe a lot of their content on their platform, I mean, Facebook and LinkedIn, and all of them need that functionality.
So you can imagine there's a lot of people that need that functionality.
So it's multilingual, it does English, Spanish, French, Portuguese, Hindu, German, Dutch, and Italian, which is a bunch of languages right off the bat, which is pretty cool.
They have, of course, two variations of their speech understanding model.
So they got VoxTrail small, and that is a 24 billion parameter for production scale deployments.
It's competitive apparently with 11 Labs Scribe, although it's way cheaper.
It's also competitive with GPT -40 Mini and Gemini 2 .5 Flash, although it's, you know, on their diagram, it's cheaper and a better word error rate than all those companies.
Then they have their VoxTroll Mini, which has a 3 billion parameter.
And this is the one that I've been talking about that like maybe Apple would be interested in, but this is for local and edge deployments.
Um, and then of course they have an ultra tree, super stripped down, very fast API version, a 3 billion model called Vox pro mini transcribe.
So this is really optimized only for transcriptions.
Um, but it says that it can outperform open AIs whisper and it's less than half the price.
Uh, so this is definitely something interesting and people right now can go and try this over on, uh, you can go for free, download the API on hugging face, or you can, uh, have the testing model is in there on their website.
Mistral's chatbot LeChat has it there.
So very, very interesting.
This is obviously one of the big AI firms out of Europe.
They have this big, you know, quote unquote, $1 billion in equity investment looking like it's going to happen from Abu Dhabi's MGX fund happening soon.
So this is kind of the perfect time for them to start rolling out these tools and perhaps getting some of their competitors or perhaps business partners interested in acquisitions or looking at making deals with them.
So if they got the Apple deal, that would be absolutely incredible whether you know, even if that's not an acquisition by Apple, which sounds like they don't really want to go in that direction.
But making some sort of partnership.
We know Apple right now is talking to more than just opening eye who's powering Apple intelligence.
Now they're looking at Anthropix cloud, it seems like Apple feels quite behind all of their investors, their board is quite upset about Apple with the their slowness to adopt these these AI features and what they've said is kind of like a failure in that department.
So acquiring a company like this might be good, but if not, they, they may be interested in working with Mistral.
So it'll be very interesting.
I'll keep you up to date on everything happening with this startup and with others, uh, as people get their hands on this new tool for Mistral and start incorporating it into products.
I think it's going to be interesting and we'll make sure to get it up on AI box in not too distant of the future.
So thanks so much for tuning in to the podcast. Uh, make sure to leave a rating and review.
If you enjoyed the episode, if you learned anything new about what's going on over at Mistral.
Thanks so much for tuning in and I will catch you in the next episode.