한국어 箭头
Podcast Cover

[Mistral AI의 새로운 혁신: 오픈 음성 모델 'VoxTroll'과 유럽 AI 시장의 미래]-[Debuting Voxtral]

Hard Fork AI · B2 · 2025-07-20

Technology
또는 웹버전으로 공부하세요

📋 Summary

Mistral AI, 오픈 음성 모델 'VoxTroll' 공개

최근 Mistral AI는 새로운 AI 모델인 **'VoxTroll'**을 발표하며 음성 인식 및 처리 분야에 큰 파장을 일으키고 있습니다. 이 모델은 단순히 텍스트를 변환하는 기능을 넘어, 오디오 파일을 이해하고 자체적인 음성으로 응답할 수 있는 오픈 모델로서, 기존의 11 Labs나 OpenAI의 Whisper와 같은 강력한 경쟁자들과 어깨를 나란히 합니다.

VoxTroll의 주요 기술적 특징과 효율성

Mistral은 세 가지 버전의 VoxTroll 모델을 공개하며 각기 다른 사용 사례를 제시했습니다.

  • VoxTroll Small (24B 파라미터): 프로덕션 규모의 배포를 위해 설계되었습니다. 벤치마크 데이터에 따르면, 이 모델은 11 Labs의 Scribe나 GPT-4o mini, Gemini 2.5 Flash와 비교했을 때 **'더 낮은 비용'**과 **'더 우수한 단어 오류율(Word Error Rate)'**을 보여줍니다.
  • VoxTroll Mini (3B 파라미터): 로컬 및 에지(Edge) 기기용으로 최적화된 모델입니다. OpenAI의 Whisper Large v3보다 뛰어난 오류율을 자랑하며, 인터넷 연결 없이도 기기 내에서 구동이 가능하다는 강력한 장점이 있습니다.
  • VoxTroll Mini Transcribe: 오직 전사(Transcription) 작업에만 최적화된 초경량 모델로, Whisper보다 성능은 높으면서 비용은 절반 이하로 낮추었습니다.

또한, 이 모델들은 다국어 처리가 가능하여 영어, 스페인어, 프랑스어, 포르투갈어, 힌디어, 독일어, 네덜란드어, 이탈리아어 등 다양한 언어를 지원합니다. 특히 LLM 기반의 백본을 통해 최대 30~40분의 오디오를 이해하고 요약하거나, 음성 명령을 실시간 API 호출로 연결하는 기능은 매우 혁신적입니다.

Apple과의 관계 및 인수설에 대한 진실

최근 업계에서는 Apple이 Mistral을 인수할 것이라는 루머가 돌았습니다. 이러한 추측은 VoxTroll과 같은 '에지 모델(Edge Model)'의 특성 때문입니다. 만약 Apple이 이러한 모델을 iPhone에 탑재한다면, 인터넷 연결 없이도 Siri가 사용자의 음성을 완벽하게 이해하고 처리할 수 있게 됩니다.

그러나 Mistral의 CEO는 구체적인 언급은 피하면서도, 회사를 매각할 의사가 없음을 분명히 했습니다. 대신 회사는 기업 공개(IPO)를 목표로 하고 있으며, 유럽의 '크라운 주얼(Crown Jewel)'로서 유럽 내에서 독자적으로 운영되기를 희망하고 있습니다. 다만, Apple이 현재 'Apple Intelligence'를 강화하기 위해 다양한 파트너와 협력을 모색하고 있는 만큼, 완전한 인수보다는 전략적 파트너십의 가능성은 여전히 열려 있습니다.

유럽 AI 시장의 중심, Mistral의 전망

Mistral AI는 유럽에서 가장 많은 자금 조달을 성공시킨 기업이며, 최근 아부다비의 MGX 펀드로부터 10억 달러 규모의 투자를 유치할 것으로 예상됩니다. 유럽의 전폭적인 지원을 받는 Mistral은 이제 단순한 모델 개발사를 넘어, 자체적인 생태계를 구축하고 있습니다.

결론적으로 VoxTroll은 오픈 소스 진영에 강력한 도구를 제공함과 동시에, 기업들이 로컬 서버나 기기에서 고성능 음성 AI를 구동할 수 있는 길을 열어주었습니다. Mistral이 향후 IPO를 통해 유럽 AI의 위상을 어떻게 드높일지, 그리고 주요 테크 기업들과 어떤 협력 관계를 맺어갈지 귀추가 주목됩니다.

🎯Key Sentences

1
All right, let's get into what mistral is doing.
좋아요, 미스트랄이 뭘 하고 있는지 한번 알아봅시다.
2
basically, what that means is like it can take an audio files
기본적으로 그 말은 오디오 파일을 받을 수 있다는 뜻이에요.
3
it's a competitor to 11 labs and a lot of these other players
11랩스나 다른 경쟁 업체들과 비슷한 곳이에요.
4
how often it messes up the words that it's saying.
이 녀석이 말하는 단어를 얼마나 자주 엉망으로 만드는지.
5
I think for a lot of companies, this is, you know, this is quite exciting
많은 회사들에게 있어서, 이건 꽤나 흥미로운 일이라고 생각해요.
모두 펼치기

📝Key Phrases

1
on the heels of
바로 뒤따라
2
rumors swirling
소문이 무성하다
3
break down
분해하다
4
run locally
로컬 환경에서 실행
5
stripped down
최소화된
모두 펼치기

📖 Transcript

Mistral AI has just come out with a brand new AI model, and it is called VoxTroll.
Now there's a bunch of interesting things about this model in particular, one that is an open speech model, the benchmarks and the data essentially of how its word error rate and its price I think are particularly interesting.
I'm going to break all of this down, especially on the heels of a potential $1 billion round of funding that Mistral is looking at doing and all of the rumors swirling about Apple acquiring the company.
What the what the founder has said about that, what the plans are for the future of this company, this is an interesting time for Mistral, no doubt, we're going to get into all of that.
But before we do, I wanted to mention if you want to try out all of the latest models from Mistral, including all the latest models from a lot of different companies, companies, the top 40 AI models, you can go check out my own startup, which is AI box .ai.
Over there we have code stroll, mistral 3b compact, pixtrel, large vision, mistral small, a ton of these interesting, well, all of the mistral models, but a ton of interesting models from Anthropic, Cohere, deep seek, Google, meta, Microsoft, NVIDIA, open AI, Quinn, grok from xai, all of the all of the interesting companies.

ListenLeap이 실제 문맥에서 학습하도록 이끌어줌

🎨 흥미로운 콘텐츠
🌍 실제 자료
📱 언제든 듣고 보기