Sunday, 4 Oct 2026 Ashwin Krishna Navami · VS 2083
Close 1 Oct: Sensex71,909.70▼0.79%Nifty 5022,421.95▼0.88%Bank Nifty54,450.75▼0.33% USD₹96.34 BTC₹82,11,616▲0.79%
--° Delhi
नभ 24 — Nabh24
BREAKING
ग्रामीण विकास सेवा संस्थान, पारू ने चलाया राहत अभियान, बाढ़ प्रभावितों के बीच बांटी राहत सामग्री; 500 से अधिक लोगों को कराया भोजनIndia में Tech Revolution की नई शुरुआत, AI और Semiconductor पर बढ़ा फोकस; ₹1.27 लाख करोड़ का Semicon 2.0CBSE Class 10 Exam 2027: Passing Marks को लेकर बड़ा अपडेट, 33% Overall Score से होंगे PassMiddle East tensions spike as Gaza strike, Sanaa raids and tanker attack unfoldEmerging Powers Reshape Global Order Amid US Repositioning, China's RiseSingeetham Srinivasa Rao’s ‘Michael Madana Kama Rajan’, a timeless classic from a master filmmakerKeralam government to begin talks on starting an academy for calligraphyWhy scientists are rethinking the chemical ‘arms race’ against fungiMukesh, Smaran help Rest of India outclass J&K by 167 runs to lift Irani CupLula vs Bolsonaro: Will Brazil join 7 Latin American nations in rightward shift? Impact on ties with US, China explained2 Jharkhand teens drown in separate incidents during Jitiya festivalAAP MLA cries foul over teacher shortage in Punjab; is suspended by partyTalks between Pakistan's Shehbaz govt, Imran Khan's PTI fail; Islamabad braces for more protestsBJP questions Congress's 'silence' on Captain Smit Machchhar, UP chief pickShould Claude get the same treatment as humans? The question opens rift between Pope and AnthropicThe Swiss watch industry is shrinking—but getting swankierKeralam’s former Minister Muneer seeks comprehensive probe into Malaysian engineer’s death in 2006
Technology

Microsoft ने लॉन्च किए नए AI Voice Models, 60 भाषाओं में Real-Time Transcription से बदलेगा Voice AI का अनुभव

Microsoft ने तीन नए AI voice models लॉन्च किए: real-time transcription के लिए MAI-Transcribe-2-Streaming, और natural आवाज के लिए MAI-Voice-2.1 व Flash।

Microsoft’s new MAI AI voice models focus on real-time transcription and natural multilingual voice generation. | Photo: Microsoft AI

नई दिल्ली, 2 अक्टूबर 2026: Microsoft ने Artificial Intelligence (AI) के क्षेत्र में बड़ा अपडेट करते हुए तीन नए voice models लॉन्च किए हैं। कंपनी ने MAI-Transcribe-2-Streaming, MAI-Voice-2.1 और MAI-Voice-2.1-Flash पेश किए हैं।

इन नए models का उद्देश्य AI को इंसानों के साथ ज्यादा तेज, natural और real-time बातचीत करने में सक्षम बनाना है। Microsoft के मुताबिक, इनका इस्तेमाल customer-service agents, multilingual assistants, interactive learning और voice-based applications में किया जा सकता है।

MAI-Transcribe-2-Streaming क्या है?

Microsoft का नया MAI-Transcribe-2-Streaming कंपनी का streaming transcription model है, जो किसी व्यक्ति के बोलते समय ही speech को text में बदलना शुरू कर देता है।

AdvertisementTech RajeshwarNeed a new website?Websites & apps built for your business — contact Tech RajeshwarContact us

आमतौर पर speech-to-text systems में user के बोलना पूरा करने के बाद transcription तैयार होती है। नए streaming model में शुरुआती transcription results बातचीत के दौरान ही मिलने लगते हैं।

Microsoft के अनुसार, model audio प्राप्त होने के 100 milliseconds से कुछ अधिक समय में शुरुआती transcription hypotheses देना शुरू कर सकता है।

“It produces its first hypotheses in just over 100ms of receiving audio.”

— Microsoft AI

60 भाषाओं में Real-Time Transcription

MAI-Transcribe-2-Streaming को 60 भाषाओं में real-time transcription के लिए तैयार किया गया है। इसमें automatic और continuous language detection की सुविधा भी दी गई है।

इसका मतलब है कि voice-based applications बातचीत के दौरान भाषा को पहचानकर speech को text में बदल सकती हैं।

AdvertisementTech RajeshwarNeed a new website?Websites & apps built for your business — contact Tech RajeshwarContact us

Microsoft का कहना है कि streaming transcription की मदद से voice agents user की बात पूरी होने से पहले ही processing या दूसरे actions की तैयारी शुरू कर सकते हैं।

AI Voice Agents के लिए क्यों महत्वपूर्ण है?

Voice-based AI में एक सामान्य बातचीत के दौरान system को कई काम करने पड़ते हैं—पहले आवाज सुनना, फिर उसे समझना, जवाब तैयार करना और अंत में आवाज में जवाब देना।

अगर इनमें से किसी चरण में ज्यादा delay हो तो बातचीत robotic महसूस हो सकती है।

Microsoft के नए models इसी latency को कम करने पर फोकस करते हैं।

  • Speech को real time में text में बदलना
  • User की बात के दौरान processing शुरू करना
  • Natural-sounding voice में जवाब देना
  • कई भाषाओं में बातचीत करना
  • Customer-service applications में तेजी से response देना
  • Live captions और voice interfaces को बेहतर बनाना

MAI-Voice-2.1 में 23 भाषाओं का support

Microsoft ने MAI-Voice-2.1 नाम का नया text-to-speech model भी पेश किया है।

कंपनी के अनुसार, यह model 23 भाषाओं और 26 locales को support करता है। इसमें एक ही voice identity को अलग-अलग भाषाओं में बनाए रखने की सुविधा दी गई है।

Supported languages में Hindi और English (India) भी शामिल हैं।

इससे multilingual AI assistants किसी user की भाषा में जवाब देने के साथ एक consistent voice बनाए रख सकते हैं।

MAI-Voice-2.1-Flash ज्यादा speed के लिए

Microsoft ने MAI-Voice-2.1-Flash को high-volume और latency-sensitive applications के लिए तैयार किया है।

कंपनी के अनुसार, Flash version लगभग 55% faster model inference देता है और comparable models की तुलना में लगभग 60% कम cost के लिए designed है।

Microsoft ने इसकी कीमत $15 प्रति 1 million characters बताई है, जबकि MAI-Voice-2.1 की कीमत $22 प्रति 1 million characters है।

Voice Cloning में भी नई सुविधा

दोनों MAI voice models short reference audio की मदद से voice matching और cloning capabilities को support करते हैं।

Microsoft के अनुसार, इसके साथ consent guardrails भी मौजूद हैं, जिनका उद्देश्य voice technology के गलत इस्तेमाल को रोकने में मदद करना है।

Voice cloning के बढ़ते इस्तेमाल के बीच consent और authorized voice use जैसे मुद्दे AI industry में महत्वपूर्ण बने हुए हैं।

किन क्षेत्रों में होगा इस्तेमाल?

Microsoft के अनुसार, developers इन models का इस्तेमाल कई तरह के applications बनाने के लिए कर सकते हैं।

  • Customer service: AI customer-service agents बातचीत के दौरान request को समझकर response तैयार कर सकते हैं।
  • Multilingual assistants: अलग-अलग भाषाओं में बातचीत करने वाले AI assistants बनाए जा सकते हैं।
  • Education: Interactive learning और tutoring applications में natural voice का इस्तेमाल किया जा सकता है।
  • Live transcription: बातचीत के दौरान speech को text में बदला जा सकता है।
  • Voice interfaces: ऐसे applications बनाए जा सकते हैं जिनमें voice मुख्य interaction method हो।

Developers के लिए उपलब्ध हैं नए Models

Microsoft के अनुसार, तीनों models Microsoft Foundry और MAI Playground के जरिए उपलब्ध हैं। कंपनी ने Vercel को भी supported platforms में शामिल किया है।

MAI-Voice-2.1 और MAI-Voice-2.1-Flash को OpenRouter के जरिए भी access किया जा सकता है, जबकि LiveKit support को आने वाले समय के लिए बताया गया है।

Microsoft का बढ़ता AI Model Ecosystem

Microsoft पिछले कुछ महीनों में अपनी MAI model family का विस्तार कर रहा है। कंपनी ने voice, transcription, image, coding और reasoning जैसे अलग-अलग क्षेत्रों में अपने AI models पेश किए हैं।

सितंबर 2026 में कंपनी ने MAI-Transcribe-2 भी पेश किया था, जबकि अब नया streaming version real-time voice interaction पर फोकस करता है।

इससे Microsoft का AI development अब केवल text-based applications तक सीमित नहीं रह गया है, बल्कि real-time voice interaction की दिशा में भी आगे बढ़ रहा है।

AI बातचीत को ज्यादा Natural बनाने की कोशिश

नए models के जरिए Microsoft voice AI के दो महत्वपूर्ण हिस्सों पर एक साथ काम कर रहा है।

एक तरफ MAI-Transcribe-2-Streaming user की आवाज को तेजी से समझने की कोशिश करता है, वहीं दूसरी तरफ MAI-Voice-2.1 और MAI-Voice-2.1-Flash natural voice में जवाब देने के लिए बनाए गए हैं।

इस combination का इस्तेमाल ऐसे AI agents बनाने में किया जा सकता है जो सुनने, समझने और बोलने के बीच कम से कम delay रखें।

“A voice agent is a loop. It has to hear, understand, decide, and speak.”

— Microsoft AI

मुख्य बातें

  • 3 नए AI models: Microsoft ने MAI-Transcribe-2-Streaming, MAI-Voice-2.1 और MAI-Voice-2.1-Flash लॉन्च किए हैं।
  • 60 भाषाएं: Streaming transcription model 60 भाषाओं में real-time transcription support करता है।
  • 23 भाषाएं: MAI-Voice-2.1 और Flash 23 भाषाओं और 26 locales को support करते हैं।
  • Hindi support: Voice models की supported languages में Hindi और English (India) शामिल हैं।
  • Low latency: Streaming model audio मिलने के 100 milliseconds से कुछ अधिक समय में शुरुआती transcription results देना शुरू कर सकता है।
  • Voice cloning: Voice models short reference audio के आधार पर voice matching capabilities support करते हैं।
  • Developer access: Models Microsoft Foundry और MAI Playground सहित कई platforms पर उपलब्ध हैं।

AI Voice Technology का अगला चरण

Real-time transcription और natural text-to-speech को एक साथ जोड़ने से voice-based AI applications के लिए नई संभावनाएं खुल सकती हैं।

Customer service, education, accessibility और multilingual communication जैसे क्षेत्रों में ऐसे AI agents विकसित किए जा सकते हैं जो बातचीत को ज्यादा natural तरीके से संभालें।

Microsoft के नए MAI models इस दिशा में कंपनी के latest AI developments में शामिल हैं और आने वाले समय में voice-based AI applications के विकास को प्रभावित कर सकते हैं।

निष्कर्ष

Microsoft के नए AI voice models real-time speech recognition और natural voice generation को एक साथ आगे बढ़ाने की कोशिश हैं।

MAI-Transcribe-2-Streaming का फोकस user की speech को बातचीत के दौरान समझने पर है, जबकि MAI-Voice-2.1 और MAI-Voice-2.1-Flash natural और तेज voice responses देने के लिए बनाए गए हैं।

इन technologies का इस्तेमाल AI assistants, customer service, education, multilingual communication और voice-driven applications में किया जा सकता है।

Frequently asked questions

Microsoft ने कौन से नए AI models लॉन्च किए हैं?

Microsoft ने MAI-Transcribe-2-Streaming, MAI-Voice-2.1 और MAI-Voice-2.1-Flash लॉन्च किए हैं।

MAI-Transcribe-2-Streaming क्या करता है?

यह बातचीत के दौरान speech को real time में text में बदलने के लिए बनाया गया streaming transcription model है।

MAI-Transcribe-2-Streaming कितनी भाषाओं को support करता है?

Microsoft के अनुसार, MAI-Transcribe-2-Streaming 60 भाषाओं में real-time transcription support करता है और इसमें automatic language detection भी है।

क्या नए Microsoft AI voice models में Hindi support है?

हां, MAI-Voice-2.1 और MAI-Voice-2.1-Flash की supported languages में Hindi और English (India) शामिल हैं।

MAI-Voice-2.1-Flash किसके लिए बनाया गया है?

इसे high-volume और latency-sensitive voice applications के लिए optimize किया गया है।

AdvertisementCallZenixAI Voice Bot, Dialer & IVR SolutionsCloud dialer and contact-center software for your teamKnow more

Spotted a mistake? See our corrections policy or write to grievance@mishraworld.com.

AdvertisementMishra World270+ free online tools — GST, tax, EMI & moreCalculators, invoices, biodata and document tools. No sign-up, works on mobile.Use free tools
MISHRA NEWS
Sections
More