Voice AI: when the machine speaks and listens
Voice has become a natural AI interface: transcribing, synthesizing, conversing aloud. This dossier tracks voice AI — from speech-to-text to synthetic voices — and the authenticity questions it raises.
Latest voice AI news
- AI audio translator with speech-to-text, LLM translation, and text-to-speech — Hacker News
- Ask HN: What are you go to LLM models for the following — Hacker News
- Claude Code 2.1.69 — Improved MCP binary content handling: tools returning PDFs, Office documents, or audio now save decoded bytes to disk with the co — Claude Code
- Claude Code 2.1.74 — Fixed voice mode silently failing on the macOS native binary for users whose terminal had never been granted microphone permissio — Claude Code
- Show HN: Fleet – drive a fleet of Claude Code/Codex agents from Telegram — Hacker News
- Show HN: PrintBlocks – API and MCP server for your thermal printer — Hacker News
- Speech to text in a crowded room with the OpenAI Realtime API — Hacker News
- Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles — the-decoder.com
- Google Gemini 3.5 AI Launches Speech-to-Text Model — The Cryptonomist
- ChatGPT Voice shifts from chatbot to proper assistant — New Atlas
Transcription and speech understanding
Speech-to-text turns audio into text with accuracy that has leapt forward: subtitles, meeting notes, dictation. Paired with a language model, it can summarize or query a recording, opening uses far beyond plain transcription.
Speech synthesis and natural voices
Text-to-speech produces increasingly natural voices, reaching expressiveness and emotion. It powers screen readers, dubbing, narration and assistants. The line with a human voice is at times hard to perceive.
Voice assistants, cloning and risks
Voice assistants are becoming conversational and responsive. But voice cloning raises real risks: impersonation scams, disinformation, identity theft. Content labeling and vigilance are needed against credible synthetic voices.
Frequently asked questions
What is voice AI?
The set of AI technologies applied to voice: transcription (speech-to-text), synthesis (text-to-speech), voice assistants and speech understanding.
Can Claude speak or transcribe?
Claude specializes in text and code; it can process a provided transcript, but transcription and speech synthesis are handled by dedicated audio tools.
Is voice cloning dangerous?
It can be: impersonation for scams, disinformation, identity theft. Hence the importance of content labeling and caution.
Claude News is published by Héra SASU. Independent media, not affiliated with Anthropic.