Realtime voice agent

실시간 음성 에이전트

Stream speech input and output so an agent can respond and handle interruptions.

···
html
<div class="scene"><div class="heading">LIVE SPEECH · TURN TAKING</div><div class="voice-row"><b>사용자</b><div id="user-wave" class="wave"></div></div><div class="voice-row"><b>에이전트</b><div id="agent-wave" class="wave"></div></div><div class="voice-status" id="voice-status">사용자 발화 수신</div></div>
css
.scene{width:min(94vw,800px);height:min(88vh,326px);padding:clamp(10px,2.7vmin,19px);border:1px solid var(--line);border-radius:13px;background:var(--surface);font:500 clamp(13px,2.4vmin,17px)/1.3 var(--font-sans,sans-serif);position:relative;overflow:hidden}.scene .mono{font-family:ui-monospace,SFMono-Regular,monospace}.scene .muted{color:var(--muted)}.scene .accent{color:var(--accent)}.scene .heading{font-weight:750;color:var(--accent);margin-bottom:clamp(5px,1.6vmin,12px)}.voice-row{height:31%;display:flex;align-items:center;gap:10px}.voice-row b{width:22%}.wave{flex:1;height:70%;display:flex;align-items:center;gap:3px}.wave i{width:5%;height:12%;background:var(--line);border-radius:4px;transition:height .2s,background .2s}.wave.active i{background:var(--accent)}.voice-status{border-left:3px solid var(--accent);padding:7px;background:var(--bg);margin-top:3px}
js
const pattern=[18,45,75,36,58,88,42,67,30,54,79,24];for(const id of ['user-wave','agent-wave'])document.getElementById(id).innerHTML=pattern.map(()=>'<i></i>').join('');const notes=['사용자 발화 수신','에이전트 음성 응답','사용자 끼어들기 · 응답 중단'];let step=0;function draw(){for(const [id,active] of [['user-wave',step!==1],['agent-wave',step===1]]){const wave=document.getElementById(id);wave.classList.toggle('active',active);wave.querySelectorAll('i').forEach((bar,i)=>bar.style.height=(active?pattern[(i+step*3)%pattern.length]:12)+'%')}document.getElementById('voice-status').textContent=notes[step];step=(step+1)%3}draw();setInterval(draw,1050)

A realtime voice agent exchanges audio chunks during a conversation. The Realtime API became generally available in 2025 with direct speech processing and tool calling to support low-latency interactions.

The demo alternates user and agent waveforms, then stops the agent output when the user interrupts. Real applications must test network latency, recognition errors, and interruption timing.

When to use

Use it for spoken support or tutoring where quick turn changes and interruptions matter.

Open as page ↗