Project 03 Voice · English & Spanish · Sources shown

Live prototype

Ask by voice. Check the source. Then hear the verified answer.

Record up to 15 seconds or start with a typed example. LinguaVoice shows the transcript and detected language, checks approved sources, and keeps the answer visible before you choose to create its spoken version. If the sources do not support an answer, it says so.

Try the guided example
  • 15 sechard recording cap
  • Ephemeralno first-party storage
  • Citedor it abstains
INPUT / LIVE AUDIO OUTPUT / GROUNDED VOICE

ASK

REPLY

EN English ES Español FR Français

One question. One evidence trail. The same language on the way back.

01 / Live system

A guided voice loop

One question from input to spoken answer.

New here? Run the example without a microphone. The page will show what happened, what to do next, and which source supports the answer.

Start with the guided example, record a short question, or type your own.
  1. 1 Ask or use the example In progress
  2. 2 Review words & language Pending
  3. 3 Check answer & source Pending
  4. 4 Choose Play Pending
First time here?

Run one safe English question. No microphone permission is needed.

Voice input 00:00 / 00:15

Explicit one-time processing. By pressing record, you agree to send up to 15 seconds of audio to OpenAI through the project API for transcription. This demo does not intentionally persist the audio; do not include confidential or regulated information.

or type instead

Up to 500 characters. Press Control/Command + Enter to send. OpenAI processes the question through the project API to select approved sources; do not include private or sensitive information.

0 / 500

Or load a question to review first

Verified output Visible before spoken

Latest question

Voice transcription or typed input will appear here.
Language pending

The answer will appear here with its evidence status. Speech starts only after the visible response has been checked by the client-side guardrail.

Step 4 · Spoken version

Press Play to send the exact short answer above to OpenAI and create an AI-generated voice. Audio is requested only on your click and is not stored by this demo.

02 / Problem

Why this interface exists

Voice is convenient. Untraceable answers are not.

A conventional voice bot can sound certain before a user has any way to inspect the answer. That is the wrong order for service information, operating procedures or internal knowledge.

LinguaVoice makes the transcript, answer and citations visible first. Voice is the final delivery layer—not a substitute for evidence.

03 / Architecture

A short, inspectable loop

Four stages. Two hard exits.

Recording time and evidence quality can both stop the flow before an unsafe response reaches speech.

  1. 01

    Capture

    Permission begins on an explicit click. Browser recording is cut off at 15 seconds and tracks close immediately.

    Binary audio · transient
  2. 02

    Transcribe

    The API returns the transcript plus language metadata. You review or correct both before retrieval begins.

    Text + language metadata
  3. 03

    Retrieve

    The question is answered against the project knowledge base. Support status, citations and abstention are returned together.

    Answer · support · citations
  4. 04

    Speak

    Only an explicit Play action can exchange a short-lived ticket for audio of the exact verified text already on screen.

    On-demand MP3 · one-use ticket

04 / Guardrails

Designed for bounded answers

The refusal is a feature.

Every layer has a narrow job, a visible state and a recoverable failure path.

G1

Evidence gate

An answer is shown as grounded only when the API marks it supported, returns citations and does not flag abstention.

G2

Deterministic refusal

If any evidence signal is missing, the client replaces the returned copy with a fixed “not enough evidence” response.

G3

Safe rendering

API text is inserted as text, never executable markup. Citation links accept only HTTP or HTTPS destinations.

G4

Ephemeral audio

No browser storage is used. Recording chunks are cleared after transcription; generated playback stays only in the open page and its object URL is revoked when replaced or closed.

G5

Human-readable state

Permission, recording, transcription, retrieval and speech are announced in a live status region without hiding the controls.

G6

Text always wins

Denied permission, missing microphone support or speech failure never blocks the typed question and visible answer path.

05 / Test matrix

Different users, same safe route

Acceptance scenarios built into the interface.

These scenarios describe the expected observable behavior, not hidden benchmark claims.

Profile / conditionActionExpected result
First-time visitorPresses record without prior permission.The browser prompt appears after the click; the interface explains the 15-second limit and storage policy beforehand.
Microphone deniedBlocks or dismisses permission.A recoverable message points to the text field; no page reload is required.
Unsupported questionAsks for information absent from the sources.The client displays a fixed abstention instead of unsupported generated copy. It is spoken only if the user then presses Play.
Multilingual userSpeaks a language returned by transcription.The transcript and language remain editable before retrieval. English and Spanish use approved sources; French, German, Italian and Portuguese return fixed safe boundaries, while an unrecognized language requires manual selection.
Keyboard-only userTabs through record, text, language and playback controls.Native controls retain visible focus, descriptive labels and predictable activation.
Reduced motionEnables the operating-system motion preference.The hero resolves to a static composition; the functional demo remains unchanged.
Slow or offline networkA request fails or reaches its timeout.The last text question can be retried from memory, while voice can be recorded again without retaining old audio.
Hostile API textA response contains markup or a non-web citation scheme.Content renders as inert text and unsafe link schemes are not made clickable.

06 / Outcome

Fast enough for conversation.
Strict enough for work.

LinguaVoice is a reusable pattern for bilingual front desks, internal procedure assistants and field-service support: voice at the edges, evidence at the center.

Discuss a grounded voice assistant