13 Case Studies 4 min read 856 words
Case Study: Voice AI Assistant for Healthcare
The Problem
A hospital network wants a voice-based AI assistant that helps nurses document patient encounters. The nurse speaks naturally; the AI produces structured clinical notes in real-time.
Constraints given in the interview:
- HIPAA compliance (PHI handling)
- Works in noisy hospital environments
- Real-time transcription (under 500ms latency)
- Must use medical terminology correctly
- Integration with existing EHR (Epic/Cerner)
The Interview Question
Solution Architecture
Speech becomes a structured clinical note in four stages, and nothing reaches the EHR until the nurse approves or edits it
Text version of this diagram
flowchart TB
subgraph Capture["Audio Capture"]
MIC[Nurse's Device] --> VAD[Voice Activity Detection]
VAD --> STREAM[Audio Stream]
end
subgraph Transcription["Real-Time Transcription"]
STREAM --> ASR[Whisper Large v3<br/>On-Prem]
ASR --> RAW[Raw Transcript]
end
subgraph Processing["Clinical Processing"]
RAW --> DIARIZE[Speaker Diarization<br/>Nurse vs Patient]
DIARIZE --> NER[Medical NER<br/>Symptoms, Meds, Vitals]
NER --> STRUCTURE[Note Structurer<br/>GPT-4o]
end
subgraph Output["EHR Integration"]
STRUCTURE --> REVIEW[Nurse Review Screen]
REVIEW --> APPROVE{Approved?}
APPROVE -->|Yes| EHR[(Epic/Cerner<br/>via FHIR)]
APPROVE -->|Edit| EDIT[Nurse Edits]
EDIT --> EHR
end
Key Design Decisions
1. On-Premise ASR for HIPAA
Answer: PHI cannot leave the hospital network without encryption and BAA. We deploy Whisper Large v3 on local GPU servers rather than using cloud APIs:
| Option | Latency | HIPAA | Cost |
|---|---|---|---|
| Cloud ASR (OpenAI) | 200ms | Requires BAA, data leaves network | $0.006/min |
| On-prem Whisper | 150ms | Full control, no data egress | $0.002/min (amortized GPU) |
On-prem wins on both latency and compliance.
2. Speaker Diarization: Who Said What
Answer: The note must distinguish "Patient reports headache" from "Nurse observes patient grimacing." We use:
# Pyannote for speaker diarization
diarization = pipeline("audio.wav")
# Output: [(0.0, 1.5, "SPEAKER_0"), (1.5, 4.2, "SPEAKER_1"), ...]
# Map speakers based on voice profile
roles = identify_roles(diarization, known_nurse_voiceprint)
# Output: {"SPEAKER_0": "nurse", "SPEAKER_1": "patient"}
The nurse's device captures their voiceprint at setup for role identification.
3. Medical NER for Structured Extraction
Answer: We need structured data, not just prose. Medical NER extracts:
Medical NER turns one spoken sentence into typed clinical fields — including the field it is honest about leaving empty
Text version of this diagram
flowchart LR
TRANSCRIPT["Patient says she has had<br/>a headache for 3 days,<br/>took Tylenol 500mg twice"]
TRANSCRIPT --> NER[Medical NER]
NER --> SYMPTOMS[Symptoms:<br/>headache, 3 days duration]
NER --> MEDS[Medications:<br/>Tylenol 500mg, BID]
NER --> VITALS[Vitals: None mentioned]
We use a fine-tuned BioBERT model for NER, not the LLM, because NER needs to be fast and deterministic.
Handling Noisy Environments
Hospitals are loud. We use multiple strategies:
- Directional microphones on nurse devices focus on nearby speech
- Noise-robust ASR models (Whisper was trained on noisy data)
- Confidence thresholds: if ASR confidence is <0.7, we flag for nurse review rather than guessing
- Keyword spotting: medical terms have custom pronunciation models
The Structured Note Format
The LLM produces SOAP-format notes:
note_prompt = f"""
Generate a clinical SOAP note from this encounter transcript.
Transcript:
{transcript_with_speakers}
Extracted entities:
- Symptoms: {symptoms}
- Medications: {medications}
- Vitals: {vitals}
Output format:
S (Subjective): Patient's reported symptoms
O (Objective): Nurse's observations and measurements
A (Assessment): Clinical impression
P (Plan): Next steps, orders
"""
EHR Integration (FHIR)
The output must be machine-readable for the EHR:
{
"resourceType": "DocumentReference",
"status": "current",
"type": {
"coding": [{"system": "http://loinc.org", "code": "34117-2", "display": "History and physical note"}]
},
"subject": {"reference": "Patient/12345"},
"author": [{"reference": "Practitioner/nurse789"}],
"content": [{
"attachment": {
"contentType": "text/plain",
"data": "base64-encoded-soap-note"
}
}],
"context": {
"encounter": {"reference": "Encounter/visit456"}
}
}
Latency Budget
| Stage | Target | Actual |
|---|---|---|
| Audio capture to VAD | 50ms | 30ms |
| ASR transcription | 200ms | 150ms |
| Diarization | 100ms | 80ms |
| NER extraction | 50ms | 40ms |
| LLM structuring | 500ms | 450ms |
| Total (end-to-end) | 900ms | 750ms |
For real-time feel, we stream partial transcripts while NER and LLM run on completed sentences.
Interview Follow-Up Questions
Q: How do you handle medical abbreviations and jargon?
A: We maintain a custom vocabulary list that maps abbreviations (PRN, BID, SOB) to full terms. This is injected into both the ASR model (for better recognition) and the LLM prompt (for correct expansion in notes).
Q: What if the nurse makes a correction mid-sentence?
A: We detect correction patterns ("actually, I mean...", "no wait, it's...") and use only the corrected version. The LLM is instructed to prefer later statements when conflicts exist.
Q: How do you ensure the AI does not miss critical information?
A: We have a "completeness check" that verifies the note includes all extracted entities. If NER found "chest pain" but the SOAP note does not mention it, we flag for nurse review. We also run a "safety critical" detector that escalates mentions of suicidal ideation, abuse, or other mandatory reporting triggers.
Key Takeaways for Interviews
- On-prem for healthcare: HIPAA often requires local processing
- Diarization is essential: who said what matters clinically
- Hybrid extraction: fast NER for structure, LLM for prose generation
- Always have human review: especially for clinical documentation
Related chapters: Multimodal Models, Reliability Patterns