06 Case Studies 3 min read 625 words
Case Study: Real-Time AI Search Engine
The Problem
A fintech startup needs to build a real-time market intelligence platform that lets analysts ask natural language questions about live market data, news, and company filings.
Constraints given in the interview:
- Data freshness: queries must reflect information from the last 5 minutes
- Scale: 10,000 concurrent users, 50,000 queries/hour
- Accuracy: financial data cannot be hallucinated
- Latency: p95 response time under 3 seconds
The Interview Question
Solution Architecture
One stream writes both a vector index and a full-text index, and the router reads both so reciprocal-rank fusion can reconcile semantic and keyword hits before reranking
Text version of this diagram
flowchart TB
subgraph Ingestion["Real-Time Ingestion Layer"]
NEWS[News Feeds] --> KAFKA[Kafka Stream]
FILINGS[SEC Filings] --> KAFKA
SOCIAL[X/Reddit APIs] --> KAFKA
KAFKA --> PROCESSOR[Stream Processor]
end
subgraph Index["Dual-Index Layer"]
PROCESSOR --> VECTOR_DB[(Vector DB<br/>Qdrant)]
PROCESSOR --> SEARCH_IDX[(Full-Text<br/>Elasticsearch)]
end
subgraph Query["Query Layer"]
USER[User Query] --> ROUTER{Query Router}
ROUTER -->|Semantic| VECTOR_DB
ROUTER -->|Keyword| SEARCH_IDX
VECTOR_DB --> FUSION[RRF Fusion]
SEARCH_IDX --> FUSION
end
subgraph Generation["Answer Generation"]
FUSION --> RERANK[Cross-Encoder Rerank]
RERANK --> LLM[GPT-4o-mini]
LLM --> RESPONSE[Sourced Answer]
end
Key Design Decisions
1. Why Kafka for Ingestion?
The interviewer wants to know you understand streaming vs batch.
Answer: Kafka provides exactly-once delivery and allows multiple consumers. We have one consumer writing to the vector DB and another to Elasticsearch. If the vector indexing falls behind, the full-text index still serves queries. This is the dual-write pattern for resilience.
2. Why Hybrid Search (Vector + Full-Text)?
Answer: Financial queries mix semantic ("sentiment around Tesla") with keyword ("TSLA 10-K filing"). Pure vector search would miss exact ticker matches. We use Reciprocal Rank Fusion (RRF) to combine results.
3. Why GPT-4o-mini Instead of GPT-4o?
Answer: For a 3-second p95 latency target at 50K queries/hour, we need fast generation. GPT-4o-mini gives us 100+ tokens/second vs 40 tokens/second for GPT-4o. The reranker handles accuracy; the LLM only synthesizes already-verified content.
Handling the Freshness Requirement
The hardest part of this problem is ensuring the index reflects data from the last 5 minutes.
Solution: TTL-Based Indexing
# Each document gets a timestamp field
doc = {
"content": "Tesla announces new factory...",
"timestamp": datetime.now(UTC),
"source": "Reuters",
"ttl_hours": 24 # Auto-delete after 24 hours
}
# Query filters to last N minutes
def search_recent(query: str, minutes: int = 60):
cutoff = datetime.now(UTC) - timedelta(minutes=minutes)
return vector_db.search(
query=query,
filter={"timestamp": {"$gte": cutoff}}
)
Cost Analysis
| Component | Monthly Cost (at 50K queries/hour) |
|---|---|
| Kafka (MSK) | $2,500 |
| Qdrant (managed) | $1,800 |
| Elasticsearch | $2,000 |
| GPT-4o-mini (generation) | $3,500 |
| Cross-encoder reranking | $800 |
| Total | $10,600/month |
Interview Follow-Up Questions
Q: How do you prevent hallucinated financial data?
A: Three layers: (1) The LLM only summarizes retrieved content, never generates facts. (2) Every claim must cite a source document. (3) A post-generation validator checks that any number in the response exists verbatim in a source.
Q: What if Kafka falls behind during a news spike?
A: We implement backpressure with consumer lag monitoring. If lag exceeds 2 minutes, we shed load on the ingestion side using sampling. Real-time queries hit a "recent" index with only the last hour of data; batch jobs backfill the full index.
Key Takeaways for Interviews
- Real-time AI search requires streaming infrastructure, not batch ETL
- Hybrid search (semantic + keyword) outperforms pure vector for structured domains
- Latency budgets drive model selection: use fast models for synthesis, save expensive models for reasoning
- Freshness is a filter, not a feature: implement at the index level, not the prompt level
Related chapters: Hybrid Search, Serving Infrastructure