· XingAI Invest AI
Treating LLM Output Like Cache Rows (Planned): Worker-Owned Inference
UI는 한국어입니다. 글 본문은 아직 영어 또는 중국어만 있습니다.
Status: planned (not shipped in V1)
This post documents a decision we accepted, not something already running in production. It belongs in the same story as our CQRS market cache: derived data should have one writer.
Today, the FastAPI backend still calls OpenAI / Gemini / Ollama inside the request lifecycle (see OpenAI-first routing). That works for V1, but it creates five pressures as traffic grows:
- Duplicate spend — ten users asking about the same symbol in one minute can mean ten identical model calls.
- Coupled failures — upstream 429s surface as user-visible errors instead of stale-but-usable cache.
- Secrets on the hot path — API keys live on the public-facing service.
- Heavier cold starts — more imports on the API process.
- Half-applied CQRS — market data already follows “worker writes, backend reads”; LLM results do not yet.
The direction
Move all LLM calls into the worker. Persist analyses in SQLite like any other derived artifact. The backend becomes a reader: cache hit returns immediately; cache miss enqueues a job and returns a small “queued” envelope (with polling URL). Free-tier quotas stay enforced at the gateway — abuse should not enqueue a thousand jobs.
We plan two phases:
- Phase A — Non-streaming analyze path moves to enqueue + poll; streaming can follow a dedicated plan.
- Phase B — After the V2 hybrid pipeline (Gemini compress → OpenAI decide), streaming and multi-stage work naturally live in the worker anyway.
SQLite WAL mode + a bounded job table with dedupe on prompt_hash keeps concurrent reads cheap while the worker writes results.
Why bother?
- Cost — dedupe and pre-warm hot symbols (watchlist, top signals) collapse redundant calls.
- Resilience — serve last good analysis with a
staleflag when upstream is unhappy. - Security posture — API keys migrate to the worker process only.
Takeaway
“LLM as a service call inside the controller” is the fastest way to V1. “LLM as cached rows with a single writer” is the shape that survives real traffic and aligns with how we already treat market data.
Further reading: ADR-010 (docs/adr/010-llm-as-cached-resource.md).