AI Platform
The AI platform (phases 0–7) has one governed call path, modality-aware telemetry, configurable models/budgets/kill-switches, and multimodal ingestion (voice, documents, creative) with semantic search.
Call path
Section titled “Call path”UI / routes ─► lib/ai/run path ├─ governed text generation ─► lib/ai/genai-invoke.ts (AI SDK + provider registry) └─ agentic copilot ─► lib/ai/copilot/* (streamText + tools) both ─► lib/ai/observe.ts ─► ai_invocations (telemetry + cost) lib/ai/budget.ts ─► daily token/USD enforcementlib/ai/genai-invoke.tskeeps the backwards-compatible raw Gemini model/tool input shape, resolves models through the AI SDK provider registry, records normalized usage/errors, and enforces budget.- Copilot routes (
/api/ai/copilot,/api/ai/whatsapp-copilot) record usage inonFinishand check budget before starting. - All
/api/ai/*routes and the central libs (insights, meeting extraction, AI-SQL, lead agent) go through the instrumented path. No hardcoded model strings remain.
Model routing & config
Section titled “Model routing & config”- Task classes:
copilot,extract,transcribe,vision(lib/ai/copilot/provider.ts), with defaults andCOPILOT_MODEL/COPILOT_MODEL_CHEAPenv overrides. - Workspace config in
workspace_settings.ai_config(models,budgets,features,caching), resolved env > DB > default inlib/ai/config.ts. - Settings → AI (
/settings/ai) shows 30-day telemetry and lets a System Manager edit models, budgets, and feature kill-switches (ai.saveConfig).
Observability
Section titled “Observability”ai_invocations records one row per model call: feature, surface, provider, model alias, modality, prompt version, status, latency, tokens (input/output/cached/reasoning), media units (images/pages/audio-seconds/video-seconds/media-bytes/sha256), cache hit, and estimated USD (lib/ai/observe.ts). Numeric fields are sanitized so a provider NaN never drops a row. Read it via ai.invocationsSummary / ai.invocationsList.
Multimodal ingestion
Section titled “Multimodal ingestion”| Modality | Library | Cron | Flag |
|---|---|---|---|
| Voice notes → transcript | lib/ai/transcribe.ts |
/api/cron/media-transcribe |
voiceTranscribe |
| Documents → structured fields | lib/ai/document-extract.ts |
/api/cron/media-extract-documents |
documentExtract |
| Creative → caption/tags/alt-text/brand-safety | lib/ai/creative.ts |
on demand | creativeTag |
| Messages/meetings → embeddings | lib/ai/index-content.ts + lib/ai/embed.ts |
/api/cron/ai-index |
retrieval |
Media is deduplicated by sha256 (media_assets), so identical bytes are never transcribed/extracted/embedded twice.
Caching
Section titled “Caching”- Response cache (
lib/ai/cache.ts, Redis): conversation insights are cached for 30 minutes keyed by content hash; hits are recorded ascache_hit=truetelemetry. - Gemini implicit prompt caching applies automatically for 2.5+ models on stable prefixes (min ~2–4k tokens).
- Media dedupe via
media_assets.sha256.
Semantic search
Section titled “Semantic search”ai_embeddings stores 3,072-dim Gemini embeddings with an owner. searchSimilar runs exact cosine search (HNSW is deferred until the corpus justifies a halfvec index). The copilot exposes search.semantic, scoped to the caller’s own indexed content.
lib/ai/evals/scorers.ts:exactMatch,contains,levenshteinSimilarity,wordErrorRate,toolCallAccuracy,jsonFieldMatch.lib/ai/evals/dataset.ts: mines the meetingsextractionvs reviewer-appliedreviewed_extractiondelta into labeled eval cases.- Run:
node scripts/with-env.mjs production -- pnpm --dir apps/web ai:eval(calls the model; exits non-zero underAI_EVAL_MIN_SCORE).
Operations
Section titled “Operations”- Register cron schedules:
npx tsx apps/web/lib/cron/setup-schedules.ts(needsQSTASH_TOKEN). Schedules:media-transcribe(15m),media-extract-documents(30m),ai-index(20m),agent-triggers(5m). - Enable feature flags in Settings → AI (
documentExtract,retrievalship off). AI_QUERY_DATABASE_URLshould point at the read-only Postgres role for AI-SQL hardening.
