Lumen - AI Infrastructure Intelligence
Monitor routing quality, spend, drift, governance, rollout safety, and cache efficiency from one production workspace.
Cost reduction
Cache hit rate v2.0
Judge fallback v2.0
Active agents
Routed last 30d
Avg quality
-Enterprise applications routed through Lumen
control scope
-Monthly budget guardrail with hard-stop policy
finance control
-PII redaction and blocked-request enforcement
privacy gate
-Requests inside latency and quality SLA
service posture
Why this model?
Policy, quality, cacheability, latency, and cost are explained for each route.
-Next best model
Model discovery canary recommendations appear before production rollout.
-Policy outcome
Privacy, residency, budget, and allowlist rules are evaluated before inference.
-Agents in production
Routing policy is learned per-agent from past conversations and eval outcomes.
Connect with an API key to load agents
Live routing decision
Most recent /v1/chat/completions decision
Incoming request
POST /v1/chat/completions
Awaiting first request…
Ranked candidates
No decisions yet
Quality score (rolling)
EWMA across all agents
Eval samples
in 30d window
Active drift alerts
—
Models monitored
across active agents
Routed model share
Daily traffic share by chosen model
Drift alerts
CUSUM-detected quality regressions
Connect to load drift alerts
Closed-loop control flow
MeasureProduction prompts & outcomes
→
EvaluateJudge Ops judges
→
DetectCUSUM drift detector
RecalibrateRouting weights updated
→
GovernQuality-bar guard
→
ScorePer-prompt rubric
Sessions
last 30d
Total spend
routed-through
Cache hit rate
prompt-cache discount
Judge fallback
eval pipeline health
Avg latency
across providers
Avg quality
judge-aggregated
Cost savings per session
Daily · vs frontier-only baseline
Routed model trend
Daily share per model
Cache & purpose breakdown v2.0
Cache hit rate by model
Cached input tokens are billed through lumen.pricing with vendor-correct discounts: Anthropic 10%, OpenAI 50%, Gemini 25%
| Model | Calls | Hit rate | Cost |
|---|---|---|---|
| Connect to load cache stats | |||
Volume & quality by purpose
Set
X-Lumen-Purpose to route per workload class| Purpose | Requests | Avg quality | Cost |
|---|---|---|---|
| Connect to load purpose data | |||
Operational self-audit v2.0
Integrity checks
Runs on every page load ·
GET /v1/audit/self-checkConnect to run self-audit
Per-prompt observability
Recent prompts
| Time | Agent | Model | Tokens | Cost | Latency | Quality | Status |
|---|---|---|---|---|---|---|---|
| Connect to load recent prompts | |||||||
Evidence rankings
Ranked by performance on the selected workload. Re-evaluated continuously by Judge Ops.
| Model | Provider | Quality | Cost / 1K | p95 latency | Samples | Status |
|---|---|---|---|---|---|---|
| Connect to load standings | ||||||
Judge Ops pipeline
Async Arq workers score samples and keep standings fresh.
1
Your preferences
—
2
Scoring criteria
5rubrics
3
Automated judge
—
4
Always-on
—
Quality by model
Recent average from automated judges
Scoring criteria
Co-created with your team
Factual accuracy
0.94
Tool-call correctness
0.91
Tone & brand voice
0.88
Refusal calibration
0.86
Citation grounding
0.82
Lumen - AI Infrastructure Intelligence
Shape live routing, provider access, and rollout safety from one governed workspace.
The infrastructure console writes directly to the admin contracts, then refreshes live lists so operators can see the impact immediately.
Policy-first routingAgent and purpose overrides
Secret-safe providersMasked references only
Progressive launchesTraffic-weighted candidates
Deployment pulseReadiness and audit chain
Launch state
-
/v1/admin/deployment/status
Provider Vault
-
secret refs, not plaintext
Policy Forge
-
per agent / purpose
Canary Lab
-
controlled rollout
Budget Burn
-
monthly guardrail
PII Blocks
-
enforced in gateway
Enterprise Control Loop
Detect, decide, enforce, and prove every model-routing action.
1. IngestApp requests arrive with agent, purpose, budget, and data-classification metadata.
2. GovernPolicies check privacy, blocked models, spend caps, region, and SLA constraints.
3. RouteBest-fit model is chosen from standings, quality bars, cacheability, and cost curves.
4. JudgeOutcomes are sampled, scored, drift-checked, and fed back into model standings.
5. ProveDecision, response, eval, budget, and policy evidence are written to the audit chain.
Eval Suite Onboarding
Upload examples, generate rubrics, benchmark models, then publish routing policy.
Loading eval onboarding
Workflow Evaluation
Agentic workflows are scored beyond final answer quality.
| Workflow | Completion | Cost | Status |
|---|---|---|---|
| Loading workflows | |||
Provider Vault
Create or rotate provider key references. Raw secrets are sent once and never rendered.
Policy Forge
Edit live routing guardrails for an agent and purpose pair.
Not simulated
Canary Lab
Send a measured slice of traffic to a candidate model.
Launch Readiness
Readiness, audit-chain status, and enabled command modules.
Connect to load deployment status
Policy Forge Records
Backed by /v1/admin/routing-policies.
| Agent | Purpose | Forced | Rules |
|---|---|---|---|
| Connect to load policies | |||
Provider Vault Records
Provider key references, rotation, allowed apps/environments, and fallback order.
| Provider | Secret ref | Rotation | Scope | Order |
|---|---|---|---|---|
| Connect to load credentials | ||||
Canary Lab Records
Candidate rollout status by agent and purpose.
| Workload | Candidate | Traffic | Status |
|---|---|---|---|
| Connect to load canaries | |||
Evaluation Center
LLM-as-judge, persisted batch history, calibrated judges, quality trend, and judge fallback.
Loading evaluation center
Topic / Pattern Clustering
Traffic clustered by topic, intent, workflow, failure, cost, quality, and security risk.
| Kind | Pattern | Requests | Cost | Action |
|---|---|---|---|---|
| Loading clusters | ||||
Applications
Superuser app drilldown pages at /dashboard/apps/{application_id}.
Connect to load applications
Requests
-
app traffic
Cost
-
budget usage
Latency
-
average response
Quality
-
judge average
Users
-
distinct metadata users
Risk
-
drift + security events
Cost / User
-
period comparison
Workflow Success
-
completed requests
Quality Bar
-
responses meeting target
Models & Providers
Model and provider mix for this app.
| Name | Kind | Requests |
|---|---|---|
| Select an app | ||
Prompts & Routing
Prompt versions and selected route outcomes.
| Name | Kind | Requests |
|---|---|---|
| Select an app | ||
Routing Intelligence
Deployment health, fallback rate, decision latency, and route quality.
| Deployment | Health | Traffic | Fallback | Quality |
|---|---|---|---|---|
| Select an app | ||||
Spend Intelligence
Burn rate, projected budget utilization, cost attribution, and savings.
Select an app
What Changed
Current period compared with the preceding period.
Select an app
Business Use Cases
Traffic translated from technical purposes into operating workflows.
| Use case | Purposes | Requests | Cost | Quality |
|---|---|---|---|---|
| Select an app | ||||
Request Log Explorer
Helicone / Datadog-style request table with filters, search, and CSV export.
| Timestamp | App | Agent | User | Model | Provider | Cost | Tokens | Latency | Status | Cache | Routing | Prompt |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Connect to load request logs | ||||||||||||
Trace Explorer
Conversation episodes grouped by root request.
| Root | Conversation | Spans | Kinds |
|---|---|---|---|
| Select an app | |||
Visual Trace
Agent steps, tool calls, provider calls, judge calls, and remediation actions.
Select a trace
Multi-agent execution trees with rolled-up cost, quality, risk, and drift.
Each simulator workflow appears as a parent-child span tree so agent steps, tool calls, governance events, and remediation actions stay inspectable as one trace.
Trace Trees
-
multi-agent roots
Span Nodes
-
rolled up branches
Agents
-
seen in traces
Max Depth
-
deepest workflow
Workflow Roots
Simulator and live A2A requests grouped by root request.
| Root | Agents | Spans | Depth |
|---|---|---|---|
| Connect to load multi-agent traces | |||
Trace Rollup
Expandable span tree with branch totals.
Select a trace
Enterprise AI health, spend, and control in one board-ready view.
Executive View summarises adoption, cost avoided, risk blocked, model quality, and audit readiness without requiring teams to inspect every route or trace.
Spend Avoided
-
routing + cache savings
Risk Blocked
-
privacy + policy events
Audit Readiness
-
evidence chain coverage
Model Opportunity
-
recommended migration
Function Adoption
Which business functions are consuming controlled AI.
| Function | Requests | Spend | Status |
|---|---|---|---|
| Loading executive view | |||
Priority Actions
Highest-value moves for the next operating review.
Loading actions