Observability Metrics

Archestra serves Prometheus metrics in OpenMetrics format. Scrape them to chart and alert on LLM cost and latency, MCP tool calls, Agent Runtime health, and background jobs.

Scraping

Each process serves its own registry, so scrape every pod:

ProcessEndpointEmits
Webhttp://<pod>:9050/metricsRequests served by the web tier: LLM, MCP, chat, Agent Runtime, HTTP
Workerhttp://<pod>:9000/metricsKnowledge Base syncs and embeddings, the task queue, scheduled tasks, and the LLM and MCP calls they make

The Helm chart runs the worker as its own Deployment by default. Sum counters across web and worker pods: an agent run from a schedule reports its LLM metrics from a worker pod.

ARCHESTRA_METRICS_PORT changes the web port. Worker pods serve metrics on their API port, 9000.

Authentication

The endpoint is open by default. Set ARCHESTRA_METRICS_SECRET to require a bearer token on web and worker pods. Requests without it get 401.

yaml
scrape_configs:
  - job_name: archestra
    authorization:
      type: Bearer
      credentials: <ARCHESTRA_METRICS_SECRET>
    kubernetes_sd_configs:
      - role: pod

Helm Pod Annotations

For a Prometheus that discovers pods by annotation, add these to the chart values:

yaml
archestra:
  podAnnotations:
    prometheus.io/scrape: "true"
    prometheus.io/port: "9050"
    prometheus.io/path: /metrics

The chart rewrites prometheus.io/port to 9000 on worker pods.

To check a pod, run curl -s http://<pod>:9050/metrics | grep "# HELP llm_". Add -H "Authorization: Bearer <secret>" when a secret is set.

Labels

Most LLM metrics share these labels:

LabelValues
provider, modelThe LLM provider and model, for example anthropic and claude-sonnet-4-5
agent_id, agent_nameThe Archestra agent that made the call. All LLM Proxy traffic reports agent_name="LLM Proxy". Knowledge Base calls report agent_name="Knowledge Base" with an empty agent_id.
agent_typeagent, llm_proxy, mcp_gateway, or profile
organization_idThe owning organization
sourceWhere the request came from, for example api, chat, chatops:slack, email, schedule-trigger, app:llm_complete, knowledge:embedding, or knowledge:reranker

MCP tool call metrics carry agent_id, agent_name, agent_type, mcp_server_name, tool_name, and status (success or error).

Agent Labels

Go to an agent's Advanced tab and add a key and value under Labels. Each agent label becomes an extra label on the LLM, MCP, and agent_runs_total metrics. Prometheus allows only letters, digits, and underscores in label names, so other characters become _: cost-center turns into cost_center.

Dimensions Without Labels

Users, skills, apps, and client-supplied IDs are not metric labels. Use these sources instead:

  • Users: per-user statistics, or the archestra.user.* span attributes.
  • Skills and apps: per-skill and per-app statistics. App LLM spend shows in the LLM metrics as source="app:llm_complete".
  • Client-supplied agent IDs from the X-Archestra-Agent-Id header: the archestra.external_agent_id span attribute. Only agent_runs_total carries it as a label.
  • Credentials: the archestra.virtual_key.id span attribute.

LLM Metrics

MetricTypeExtra LabelsMeasures
llm_request_duration_secondsHistogramstatus_codeLLM request duration
llm_tokens_totalCountertype (input, output)Tokens used. input excludes prompt-cache tokens.
llm_cache_tokens_totalCountercache_type (read, write)Prompt-cache tokens read from or written to the provider cache
llm_token_usageHistogramInput plus output tokens per request
llm_cost_totalCounterbilling_mode, auth_methodEstimated list-price cost in USD. Needs model pricing.
llm_cache_cost_totalCounterCost in USD of prompt-cache reads and writes
llm_cache_savings_totalCounterUSD saved by cache reads at the discounted price
llm_time_to_first_token_secondsHistogramStreaming: from the upstream call to the provider's first chunk
llm_time_to_first_byte_secondsHistogramStreaming: from request receipt to the first byte sent to the client
llm_tokens_per_secondHistogramOutput token throughput
llm_blocked_tools_totalCounterTool calls blocked by tool policies
agent_runs_totalCounterexternal_agent_idUnique runs, counted by the X-Archestra-Run-Id header
llm_active_usersGaugewindow (24h, 7d) onlyDistinct users with at least one LLM request in the window

billing_mode is metered for per-token billing and subscription for flat-rate subscription credentials. Billed spend is sum(llm_cost_total{billing_mode="metered"}). auth_method is provider_key, virtual_key, passthrough_virtual_key, jwks, oauth_client_credentials, oauth_user, internal, or unknown.

llm_time_to_first_byte_seconds minus llm_time_to_first_token_seconds is the time Archestra spends before the upstream call: authentication, guardrails, and tool policies.

Every replica reports the same llm_active_users value, so aggregate it with max(), not sum(). ARCHESTRA_METRICS_ACTIVE_USERS_REFRESH_INTERVAL_MS sets the refresh interval; 0 turns the metric off.

LLM and MCP metrics carry trace exemplars, so a Grafana panel can link a data point to its trace. See Exemplars.

MCP Metrics

MetricTypeMeasures
mcp_tool_calls_totalCounterTool calls through the MCP Gateway
mcp_tool_call_duration_secondsHistogramTool call duration
mcp_request_size_bytesHistogramTool call argument size
mcp_response_size_bytesHistogramTool call result size
mcp_server_deployment_statusGaugeState of each self-hosted MCP server, by server_name and state

mcp_server_deployment_status is 1 for the server's current state: not_created, pending, running, failed, succeeded, hibernated, or waking. count(mcp_server_deployment_status{state="running"} == 1) counts running servers.

Agent Runtime Health

MetricTypeLabelsMeasures
agent_runtime_runs_started_totalCounterAgent Runtime runs that reached a running backend
agent_runtime_runs_terminated_totalCounteroutcomeFinished runs: completed, failed, stopped_by_user, expired_ttl, or expired_idle
agent_runtime_provision_duration_secondsHistogramStartup time, including scheduling and image pulls
agent_runtime_steers_totalCountersteer_modeSteering messages delivered into running sessions
agent_runtime_completion_deliveries_totalCounterinterface, outcomeFinal replies delivered to chatops or email
agent_runtime_health_tasksGaugeagent_id, backend, conditionCurrent task counts per condition
agent_runtime_health_age_secondsGaugeagent_id, backend, conditionAge of the oldest task per condition
agent_runtime_health_collection_timestamp_secondsGaugeTime of the last health snapshot

agent_runtime_health_tasks conditions are working, submitted, input_required, auth_required, failed_recent (the last 15 minutes), and completion_pending. Conditions overlap: a task waiting for authentication also counts as working. agent_runtime_health_age_seconds conditions are heartbeat, submitted, and completion_pending.

Every replica reports the same health gauge values, so aggregate them with max, not sum. Alert on failed scrapes too, because missing samples do not mean healthy runs. A fresh heartbeat shows the orchestration is alive, not that the agent is making progress.

This alert finds runs whose heartbeat is more than two minutes old:

promql
max by (agent_id, backend) (
  agent_runtime_health_age_seconds{condition="heartbeat"}
) > 120

Add a pending period (for: 5m) to skip transient spikes.

Knowledge Base Metrics

MetricLabelsMeasures
rag_connector_syncs_total, rag_connector_sync_duration_secondsconnector_type, statusConnector syncs (success, failed, or partial) and their duration
rag_documents_processed_total, rag_documents_ingested_total, rag_chunks_created_totalconnector_typeDocuments read, documents added or updated, and chunks created
rag_documents_without_text_totalconnector_typeDocuments skipped because they have no extractable text
rag_ocr_pages_totalconnector_type, outcomeScanned PDF pages sent to OCR
rag_embedding_batches_total, rag_embedding_documents_totalstatusEmbedding batches and documents
rag_queries_total, rag_query_duration_seconds, rag_query_results_countsearch_typeSearches, their end-to-end duration, and results returned
rag_search_lane_timeout_totallaneSearch lanes cut by the database statement timeout
rag_quote_verification_totalresultChat-answer quotes checked against their cited source
rag_permission_syncs_totalconnector_type, statusPermission sync passes
rag_permission_sync_*connector_typePermission sync gaps: group failures, dropped principals, ACL over-approximations, unreadable containers, restriction fallbacks, and skipped identity lookups
rag_access_token_truncations_totalkindUsers whose group memberships exceeded the per-user cap at query time
rag_knowledge_query_unresolved_identity_totalSearches limited to organization-wide documents because the caller had no email

Skill and Sandbox Metrics

MetricLabelsMeasures
skill_activations_totalactivation_typeSkill activations: slash_command, chat_attachment, load_skill, or delegation
skill_context_tokens_totalactivation_typeTokens that skill activations added to model context
sandbox_commands_total, sandbox_command_duration_secondsstatusCode Sandbox commands: ok, script_error, timeout, or runtime_error
sandbox_runtime_errors_totalcodeSandbox runtime errors. engine_unreachable means the sandbox engine is down.
sandbox_runtime_statusstatus1 for the runtime's current status: disabled, initializing, ready, error, or stopped

Background Task Metrics

MetricLabelsMeasures
task_queue_tasks_enqueued_total, task_queue_tasks_completed_totaltask_typeTasks queued and completed
task_queue_tasks_failed_totaltask_typeFailed attempts. The task may be retried.
task_queue_tasks_dead_totaltask_typeTasks that ran out of retries
task_queue_task_duration_secondstask_typeTask duration
task_queue_active_taskstask_typeTasks running now
task_queue_stuck_tasks_reset_totalStuck tasks returned to the queue
schedule_trigger_runs_totalagent_name, statusScheduled task runs: success, failed, or cancelled

Common task_type values are connector_sync, batch_embedding, permission_sync, and schedule_trigger_run_execute.

Platform Metrics

MetricLabelsMeasures
http_request_duration_seconds, http_request_summary_secondsmethod, route, status_codeAPI request duration, as a histogram and as a summary with quantiles
database_pool_connectionsstateDatabase pool connections: total, idle, and waiting
database_pool_size_limitMaximum pool size per process
audit_write_failures_totalsource, resource_typeAudit log rows that failed to save
file_storage_orphaned_objects_totalprovider, scopeStored files left behind by a failed delete
chat_message_feedback_totalfeedbackThumbs up, thumbs down, and cleared ratings on chat answers
process_*, nodejs_*Standard Node.js process metrics: CPU, memory, heap, event loop lag, and garbage collection

Example Queries

promql
# Billed LLM spend per agent over the last day
sum by (agent_name) (increase(llm_cost_total{billing_mode="metered"}[1d]))

# 95th percentile LLM latency per model
histogram_quantile(0.95, sum by (le, model) (rate(llm_request_duration_seconds_bucket[5m])))

# MCP tool error rate per server
sum by (mcp_server_name) (rate(mcp_tool_calls_total{status="error"}[5m]))
  / sum by (mcp_server_name) (rate(mcp_tool_calls_total[5m]))