Observability & Alerting
Structured logging, Prometheus metrics, distributed tracing, and threshold-based anomaly alerting — built into the governance pipeline.
Three Layers
GaaS observability is designed around three layers that work independently. You can run with just structured logs, add Prometheus scraping, and optionally enable OpenTelemetry tracing — each layer requires zero code changes.
Structured Logging
JSON-formatted structured logs via structlog. Every log entry carries a request_id for correlation across the pipeline. Set GAAS_LOG_LEVEL to control verbosity.
Prometheus Metrics
Scrape GET /metrics for HTTP request counts, latency histograms, pipeline stage durations, decision verdicts, error rates, circuit breaker state, and more. No authentication required on this endpoint.
Distributed Tracing
Optional OpenTelemetry integration. Set GAAS_TRACING_ENABLED=true and GAAS_OTEL_ENDPOINT to send spans for every pipeline stage to your collector. Zero overhead when disabled.
Anomaly Alerting
Beyond passive metrics, GaaS actively monitors the health of its own platform and raises an alert when a threshold is exceeded. The anomaly monitor evaluates metrics on every governance decision and through a periodic background check (every 60 seconds), and reports each alert to the GaaS operations team. These are platform-health alerts, not alerts about one organization: your organization is notified by email, automatically, when one of its actions is blocked or an escalation is waiting for review.
Event Types
| Event | Trigger | Severity |
|---|---|---|
circuit_breaker_tripped |
Learning calibration circuit breaker freezes due to quality degradation | Critical |
high_error_rate |
More than 10 application errors in a 5-minute window | Critical |
block_rate_anomaly |
Block rate exceeds 50% over recent decisions (min 10 sample size) | Warning |
pipeline_degraded |
Any single pipeline execution exceeds 5,000ms | Warning |
rate_limit_spike |
More than 50 rate-limit rejections in a 5-minute window | Warning |
escalation_queue_growth |
Pending escalation count ≥ 10 and growing (checked every 60s) | Warning |
Metrics History
In addition to real-time Prometheus metrics, your organization's governance metrics are available for each past hour, computed from its own decisions and escalations:
GET /v1/dashboard/metrics/history?hours=24&limit=100
Returns one entry for each UTC hour that had decisions or escalations, oldest first: decision counts by verdict, block rate, average risk, deliberation rate, latency percentiles (p50, p99), escalations created, escalations awaiting review at the end of the hour, and the override and false-block rates over the 7 days up to it. Shadow-mode decisions are not counted.
The window is the last 30 days unless you set it: hours (1–8760) counts back from end, or give start and end as ISO 8601 date-times. limit (default 720) and offset page through the hours. Only your organization's data is included.
Prometheus Metrics Reference
| Metric | Type | Labels |
|---|---|---|
gaas_http_requests_total |
Counter | method, path_template, status_code |
gaas_http_request_duration_seconds |
Histogram | method, path_template |
gaas_active_requests |
Gauge | — |
gaas_pipeline_stage_duration_seconds |
Histogram | stage |
gaas_pipeline_total_duration_seconds |
Histogram | — |
gaas_decisions_total |
Counter | verdict, pipeline_mode |
gaas_errors_total |
Counter | error_code |
gaas_rate_limit_hits_total |
Counter | scope |
gaas_circuit_breaker_state |
Gauge | breaker_name (0=normal, 1=frozen) |
gaas_escalations_created_total |
Counter | — |
gaas_deliberation_triggered_total |
Counter | — |
Environment Variables
| Variable | Default | Description |
|---|---|---|
GAAS_LOG_LEVEL |
info |
Structlog level (debug, info, warning, error) |
GAAS_TRACING_ENABLED |
false |
Enable OpenTelemetry tracing |
GAAS_OTEL_ENDPOINT |
localhost:4317 |
OTLP gRPC collector endpoint |
GAAS_PAGERDUTY_ROUTING_KEY |
— | PagerDuty Events API v2 routing key for sub-processor change notices |
Related Pages
- Conversational Dashboard — review decisions, escalations and agent activity through natural language
- Getting Started — onboard, integrate, and go live
- Intent Declaration API — the pipeline that generates the metrics